Scalable System-on-Chip

The scalable SOC design with an integrated memory architecture addresses the lack of design reuse in traditional SOC designs by allowing easy scaling and software compatibility across varying complexity levels, enhancing efficiency and reducing redundant efforts.

JP2026048651APending Publication Date: 2026-03-17APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Traditional system-on-chip (SOC) designs are not scalable and require redundant design efforts for different applications, lacking opportunities for design reuse and software compatibility across varying complexity levels.

Method used

A scalable SOC design with an integrated memory architecture that allows heterogeneous agents to share a unified address space, enabling easy scaling of complexity and facilitating software compatibility across different resource versions.

Benefits of technology

Enables efficient integration and scaling of SOC designs from small to large applications, reducing redundant design efforts and ensuring software compatibility and consistent functionality across different hardware configurations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026048651000001_ABST
    Figure 2026048651000001_ABST
Patent Text Reader

Abstract

The system provides a scalable system-on-a-chip with integrated memory that can access heterogeneous agents within the system. [Solution] A system-on-a-chip (SoC) 10 comprising a plurality of processor cores 16 including a general-purpose processor and a graphics processing unit (GPU), a plurality of peripheral devices 20A to 20p, one or more memory controllers 22A to 22m, and an interconnect fabric 28, wherein the interconnect fabric includes at least two networks having heterogeneous operating characteristics, and these are integrated onto two or more packaged dies to support scaling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments described herein relate to digital systems, and more particularly, to systems having an integrated memory accessible to heterogeneous agents within the system.

Background Art

[0002] In the design of modern computing systems, it has become increasingly common to integrate a wide variety of system hardware components that were previously implemented as individual silicon components onto a single silicon die. For example, a complete computer system may have once included an individually packaged microprocessor mounted on a backplane and coupled to a chipset that interfaces the microprocessor to other devices such as system memory, a graphics processor, and other peripheral devices. In contrast, the development of semiconductor process technology has enabled the integration of many of these discrete devices. The result of such integration is commonly referred to as a "system-on-chip" (SOC).

[0003] Traditionally, SOCs for different applications have been built, designed, and implemented individually. For example, an SOC for a smartwatch device may have stringent power consumption requirements because the form factor of such a device limits the available battery size and therefore the maximum usage time of the device. At the same time, the small size of such a device may limit the number of peripherals that the SOC needs to support, as well as the computing requirements of the application that the SOC runs on. In contrast, an SOC for a mobile phone application has a larger available battery and therefore a larger power budget, but is also expected to have more complex peripherals and greater graphics and general computing requirements. Therefore, such an SOC is expected to be larger and more complex than designs for smaller devices. This comparison can be arbitrarily extended to other applications. For example, wearable computing solutions such as augmented and / or virtual reality systems may be expected to present greater computing requirements than devices for less complex devices, as well as devices for desktop and / or rack-mount computer systems.

[0004] Traditional, individually architectured approaches to SOCs offer little opportunity for design reuse, and design effort is duplicated across multiple SOC implementations.

[0005] For a detailed explanation, please refer to the attached diagrams briefly described below. [Brief explanation of the drawing]

[0006] [Figure 1] This is a block diagram of one embodiment of a system-on-a-chip (SOC).

[0007] [Figure 2] This is a block diagram of a system including one embodiment of multiple networks that interconnect agents.

[0008] [Figure 3] This is a block diagram of one embodiment of a network that uses a ring topology.

[0009] [Figure 4] This is a block diagram of one embodiment of a network that uses a mesh topology.

[0010] [Figure 5] This is a block diagram of one embodiment of a network that uses a tree topology.

[0011] [Figure 6] This is a block diagram of one embodiment of a system-on-a-chip (SOC) having multiple networks for one embodiment.

[0012] [Figure 7] This is a block diagram of one embodiment of a system-on-a-chip (SOC), showing one of the independent networks shown in Figure 6 for one embodiment.

[0013] [Figure 8] This is a block diagram of one embodiment of a system-on-a-chip (SOC), showing another of the independent networks shown in Figure 6 for one embodiment.

[0014] [Figure 9] This is a block diagram of one embodiment of a system-on-a-chip (SOC), showing yet another independent network shown in Figure 6 for one embodiment.

[0015] [Figure 10] This is a block diagram of one embodiment of a multi-die system including two semiconductor dies.

[0016] [Figure 11] This is a block diagram of one embodiment of an input / output (I / O) cluster.

[0017] [Figure 12] A block diagram of one embodiment of a processor cluster.

[0018] [Figure 13] A pair of tables showing virtual channels, traffic types, and networks shown in FIGS. 6-9, used for one embodiment.

[0019] [Figure 14] A flowchart showing one embodiment of starting a transaction on a network. <s|

[0020] [Figure 15] A block diagram of one embodiment of a system including an interrupt controller and a plurality of cluster interrupt controllers corresponding to a plurality of clusters of processors.

[0021] <s| [Figure 16] A block diagram of one embodiment of a system-on-chip (SOC) that can implement one embodiment of the system shown in FIG. 15.

[0022] [Figure 17] A block diagram of one embodiment of a state machine that can be implemented in one embodiment of an interrupt controller.

[0023] [Figure 18] A flowchart showing the operation of one embodiment of an interrupt controller that performs a soft or hard iteration of interrupt distribution.

[0024] [Figure 19] A flowchart showing the operation of one embodiment of a cluster interrupt controller.

[0025] [Figure 20] A block diagram of one embodiment of a processor.

[0026] [Figure 21] This is a block diagram of one embodiment of a reorder buffer.

[0027] [Figure 22] Figure 20 is a flowchart illustrating the operation of one embodiment of the interrupt acknowledgment control circuit.

[0028] [Figure 23] This is a block diagram of multiple SOCs that can implement one embodiment of the system shown in Figure 15.

[0029] [Figure 24] Figure 23 is a flowchart illustrating the operation of one embodiment of the primary interrupt controller.

[0030] [Figure 25] Figure 23 is a flowchart showing the operation of one embodiment of the secondary interrupt controller.

[0031] [Figure 26] This is a flowchart showing one embodiment of a method for handling interrupts.

[0032] [Figure 27] This is a block diagram of one embodiment of a cache-coherent system implemented as a system-on-a-chip (SOC).

[0033] [Figure 28] This block diagram shows one embodiment of a three-hop protocol for coherent transfer of cache blocks.

[0034] [Figure 29] This block diagram shows one embodiment of managing competition between a fill of one coherent transaction and a snoop of another coherent transaction.

[0035] [Figure 30]This block diagram shows one embodiment of managing competition between snooping one coherent transaction and acknowledging another coherent transaction.

[0036] [Figure 31] This is a block diagram of a part of one embodiment of a coherent agent.

[0037] [Figure 32] This is a flowchart showing the operation of one embodiment of how a coherence controller processes requests.

[0038] [Figure 33] This flowchart illustrates the operation of one embodiment of a coherent agent that transmits a request to a memory controller to process the completion associated with the request.

[0039] [Figure 34] This is a flowchart illustrating the operation of one embodiment of a coherent agent that receives snoops.

[0040] [Figure 35] A block diagram showing a chain of competing requests to a cache block according to one embodiment.

[0041] [Figure 36] This is a flowchart showing one embodiment of a coherent agent that absorbs Snoop.

[0042] [Figure 37] This is a block diagram showing one embodiment of a non-cacheable request.

[0043] [Figure 38] This flowchart illustrates the operation of one embodiment of a coherence controller that generates snoops based on the cacheable and non-cacheable properties of a request.

[0044] [Figure 39] This table shows multiple cache states according to one embodiment of the coherence protocol.

[0045] [Figure 40] This is a table showing several messages that may be used in one embodiment of the coherence protocol.

[0046] [Figure 41] This flowchart illustrates the operation of one embodiment of a coherence controller that handles changes to exclusive conditional requests.

[0047] [Figure 42] This flowchart illustrates the operation of one embodiment of a coherence controller that reads directory entries and generates snoops.

[0048] [Figure 43] This flowchart shows the operation of one embodiment of a coherence controller that handles exclusive no-data requests.

[0049] [Figure 44] A block diagram shows exemplary elements of a system-on-a-chip according to several embodiments.

[0050] [Figure 45] A block diagram illustrating exemplary elements of the interaction between an I / O agent and a memory controller in several embodiments.

[0051] [Figure 46A] A block diagram illustrating the elements of an I / O agent configured to handle write transactions, according to several embodiments.

[0052] [Figure 46B]A block diagram illustrating the elements of an I / O agent configured to process read transactions, according to several embodiments.

[0053] [Figure 47] This flowchart illustrates an example of processing read transaction requests from peripheral components according to several embodiments.

[0054] [Figure 48] This flowchart illustrates an exemplary method for processing read transaction requests by an I / O agent, according to several embodiments.

[0055] [Figure 49] A block diagram of one embodiment of a system having two interconnected integrated circuits is shown.

[0056] [Figure 50] A block diagram of one embodiment of an integrated circuit having an external interface is shown.

[0057] [Figure 51] The block diagram shows a system with two integrated circuits that utilize interface wrappers to route the pin assignments of their respective external interfaces.

[0058] [Figure 52] A block diagram of one embodiment of an integrated circuit having an external interface that utilizes pin bundles is shown.

[0059] [Figure 53A] Two examples of two integrated circuits coupled together using complementary interfaces are shown.

[0060] [Figure 53B] Two additional examples of two integrated circuits coupled together are shown.

[0061] [Figure 54] A flowchart of one embodiment of a method for transferring data between two coupled integrated circuits is shown.

[0062] [Figure 55] This diagram shows a flowchart of one embodiment of a method for routing signal data between an external interface within an integrated circuit and an on-chip router.

[0063] [Figure 56] This is a block diagram of one embodiment of a plurality of system-on-a-chip (SOCs), where a given SOC includes a plurality of memory controllers.

[0064] [Figure 57] This is a block diagram showing one embodiment of the memory controller and its physical / logical arrangement on an SOC.

[0065] [Figure 58] This is a block diagram of one embodiment of a binary decision tree for determining which memory controller serves a specific address.

[0066] [Figure 59] This is a block diagram showing one embodiment of multiple memory location configuration registers.

[0067] [Figure 60] This is a flowchart showing the operation of one embodiment of an SOC during boot / power-on.

[0068] [Figure 61] This is a flowchart illustrating the operation of one embodiment of a SoC for routing memory requests.

[0069] [Figure 62] This is a flowchart showing the operation of one embodiment of a memory controller in response to a memory request.

[0070] [Figure 63] This is a flowchart illustrating the operation of one embodiment of a monitoring system for determining memory folding.

[0071] [Figure 64] This is a flowchart showing the operation of one embodiment of folding a memory slice.

[0072] [Figure 65] This is a flowchart showing the operation of one embodiment of unpacking a memory slice.

[0073] [Figure 66] This is a flowchart showing one embodiment of a memory folding method.

[0074] [Figure 67] This is a flowchart illustrating one embodiment of a method for hashing memory addresses.

[0075] [Figure 68] This is a flowchart showing one embodiment of a method for forming a compressed pipe address.

[0076] [Figure 69] This is a block diagram of one embodiment of an integrated circuit design that supports both whole and partial instances.

[0077] [Figure 70] Figure 69 shows various embodiments of the whole instance and partial instance of the integrated circuit. [Figure 71] Figure 69 shows various embodiments of the whole instance and partial instance of the integrated circuit. [Figure 72] Figure 69 shows various embodiments of the whole instance and partial instance of the integrated circuit.

[0078] [Figure 73]Figure 69 is a block diagram of one embodiment of an integrated circuit, in which each sub-region of the integrated circuit has a local clock source.

[0079] [Figure 74] Figure 69 is a block diagram of one embodiment of an integrated circuit, which has local analog pads in each sub-region of the integrated circuit.

[0080] [Figure 75] Figure 69 is a block diagram of one embodiment of an integrated circuit, in which there are block-out regions at the corners of each sub-region, and regions for interconnect "bumps" that exclude the areas near the edges of each sub-region.

[0081] [Figure 76] This is a block diagram showing one embodiment of a stub and a corresponding circuit component.

[0082] [Figure 77] A block diagram showing one embodiment of a pair of integrated circuits and specific further details of the pair of integrated circuits.

[0083] [Figure 78] This is a flowchart illustrating one embodiment of an integrated circuit design method.

[0084] [Figure 79] This is a block diagram showing the testbench layout for testing the entire instance and partial instances.

[0085] [Figure 80] This is a block diagram showing the test bench layout for component-level testing.

[0086] [Figure 81] This is a flowchart illustrating one embodiment of a method for designing and manufacturing integrated circuits.

[0087] [Figure 82]This flowchart shows one embodiment of a method for manufacturing an integrated circuit.

[0088] [Figure 83] This is a block diagram of one embodiment of the system.

[0089] [Figure 84] This is a block diagram of one embodiment of a computer-accessible storage medium.

[0090] While the embodiments described in this disclosure may be subject to various modifications and alternative forms, specific embodiments are shown in the drawings as examples and described in detail herein. However, it should be understood that the drawings and the detailed description relating to them are not intended to limit the embodiments to any particular form disclosed, but rather to cover all modifications, equivalents, and alternative forms that fall within the spirit and scope of the appended claims. The titles used herein are for structural purposes only and are not intended to limit the scope of the description. [Modes for carrying out the invention]

[0091] A System of Computer (SOC) may include most of the elements necessary to implement a complete computer system, although some elements (e.g., system memory) may be located outside the SOC. For example, an SOC may include one or more general-purpose processor cores, one or more graphics processing units, and one or more other peripheral devices (such as application-specific accelerators, I / O interfaces, or other types of devices) separate from the processor cores and graphics processing units. The SOC may further include one or more memory controller circuits configured to interface with system memory, and an interconnect fabric configured to provide communication between the memory controller circuits (one or more), the processor cores (one or more), the graphics processing units (one or more), and the peripheral devices (one or more).

[0092] The design requirements of a given SOC are often determined by the power and performance requirements of the specific application the SOC targets. For example, an SOC for a smartwatch device may have stringent power consumption requirements because the form factor of such a device limits the available battery size and therefore the maximum usage time of the device. At the same time, the small size of such a device may limit the number of peripherals that the SOC needs to support, as well as the computational requirements of the application that the SOC runs on. In contrast, an SOC for a mobile phone application has a larger available battery and therefore a larger power budget, but is also expected to have more complex peripherals and greater graphics and general computing requirements. Therefore, such an SOC is expected to be larger and more complex than a design for a smaller device.

[0093] This comparison can be arbitrarily extended to other applications. For example, wearable computing solutions such as augmented and / or virtual reality systems may be expected to present greater computing requirements than less complex devices, as well as devices for desktop and / or rack-mount computer systems.

[0094] As systems are built for larger applications, multiple chips can be used together to scale performance and form a “system of chips.” In this specification, whether these systems consist of a single physical chip or multiple physical chips, these systems will continue to be referred to as a “SOC.” The principles in this disclosure are equally applicable to multi-chip SOCs and single-chip SOCs.

[0095] The inventors' insight in this disclosure is that the computational requirements and corresponding SOC complexity for the various applications described above tend to scale from small to large. If an SOC can be designed to easily scale in physical complexity, the core SOC design can be easily adapted for a wide variety of applications, leveraging design reuse and reducing redundant effort. Such an SOC also provides a consistent view of functional blocks, such as processing cores or media blocks, making their integration into the SOC easier and adding further effort reduction. That is, the same functional block (or "IP") design can be used in SOCs from small to large without essentially modifying it. Furthermore, if such an SOC design can scale in a manner that is largely or completely transparent to the software running on the SOC, the development of software applications that can easily scale across different resource versions of the SOC becomes significantly simplified. The application can be written only once and automatically and correctly function on many different systems, also from small to large. When the same software scales across different resource versions, the software provides the same interface to the user, which is another advantage of scaling.

[0096] This disclosure envisions such a scalable SOC design. In particular, the core SOC design may include a set of processor cores, graphics processing units, memory controller circuits, peripheral devices, and an interconnect fabric configured to interconnect them. Furthermore, the processor cores, graphics processing units, and peripheral devices may be configured to access system memory via an integrated memory architecture. The integrated memory architecture includes an integrated address space, which enables heterogeneous agents in the system (processors, graphics processing units, peripherals, etc.) to cooperate simply and efficiently. That is, rather than allocating a private address space to the graphics processing unit and requiring data to be copied to and from that private address space, the graphics processing unit, processor cores, and other peripheral devices can, in principle, share access to any memory address accessible by the memory controller circuit (in some embodiments, they are subject to a privilege model or other security features that restrict access to certain types of memory content). In addition, the integrated memory architecture provides the same memory semantics (e.g., a common set of memory semantics) as the complexity of the SOC scales to meet the requirements of different systems. For example, memory semantics may include memory ordering properties, quality of service (QoS) support and attributes, memory management unit definitions, cache coherency functions, etc. The integrated address space may be a physical address space distinct from the virtual address space, or it may be a physical address space, or both.

[0097] The architecture remains the same as that which scales the SOC, but various implementation options are available. For example, virtual channels can be used as part of QoS support, but if not all QoS is guaranteed in a given system, a subset of supported virtual channels may be implemented. Different interconnect fabric implementations may be used depending on the bandwidth and latency characteristics required in a given system. In addition, some features may not be necessary in smaller systems (for example, address hashing to balance memory traffic to various memory controllers may not be necessary in a single-memory-controller system). The hash algorithm may not be important when there are a few memory controllers (e.g., two or four), but it contributes more significantly to system performance when a larger number of memory controllers are used.

[0098] Furthermore, some components may be designed with scalability in mind. For example, a memory controller may be designed to scale up by adding additional memory controllers to the fabric, each having its own address space, memory cache, and coherency tracking logic.

[0099] More specifically, embodiments of SOC designs are disclosed that can be easily scaled down and up in terms of complexity. For example, in an SOC, the processor cores, graphics processing units, fabric, and other devices can be arranged and configured so that the size and complexity of the SOC can be easily reduced before manufacturing by "shearing" the SOC along defined axes such that the resulting design contains only a subset of the components defined in the original design. Alternatively, when buses extending from the removed portion of the SOC are properly terminated, a reduced-complexity version of the original SOC design can be obtained with relatively little design and verification effort. An integrated memory architecture can facilitate the deployment of applications in the reduced-complexity design, and in some cases, can easily operate without substantial modification.

[0100] As described above, the disclosed embodiments of the SOC design can be configured to scale up complexity. For example, multiple instances of a single-die SOC design can be interconnected, resulting in a system with resources that are two, three, four, or more multiples greater than those of the single-die design. In this case as well, the integrated memory architecture and consistent SOC architecture can facilitate the development and deployment of software applications that scale to use the additional computing resources provided by these multi-die system configurations.

[0101] Figure 1 is a block diagram of one embodiment of a scalable SOC 10 coupled to one or more memories such as memories 12A to 12m. The SOC 10 may include a plurality of processor clusters 14A to 14n. Each processor cluster 14A to 14n may include one or more processors (P) 16 coupled to one or more caches (e.g., cache 18). The processors 16 may include general-purpose processors (e.g., central processing units or CPUs), as well as other types of processors such as graphics processing units (GPUs). The SOC 10 may include one or more other agents 20A to 20p. One or more other agents 20A to 20p may include, for example, a wide variety of peripheral circuits / devices, and / or bridges such as input / output agents (IOAs) coupled to one or more peripheral devices / circuits. The SOC 10 may include one or more memory controllers 22A to 22m, each coupled to a separate memory device or circuit 12A to 12m during use. In one embodiment, each memory controller 22A-22m may include a coherency controller circuit (more simply, a “coherency controller” or “CC”) coupled to a directory (the coherency controller and directory are not shown in Figure 1). Furthermore, a die-to-die (D2D) circuit 26 is shown within the SOC 10. The memory controllers 22A-22m, other agents 20A-20p, the D2D circuit 26, and processor clusters 14A-14n may be coupled to an interconnect 28 for communication between the various components 22A-22m, 20A-20p, 26, and 14A-14n. As indicated by their names, the components of the SOC 10 may, in one embodiment, be integrated on a single integrated circuit “chip”. In other embodiments, the various components may be external to the SOC 10 on other chips, or may be separate components. Any amount of integrated or separate components may be used. In one embodiment, a subset of processor clusters 14A to 14n and memory controllers 22A to 22m may be implemented on one of a plurality of integrated circuit chips coupled together to form the components shown in SOC 10 in Figure 1.

[0102] The D2D circuit 26 may be an off-chip interconnect, which may be coupled to an interconnect fabric 28 and configured to couple the interconnect fabric 28 to a corresponding interconnect fabric 28 on another instance of the SOC 10. The interconnect fabric 28 and the off-chip interconnect 26 provide an interface that transparently connects one or more memory controller circuits, processor cores, graphics processing units, and peripheral devices in either a single instance of an integrated circuit or two or more instances of an integrated circuit. That is, via the D2D circuit 26, the interconnect fabric 28 extends across two or more integrated circuit dies, and communication is routed between the source and destination transparently to the source and destination locations on the integrated circuit dies. The interconnect fabric 28 extends across two or more integrated circuit dies using hardware circuitry (e.g., the D2D circuit 26) to automatically route communication between the source and destination, regardless of whether the source and destination are on the same integrated circuit die.

[0103] Therefore, the D2D circuit 26 supports the scalability of SOC10 to two or more instances of SOC10 in the system. When two or more instances are included, the integrated memory architecture, including the unified address space, extends across two or more instances of integrated circuit dies, transparent to the software running on the processor cores, graphics processing units, or peripheral devices. Similarly, in the case of a single instance of an integrated circuit die in the system, the integrated memory architecture, including the unified address space, maps to a single instance, transparent to the software. When two or more instances of integrated circuit dies are included in the system, the system set of processor cores 16, graphics processing units, peripheral devices 20A-20p, and interconnect fabric 28 is distributed across two or more integrated circuit dies, also transparent to the software.

[0104] As described above, processor clusters 14A to 14n may include one or more processors 16. A processor 16 can function as the central processing unit (CPU) of the SOC 10. The system's CPU includes one or more processors that run the system's primary control software, such as the operating system. Generally, the software run by the CPU during use can control other components of the system to achieve desired functions of the system. A processor may also run other software, such as application programs. Application programs may provide user functions and may rely on the operating system for low-level device control, scheduling, memory management, etc. Therefore, a processor may also be referred to as an application processor. In addition, a processor 16 in a given cluster 14A to 14n may, as previously stated, be a GPU and may implement a graphics instruction set optimized for rendering, shading, and other operations. Clusters 14A to 14n may further include other hardware, such as a cache 18 and / or interfaces to other components of the system (e.g., an interface to the interconnect 28). Other coherent agents may include processors that are neither CPUs nor GPUs.

[0105] Generally, a processor may include any circuitry and / or microcode configured to execute instructions defined within the instruction set architecture implemented by the processor. A processor may also include a processor core implemented as a system-on-a-chip (SOC10) or at other integration levels on an integrated circuit along with other components. A processor may further include individual microprocessors, processor cores and / or microprocessors integrated into a multi-chip module implementation, processors implemented as multiple integrated circuits, and so on. The number of processors 16 in a given cluster 14A-14n may differ from the number of processors 16 in another cluster 14A-14n. Generally, one or more processors may be included. Furthermore, processors 16 may differ in terms of their microarchitecture implementation, performance and power characteristics, etc. In some cases, processors may also differ in terms of the instruction set architecture they implement, their functions (e.g., CPU, graphics processing unit (GPU) processor, microcontroller, digital signal processor, image signal processor, etc.).

[0106] The cache 18 can have any capacity and configuration, such as set-associative, direct-mapped, or fully associative. The cache block size may be any desired size (e.g., 32 bytes, 64 bytes, 128 bytes, etc.). A cache block can be a unit of allocation and deallocation in the cache 18. Furthermore, in this embodiment, a cache block can be a coherence-maintained address space (e.g., an aligned coherence-grained segment of a memory unit). A cache block may also be referred to as a cache line.

[0107] The memory controllers 22A-22m may generally include circuitry for receiving memory operations from other components of the SOC 10 and accessing memories 12A-12m to complete the memory operations. The memory controllers 22A-22m can be configured to access any type of memory 12A-12m. More specifically, memories 12A-12m can be any type of memory device that can be mapped as random access memory. For example, memories 12A-12m may be static random access memory (SRAM), double data rate (DRAM) such as synchronous DRAM (SDRAM) including dynamic RAM (DDR, DDR2, DDR3, DDR4, etc.), non-volatile memory, graphics DRAM such as graphics DDR DRAM (GDDR), and high-bandwidth memory (HBM). Low-power / mobile versions of DDR DRAM (e.g., LPDDR, mDDR, etc.) may be supported. The memory controllers 22A-22m may include queues for memory operations to order (and potentially reorder) operations and present them to the memories 12A-12m. The memory controllers 22A-22m may further include data buffers for storing write data awaiting writing to memory and read data awaiting the return of the memory operation to its source (if the data is not provided by the snoop). In some embodiments, the memory controllers 22A-22m may include a memory cache for storing recently accessed memory data. In SOC implementations, for example, the memory cache can reduce power consumption in the SOC by avoiding re-access of data from the memories 12A-12m when it is expected to be accessed again soon. In some cases, the memory cache may also be referred to as a system cache, in contrast to private caches such as cache 18 or a cache in the processor 16 that serve only specific components. Furthermore, in some embodiments, the system cache does not need to be located within the memory controllers 22A-22m.Therefore, there may be one or more levels of cache between the processor core, graphics processing unit, peripheral devices, and system memory. One or more memory controller circuits 22A to 22m may include their respective memory caches inserted between the interconnect fabric and system memory, each memory cache being one of the one or more levels of cache.

[0108] Other agents 20A-20p may generally include various additional hardware functions (e.g., “Peripherals,” “Peripheral Devices,” or “Peripheral Circuits”) included in the SOC C10. For example, Peripherals may include video peripherals such as image signal processors configured to process image acquisition data from cameras or other image sensors, video encoders / decoders, scalers, rotators, blenders, etc. Peripherals may include audio peripherals such as microphones, speakers, interfaces to microphones and speakers, audio processors, digital signal processors, mixers, etc. Peripherals may include interface controllers for various interfaces outside the SOC10, including interfaces such as Universal Serial Bus (USB), Peripheral Component Interconnect (PCI) including PCI Express (PCIe), serial and parallel ports. Peripherals may include networking peripherals such as Media Access Controllers (MACs). Any set of hardware may be included. Other agents 20A-20p may also, in one embodiment, include bridges to a set of peripherals such as IOAs described below. In one embodiment, the peripheral device includes one or more of the following: an audio processing device, a video processing device, a machine learning accelerator circuit, a matrix arithmetic accelerator circuit, a camera processing circuit, a display pipeline circuit, a non-volatile memory controller, a peripheral component interconnect controller, a security processor, or a serial bus controller.

[0109] The interconnect 28 may be any communication interconnect and protocol for communication between components of the SOC 10. The interconnect 28 may be bus-based, including a hierarchical bus with a shared bus configuration, a crossbar configuration, and bridges. The interconnect 28 may be packet-based or circuit-switched, and may be hierarchical with bridges, crossbars, point-to-point, or other interconnects. In one embodiment, the interconnect 28 may include a plurality of independent communication fabrics.

[0110] In one embodiment, if the system includes two or more instances of an integrated circuit die, the system may further include at least one interposer device configured to connect a bus of interconnect fabrics across the two or more integrated circuit dies. In one embodiment, a given integrated circuit die includes a power manager circuit configured to manage the local power state of the given integrated circuit die. In one embodiment, when the system includes two or more instances of an integrated circuit die, each individual power manager is configured to manage the local power state of the integrated circuit die, and at least one of the two or more integrated circuit dies includes another power manager circuit configured to synchronize the power manager circuits.

[0111] In general, the number of each component 22A-22m, 20A-20p, and 14A-14n may vary from embodiment to embodiment, and any number may be used. As indicated by the postfixes "m", "p", and "n", the number of one type of component may differ from the number of another type of component. However, the number of a given type may be the same as the number of other types. Furthermore, although the system in Figure 1 is shown with multiple memory controllers 22A-22m, embodiments having one memory controller 22A-22m are also conceivable.

[0112] The concept of scalable SOC design is easy to explain but difficult to implement. Numerous technological innovations have been developed in support of this effort, which are described in more detail below. In particular, Figures 2 to 14 include further details of embodiments of the communication fabric 28. Figures 15 to 26 show embodiments of the scalable interrupt structure. Figures 27 to 43 show embodiments of a scalable cache coherency mechanism that can be implemented among coherent agents in the system, including processor clusters 14A to 14n and one or more directory / coherency control circuits. In one embodiment, the directory and coherency control circuits are distributed among a plurality of memory controllers 22A to 22m, and each directory and coherency control circuit is configured to manage the cache coherency of a portion of the address space mapped to memory devices 12A to 12m to which a given memory controller is coupled. Figures 44 to 48 show embodiments of IOA bridges for one or more peripheral circuits. Figures 49 to 55 show further details of embodiments of the D2D circuit 26. Figures 56 to 68 show embodiments of a hashing scheme that distributes the address space across multiple memory controllers 22A to 22m. Figures 69 to 82 show embodiments of a design method that supports multiple tape-outs of a scalable SOC 10 for different systems, based on the same design database.

[0113] The various embodiments described below and the embodiments described above may be used in any desired combination to form embodiments of the present disclosure. Specifically, any subset of embodied features from any of the embodiments may be combined to form embodiments that include not all of the features described in any given embodiment and / or not all of the embodiments. All such embodiments are intended embodiments of the scalable SOC described herein. fabric

[0114] Figures 2 to 14 show various embodiments of the interconnect fabric 28. Based on this description, a system is envisioned comprising: a plurality of processor cores; a plurality of graphics processing units; a plurality of peripheral devices distinct from the processor cores and graphics processing units; one or more memory controller circuits configured to interface with system memory; one or more memory controller circuits; and an interconnect fabric configured to provide communication between the processor cores, graphics processing units, and peripheral devices, wherein the interconnect fabric includes at least two networks having heterogeneous operating characteristics. In one embodiment, the interconnect fabric includes at least two networks having heterogeneous interconnect topologies. The at least two networks may include coherent networks interconnecting the processor cores and one or more memory controller circuits. More specifically, the coherent networks interconnect coherent agents, and the processor cores may be coherent agents, or the processor cluster may be coherent agents. The at least two networks may include relaxed sequence networks coupled to the graphics processing units and one or more memory controller circuits. In one embodiment, the peripheral device includes a subset of devices, the subset including one or more machine learning accelerator circuits or relaxed sequence bulk media devices, and the relaxed sequence network is further coupled to the subset of devices to one or more memory controller circuits. At least two networks may include input / output networks coupled to interconnect the peripheral device and one or more memory controller circuits. The peripheral device includes one or more real-time devices.

[0115] In one embodiment, at least two networks include a first network that includes one or more characteristics for reducing latency compared to a second network of the at least two networks. For example, one or more characteristics may include shorter routes than the second network on the surface area of ​​the integrated circuit. One or more characteristics may include routing for the first interconnect in a metal layer that provides lower latency characteristics than routing for the second interconnect.

[0116] In one embodiment, at least two networks include a first network that includes one or more characteristics to increase bandwidth compared to a second network of the at least two networks. For example, one or more characteristics include wider interconnects compared to the second network. One or more characteristics include wiring in a metal layer further from the surface of the substrate on which the system is mounted than wiring for the second network.

[0117] In one embodiment, the interconnect topology employed by at least two networks includes at least one of a star topology, mesh topology, ring topology, tree topology, fat tree topology, hypercube topology, or a combination of one or more topologies. In another embodiment, at least two networks are physically and logically independent. In yet another embodiment, at least two networks are physically separated in a first operating mode, and the first network and the second network are virtual in a second operating mode and share a single physical network.

[0118] In one embodiment, the SOC is integrated on a semiconductor die. The SOC comprises a plurality of processor cores, a plurality of graphics processing units, a plurality of peripheral devices, one or more memory controller circuits, and an interconnect fabric configured to provide communication between the one or more memory controller circuits and the processor cores, graphics processing units, and peripheral devices, wherein the interconnect fabric comprises at least a first network and a second network, the first network including one or more characteristics for reducing latency compared to the second network of at least two networks. For example, one or more characteristics include a route for the first network on the surface of the semiconductor die that is shorter than the route of the second network. In another example, one or more characteristics include wiring in a metal layer having lower latency characteristics than the wiring layer used for the second network. In one embodiment, the second network includes one or more second characteristics for increasing bandwidth compared to the first network. For example, one or more second characteristics may include wider interconnects compared to the second network (e.g., more wires per interconnect than in the first network). One or more second characteristics may include wiring within a metal layer that is denser than the wiring layer used for the first network.

[0119] In one embodiment, a system-on-a-chip (SOC) may include multiple independent networks. The networks may be physically independent (e.g., having dedicated wires and other circuits forming the network) or logically independent (e.g., communications supplied by agents within the SOC may be logically defined to be transmitted on a selected network among the multiple networks and not affected by transmissions on other networks). In some embodiments, network switches may be included to transmit packets on a given network. The network switches may be physically part of the network (e.g., each network may have its own dedicated network switch). In other embodiments, network switches may be shared among physically independent networks, thus ensuring that communications received on one of the networks remain on that network.

[0120] High bandwidth can be achieved through parallel communication over different networks by providing physically and logically independent networks. Furthermore, different traffic can be transmitted over different networks, and therefore a given network can be optimized for a given type of traffic. For example, a processor such as a central processing unit (CPU) in a SOC may be sensitive to memory latency and can cache data that is expected to be coherent between the processor and memory. Therefore, a CPU network can be provided, with the CPU and memory controller in the system acting as agents. The CPU network can be optimized to provide low latency. For example, in one embodiment, there may be virtual channels for low-latency requests and bulk requests. Low-latency requests may be more favorable than bulk requests when transporting them around the fabric and by the memory controller. The CPU network can also support cache coherency using messages and protocols defined to communicate coherently. Another network may be an input / output (I / O) network. This network may be used by various peripheral devices ("peripherals") to communicate with memory. The network can support the bandwidth required by the peripherals and can also support cache coherency. However, I / O traffic can often have significantly higher latency than CPU traffic. By separating I / O traffic from CPU traffic and separating it from memory traffic, the impact of I / O traffic on CPU traffic can be reduced. The CPU may be included as an agent on the I / O network to manage coherence and communicate with peripherals. Yet another network may, in one embodiment, be a relaxed ordering network. Both the CPU and the I / O network can support an ordering model between communications on their networks that provides the ordering expected by the CPU and peripherals.However, the relaxed ordering network may be non-coherent and may not impose many ordering constraints. The relaxed ordering network may be used by a graphics processing unit (GPU) to communicate with a memory controller. Thus, the GPU may have dedicated bandwidth in the network and may not be constrained by the ordering required by the CPU and / or peripherals. Other embodiments may employ any subset and / or any additional networks of the above network as desired.

[0121] A network switch may be a circuit configured to receive communications on a network and forward communications on the network in the direction of their destination. For example, communications supplied by a processor may be transmitted to a memory controller that controls memory mapped to the addresses of the communications. In each network switch, communications can be transmitted forward toward the memory controller. If the communications are reads, the memory controller can return the data to the source, and each network switch can forward the data on the network toward the source. In one embodiment, the network may support multiple virtual channels. The network switch may employ dedicated resources (e.g., buffers) for each virtual channel so that communications on the virtual channels may remain logically independent. The network switch may also employ arbitration circuits to select which communications to forward on the network from among the buffered communications. Virtual channels may be channels that physically share the network but are logically independent on the network (e.g., communications on one virtual channel do not block communications on another virtual channel).

[0122] An agent can generally be any device (e.g., a processor, peripheral, memory controller, etc.) capable of supplying and / or sinking communications on a network. A source agent generates (supplies) communications, and a destination agent receives (sinks) them. A given agent may be a source agent for some communications and a destination agent for others.

[0123] Referring to the drawings, Figure 2 is a general diagram showing physically and logically independent networks. Figures 3-5 are examples of various network topologies. Figure 6 is an example of a SOC with multiple physically and logically independent networks. Figures 7-9 show the various networks in Figure 6 separately for further clarity. Figure 10 is a block diagram of a system including two semiconductor dies, showing the scalability of the network for multiple instances of the SOC. Figures 11 and 12 are exemplary agents shown in more detail. Figure 13 shows various virtual channels and communication types, as well as which virtual channels and communication types apply to which networks in Figure 6. Figure 14 is a flowchart of the method. Further details will be explained below based on the drawings.

[0124] Figure 2 is a block diagram of a system including one embodiment of multiple networks interconnecting agents. In Figure 1, agents A10A, A10B, and A10C are shown, but in various embodiments, any number of agents may be included. Agents A10A to A10B are connected to network A12A, and agents A10A and A10C are connected to network A12B. In various embodiments, any number of networks A12A to A12B may be included. Network A12A includes multiple network switches, including network switches A14A, A14AB, A14AM, and A14AN (collectively network switch A14A). Similarly, network A12B includes multiple network switches, including network switches A14BA, A14BB, A14BM, and A14BN (collectively network switch A14B). Different networks A12A to A12B may include different numbers of network switches A14A, and A12A to A12B may include physically separate connections ("wires", "buses", or "interconnects") as various arrows in Figure 2.

[0125] Each network A12A-A12B has its own physically and logically separate interconnects and network switches, and therefore networks A12A-A12B are physically and logically separate. Communications on network A12A are not affected by communications on network A12B, and vice versa. Even the bandwidth on the interconnects within each network A12A-A12B is separate and independent.

[0126] Optionally, agents A10A to A10C may include or be coupled to network interface circuits (reference numerals A16A to A16C, respectively). Some agents A10A to A10C may include or be coupled to network interfaces A16A to A16C, while other agents A10A to A10C may not include or be coupled to network interfaces A16A to A16C. Network interfaces A16A to A16C may be configured to transmit and receive traffic on network A12A to A12B on behalf of the corresponding agents A10A to A10C. Network interfaces A16A to A16C may be configured to convert or modify communications issued by the corresponding agents A10A to A10C to conform to the protocol / format of network A12A to A12B, remove modifications, or convert received communications to the protocol / format used by agents A10A to A10C. Therefore, network interfaces A16A to A16C may be used for agents A10A to A10C that are not specifically designed to interface directly to networks A12A to A12B. In some cases, agents A10A to A10C may communicate over two or more networks (for example, agent A10A communicates over both networks A12A to A12B in Figure 1). The corresponding network interface A16A may be configured to separate traffic issued by agent A10A to networks A12A to A12B depending on which network A12A to A12B each communication is assigned to, and network interface A16A may be configured to combine traffic received from networks A12A to A12B for the corresponding agent A10A. Any mechanism may be used to determine that networks A12A to A12B carry a given communication (for example, based on the type of communication, the destination agents A10B to A10C for the communication, the address, etc., in various embodiments).

[0127] The network interface circuit is optional and is not often required for agents to directly support networks A12A-A12B; therefore, for simplicity, the network interface circuit is omitted from the rest of the drawings. However, it should be understood that the network interface circuit may be employed by any agent, a subset of agents, or all agents in any of the illustrated embodiments.

[0128] In one embodiment, the system in Figure 2 may be implemented as a System-on-a-Chip (SOC), and the components shown in Figure 2 may be formed on a single semiconductor substrate die. The circuitry included in the SOC may include a plurality of agents A10C and a plurality of network switches A14A to A14B coupled to the plurality of agents A10A to A10C. The plurality of network switches A14A to A14B are interconnected to form a plurality of physically and logically independent networks A12A to A12B.

[0129] Since networks A12A to A12B are physically and logically independent, different networks can have different topologies. For example, a given network may have a ring, mesh, tree, star, full connect set of network switches (e.g., switches directly connected to each other within the network), a shared bus with multiple agents coupled to the bus, or a hybrid of any one or more of these topologies. Each network A12A to A12B may employ a topology that provides, for example, desired bandwidth and latency attributes to the network, or provides any desired attributes to the network. Thus, generally speaking, a SOC may include a first network constructed according to a first topology and a second network constructed according to a second topology different from the first topology.

[0130] Figures 3 to 5 show exemplary topologies. Figure 3 is a block diagram of one embodiment of a network connecting agents A10A to A10C using a ring topology. In the example in Figure 3, the ring is formed from network switches A14AA to A14AH. Agent A10A is connected to network switch A14AA. Agent A10B is connected to network switch A14AB, and agent A10C is connected to network switch A14AE.

[0131] In a ring topology, each network switch A14AA~A14AH may be connected to two other network switches A14AA~A14AH, and the switches form a ring such that any network switch A14AA~A14AH can reach any other network switch in the ring by transmitting communications in the direction of other network switches on the ring. A given communication can reach a target network switch by passing through one or more intermediate network switches in the ring. When a given network switch A14AA~A14AH receives a communication from an adjacent network switch A14AA~A14AH on the ring, the given network switch can inspect the communication and determine that the destination of the communication is agents A10A~A10C to which the given network switch is connected. If so, the given network switch may terminate the communication and forward the communication to the agent. Otherwise, a given network switch can forward communications to the next network switch in the ring (for example, other network switches A14AA~A14AH adjacent to the given network switch but not the adjacent network switch from which the given network switch received communications). Network switches adjacent to a given network switch may be network switches to which the given network switch can directly transmit communications without the communications going through any intermediate network switches.

[0132] Figure 4 is a block diagram of one embodiment of a network that uses a mesh topology to connect agents A10A to A10P. As shown in Figure 4, the network may include network switches A14AA to A14AH. Each network switch A14AA to A14AH is coupled to two or more other network switches. For example, as shown in Figure 4, network switch A14AA is coupled to network switches A14AB and A14AE, network switch A14AB is coupled to network switches A14AA, A14AF, and A14AC, and so on. Thus, different network switches in a mesh network may be coupled to different numbers of other network switches. Furthermore, while the embodiment in Figure 4 has a relatively symmetrical structure, other mesh networks may be asymmetrical, for example, depending on the various traffic patterns that are expected to prevail on the network. Each network switch A14AA~A14AH can use one or more attributes of the received communication to determine which neighboring network switches A14AA~A14AH will transmit the communication to (unless the agents A10A~A10P to which the receiving network switches A14AA~A14AH are connected are the destinations of the communication, in which case the receiving network switches A14AA~A14AH can terminate the communication on the network and provide it to the destination agents A10A~A10P). For example, in one embodiment, network switches A14AA~A14AH may be programmed during system initialization to route communications based on various attributes.

[0133] In one embodiment, communication may be routed based on the destination agent. The routing may be configured to transport communication through the minimum number of network switches ("shortest path") between the source agent and the destination agent that can be supported in the mesh topology. Alternatively, different communications from a given source agent to a given destination agent may take different paths through the mesh. For example, latency-sensitive communications may be transmitted over shorter paths, while less critical communications may take different paths to avoid consuming bandwidth on shorter paths, for example, different paths may not be as heavily loaded during use.

[0134] The mesh in Figure 4 may be an example of a partially connected mesh, where at least some communications may pass through one or more intermediate network switches within the mesh. A fully connected mesh may have connections from each network switch to each other, and therefore any communications may be transmitted without traversing any intermediate network switches. In various embodiments, any level of interconnectivity can be used.

[0135] Figure 5 is a block diagram of one embodiment of a network that uses a tree topology to connect agents A10A to A10E. Network switches A14A to A14AG are interconnected in this example to form a tree. The tree is a form of hierarchical network in which edge network switches (e.g., A14A, A14AB, A14AC, A14AD, and A14AG in Figure 5) connect to agents A10A to A10E and intermediate network switches (e.g., A14AE and A14AF in Figure 5) connect only to other network switches. A tree network can be used, for example, when a particular agent is often the destination of communications issued by other agents, or is often the source agent of communications. Thus, for example, the tree network in Figure 5 could be used for agent A10E, which is the primary source or destination for communications. For example, agent A10E may be a memory controller, which is often the destination of memory transactions.

[0136] Many other possible topologies exist that may be used in other embodiments. For example, a star topology has a source / destination agent at the "center" of the network, and other agents can be connected directly to the central agent or via a set of network switches. A star topology may be used when the central agent is often the source or destination of communications, similar to a tree topology. A shared bus topology may also be used, or a hybrid of any two or more of the topologies may be used.

[0137] Figure 6 is a block diagram of one embodiment of a system-on-a-chip (SOC) A20 having multiple networks for one embodiment. For example, SOC A20 may be an instance of SOC 10 in Figure 1. In the embodiment of Figure 6, SOC A20 includes multiple processor clusters (P clusters) A22A to A22B, multiple input / output (I / O) clusters A24A to A24D, multiple memory controllers A26A to A26D, and multiple graphics processing units (GPUs) A28A to A28D. As implied by the name (SOC), the components shown in Figure 6 (except for memory A30A to A30D in this embodiment) may be integrated on a single semiconductor die or "chip". However, other embodiments may employ two or more dies combined or packaged in any desired manner. In addition, while a specific number of P clusters A22A-A22B, I / O clusters A24-A24D, memory controllers A26A-A26D, and GPUs A28A-A28D are shown in the example in Figure 6, the number and arrangement of any of the above components may be changed, and may be more or less than the number shown in Figure 6. The memory A30A-A30D are coupled to the SOC A20, and more specifically, to the memory controllers A26A-A26D, respectively, as shown in Figure 6.

[0138] In the illustrated embodiment, SOC A20 includes three physically and logically independent networks formed from a plurality of network switches A32, A34, and A36 as shown in Figure 6, and interconnects between them, indicated as arrows between the network switches and other components. Other embodiments may include more or fewer networks. Network switches A32, A34, and A36 may be instances of network switches similar to, for example, network switches A14A to A14B described above with respect to Figures 2 to 5. The plurality of network switches A32, A34, and A36 are coupled to a plurality of P clusters A22A to A22B, a plurality of GPUs A28A to A28D, a plurality of memory controllers A26 to A25B, and a plurality of I / O clusters A24A to A24D, as shown in Figure 6. P clusters A22A-A22B, GPUs A28A-A28B, memory controllers A26A-A26B, and I / O clusters A24A-A24D are all examples of agents communicating across various networks of SOC A20. Other agents may be included as needed.

[0139] In Figure 6, the central processing unit (CPU) network is formed from a first subset of multiple network switches (e.g., network switch A32), and the interconnects between them are shown as short dashed / long dashed lines, such as reference code A38. The CPU network connects the P clusters A22A-A22B and the memory controllers 26A-A26D. The I / O network is formed from a second subset of multiple network switches (e.g., network switch A34) and the interconnects between them, shown as solid lines, such as reference code A40. The I / O network connects the P clusters A22A-A22B, the I / O clusters A24A-A24D, and the memory controllers A26A-A26B. The relaxation sequence network is formed from a third subset of multiple network switches (e.g., network switch A36) and the interconnects between them, shown as short dashed lines, such as reference code A42. The relaxed ordering network connects the GPUs 2A8A-A28D and the memory controllers A26A-A26D. In one embodiment, the relaxed ordering network may also connect a selection of I / O clusters A24A-A24D. As described above, the CPU network, I / O network, and relaxed ordering network are independent of each other (e.g., logically and physically). In one embodiment, the protocols on the CPU network and I / O network support cache coherency (e.g., the network is coherent). The relaxed ordering network does not need to support cache coherency (e.g., the network is non-coherent). The relaxed ordering network also has reduced ordering constraints compared to the CPU network and I / O network. For example, in one embodiment, a set of virtual channels and subchannels within the virtual channels are defined for each network. For the CPU and I / O networks, communication within the same virtual channels and subchannels between the same source and destination agents can be ordered. Relaxed ordering networks allow for the ordering of communications between the same source agent and destination agent.In one embodiment, only communications between the same source agent and destination agent to the same address (at a given granularity, such as a cache block) may be ordered. Because less strict ordering is performed on relaxed ordering networks, transactions may be allowed to complete out of order, for example, if a newer transaction is ready to complete before an older transaction, thus potentially achieving higher bandwidth on average.

[0140] The interconnect between network switches A32, A34, and A36 can have any form and configuration in various embodiments. For example, in one embodiment, the interconnect may be a point-to-point unidirectional link (e.g., a bus or serial link). Packets may be transmitted over the link, and the packet format may include data indicating the virtual channel and subchannel on which the packet is progressing, a memory address, source and destination agent identifiers, data (where appropriate), etc. Multiple packets may form a given transaction. A transaction may be a complete communication between a source agent and a target agent. For example, a read transaction may, depending on the protocol, include a read request packet from the source agent to the target agent, one or more coherence message packets between the caching agent and the target agent and / or the source agent if the transaction is coherent, a data response packet from the target agent to the source agent, and optionally a completion packet from the source agent to the target agent. A write transaction may include a write request packet from the source agent to the target agent, one or more coherence message packets (similar to a read transaction if the transaction is coherent), and possibly a completion packet from the target agent to the source agent. In one embodiment, the write data may be included in the write request packet or transmitted from the source agent to the target agent in a separate write data packet.

[0141] The agent arrangement in Figure 6 can, in one embodiment, represent the physical arrangement of agents on the semiconductor die forming the SOC A20. That is, Figure 6 can be viewed as the surface area of ​​the semiconductor die, and the locations of the various components in Figure 6 can approximate their physical locations by their area. For example, I / O clusters A24A to A24D can be arranged within the semiconductor die region represented by the top of the SOC A20 (as oriented in Figure 6). P clusters A22A to A22B can be arranged in the region represented by the portion of the SOC A20 below and between the arrangement of I / O clusters A24A to A24D, as oriented in Figure 6. GPUs A24A to A28D may be located in the center and extend toward the region represented by the bottom of the SOC A20 as oriented in Figure 6. Memory controllers A26A to A26D can be arranged on the regions represented by the right and left of the SOC A20, as oriented in Figure 6.

[0142] In one embodiment, the SOC A20 can be designed to directly couple to one or more other instances of the SOC A20, logically coupling a given network on each instance into a single network over which an agent on one die can logically communicate with agents on different dies in the same way that an agent on one die communicates with another agent on the same die. Latencies may differ, but communication can be performed in the same way. Thus, as shown in Figure 6, the network extends to the bottom of the SOC A20 as oriented in Figure 6. Interface circuits not shown in Figure 6 (e.g., serializer / deserializer (SERDES) circuits) may be used to communicate with another die across die boundaries. Thus, the network can be scalable to two or more semiconductor dies. For example, two or more semiconductor dies may be configured as a single system where the presence of multiple semiconductor dies is transparent to the software running on a single system. In one embodiment, the delay in die-to-die communication can be minimized, as a form of software transparency for multi-die systems, so that inter-die communication does not typically result in significantly additional latency compared to intra-die communication. In another embodiment, the network may be a closed network that communicates only within dies.

[0143] As described above, different networks may have different topologies. In the embodiment of Figure 6, for example, the CPU and I / O networks may implement a ring topology, and the relaxed sequence may implement a mesh topology. However, other topologies may be used in other embodiments. Figures 7, 8, and 9 show portions of the SOC A30 that include different networks, namely the CPU (Figure 7), I / O (Figure 8), and relaxed sequence (Figure 9). As seen in Figures 7 and 8, network switches A32 and A34 form a ring when coupled to corresponding switches on other dies, respectively. If only a single die is used, a connection may be made between the two network switches A32 or A34 at the bottom of the SOC A20, as oriented in Figures 7 and 8 (e.g., via an external connection on the pins of the SOC A20). Alternatively, the two lower network switches A32 or A34 may have a link between them that can be used in a single-die configuration, or the network may operate in a daisy-chain topology.

[0144] Similarly, Figure 9 shows the connection of network switch A36 in a mesh topology between GPUs A28A-A28D and memory controllers A26A-A26D. As mentioned above, in one embodiment, one or more of the I / O clusters A24A-A24D may be coupled to a relaxed sequence where the network is good. For example, I / O clusters A24A-A24D including video peripherals (e.g., display controllers, memory scalers / rotators, video encoders / decoders, etc.) may have access to the relaxed sequence network for video data.

[0145] Network switch A36 near the bottom of SOC A30, as oriented in Figure 9, may include connections that can be routed to another instance of SOC A30, allowing the CPU to extend across multiple dies, as described above with respect to mesh and I / O networks. In a single-die configuration, paths extending outside the chip may not be used. Figure 10 is a block diagram of a two-die system in which each network extends across two SOC dies A20A-A20B, forming a network that is logically the same even when they extend across two dies. Network switches A32, A34, and A36 have been removed in Figure 10 for simplification, and the relaxation sequence network is simplified to lines, but in one embodiment it may be a mesh. The I / O network A44 is shown as a solid line, the CPU network A46 as a dashed line, and the relaxation sequence network A48 as a dashed line. The ring structure of networks A44 and A46 is also evident in Figure 10. Although two dies are shown in Figure 10, other embodiments may employ three or more dies. In various embodiments, the network may be daisy-chained together, fully connected with point-to-point links between teach die pairs, or may be any other connection structure.

[0146] In one embodiment, physically separating the I / O network from the CPU network may help the system provide low-latency memory access by processor clusters A22A-A22B, as I / O traffic can be transferred to the I / O network. The network uses the same memory controller to access memory, and therefore the memory controller may be designed to prioritize memory traffic from the CPU network to some extent over memory traffic from the I / O network. Processor clusters A22-A22B may also be part of the I / O network to access device space within I / O clusters A24A-A24D (e.g., in programmed input / output (PIO) transactions). However, memory transactions initiated by processor clusters A22A-A22B may be transmitted over the CPU network. Thus, CPU clusters A22A-A22B may be an example of an agent coupled to at least two of several physically and logically independent networks. The agent may be configured to generate the transaction to be transmitted and, based on the type of transaction (e.g., memory or PIO), select at least two of a plurality of physically and logically independent networks to transmit the transaction.

[0147] Various networks may include different numbers of physical and / or virtual channels. For example, an I / O network may have multiple request and completion channels, while a CPU network may have one request channel and one completion channel (or vice versa). When there are two or more requests transmitted on a given request channel, they may be determined in any desired manner (e.g., by the type of request, by the priority of the requests, to balance bandwidth across physical channels, etc.). Similarly, I / O networks and CPU networks may include snoop virtual channels for carrying snoop requests, but a relaxed sequence network, being non-coherent in this embodiment, may not include snoop virtual channels.

[0148] Figure 11 is a block diagram of one embodiment of an input / output (I / O) cluster A24A, shown in more detail. Other I / O clusters A24B to A24D may be similar. In the embodiment of Figure 11, I / O cluster A24A includes peripherals A50 and A52, a peripheral interface controller A54, a local interconnect A56, and a bridge A58. Peripheral device A52 may be coupled to an external component A60. Peripheral interface controller A54 may be coupled to a peripheral interface A62. Bridge A58 may be coupled to a network switch A34 (or a network interface coupled to network switch A34).

[0149] Peripherals A50 and A52 may include any set of additional hardware functions (other than, for example, the CPU, GPU, and memory controller) included in the SOC A20. For example, peripherals A50 and A52 may include video peripherals such as image signal processors configured to process image acquisition data from a camera or other image sensor, video encoders / decoders, scalers, rotators, blenders, and display controllers. Peripherals may include audio peripherals such as microphones, speakers, interfaces to microphones and speakers, audio processors, digital signal processors, and mixers. Peripherals may include networking peripherals such as media access controllers (MACs). Peripherals may include other types of memory controllers such as non-volatile memory controllers. Some peripherals A52 may include on-chip and off-chip components A60. The peripheral interface controller A54 may include interface controllers for various external interfaces A62 of the SOC A20, including interfaces such as Universal Serial Bus (USB), PCI Express (PCIe), Peripheral Components Interconnect (PCI), serial and parallel ports, etc.

[0150] The local interconnect A56 may be an interconnect through which various peripherals A50, A52, and A54 communicate. The local interconnect A56 may differ from the system-wide interconnect (e.g., CPU, I / O, and relaxed network) shown in Figure 6. The bridge A58 may be configured to translate communication over the local interconnect to communication over the system-wide interconnect and vice versa. In one embodiment, the bridge A58 may be coupled to one of the network switches A34. The bridge A58 can also manage the ordering of transactions issued by peripherals A50, A52, and A54. For example, the bridge A58 may use a network-supported cache coherency protocol to ensure transaction ordering on behalf of peripherals A50, A52, and A54, etc. Different peripherals A50, A52, and A54 may have different ordering requirements, and the bridge A58 may be configured to adapt to these different requirements. In some embodiments, bridge A58 can also implement various performance enhancement features. For example, bridge A58 may prefetch data for a given request. Bridge A58 can capture coherent copies of cache blocks (e.g., in an exclusive state) that are targeted by one or more transactions from peripherals A50, A52, and A54, enabling the transaction to complete locally and perform ordering. Bridge A58 can speculatively capture exclusive copies of one or more cache blocks targeted by a subsequent transaction, and if the exclusive state is successfully maintained until the subsequent transaction can be completed (e.g., after satisfying any ordering constraints with the previous transaction), the subsequent transaction can use the cache blocks to complete the transaction. Thus, in one embodiment, multiple requests within a cache block can be served from the cached copy.Various details can be found in U.S. Provisional Patent Application No. 63 / 170,868, filed on April 5, 2021; No. 63 / 175,868, filed on April 16, 2021; and No. 63 / 175,877, filed on April 16, 2021. These patent applications are incorporated herein by reference in their entirety. In the event of any incorporation of material conflict with material expressly described herein, the material expressly described herein shall prevail.

[0151] Figure 12 is a block diagram of one embodiment of processor cluster A22A. Other embodiments may be similar. In the embodiment of Figure 12, processor cluster A22A includes one or more processors A70 coupled to a last-level cache (LLC) A72. The LLC A72 may optionally include interface circuits for interface connection to network switches A32 and A34 for transmitting transactions over the CPU network and I / O network.

[0152] Processor A70 may include any circuitry and / or microcode configured to execute instructions defined in the instruction set architecture implemented by Processor A70. Processor A70 may have any microarchitecture implementation, performance and power characteristics, etc. For example, the processor may be in-order execution, out-of-order execution, superscalar, superpipelined, etc.

[0153] Any caches within LLC A72 and processor A70 can have any capacity and configuration, such as set-associative, direct-mapped, or fully associative. The cache block size may be any desired size (e.g., 32 bytes, 64 bytes, 128 bytes, etc.). A cache block may be a unit of allocation and deallocation in LLC A70. Furthermore, a cache block may be a unit on which coherency is maintained in this embodiment. A cache block may also be referred to as a cache line. In one embodiment, a distributed directory-based coherency scheme may be implemented using coherency points in each memory controller A26 in the system, where the coherency points are applied to memory addresses mapped to the memory controllers. The directory can track the state of cached cache blocks in any coherent agent. The coherency scheme may be scalable to many memory controllers across multiple semiconductor dies.For example, the coherency scheme has the following features: precise directories for snoop filtering and conflict resolution in coherent agents and memory agents, ordering points (access order) determined by the memory agent, serialization points moving between coherent agents and memory agents, secondary completion (invalidation acknowledgment) collection in the requesting coherent agent, tracked by completion counts provided by the memory agent, fill / snoop and snoop / victim-ack conflict resolution processed in the coherent agent via directory states provided by the memory agent, separate primary / secondary shared states to restrict flight snoops to the same address / target to assist in conflict resolution, absorption of conflict snoops in the coherent agent to avoid deadlocks without additional nack / conflict / retry messages or actions, and minimization of serialization (one additional nack per access mechanism for transferring ownership through the conflict chain). It may employ one or more of the following: additional message latency, message minimization (direct messages between relevant agents with no additional messages to handle competition / contest (e.g., no messages returning to memory agents)), store conditions without excessive invalidation on failure due to competition, minimal data transfer (only in dirty cases) and exclusive ownership claims intended to correct the entire cache line with the relevant cache / directory state, separate snoopback and snoopforward message types to handle both cacheable and non-cacheable flows (e.g., 3-hop and 4-hop protocols). Further details can be found in U.S. Provisional Patent Application No. 63 / 077,371, filed September 11, 2020, which is incorporated herein by reference in its entirety. In the event that any incorporated material conflicts with any material expressly described herein, the material expressly described herein shall prevail.

[0154] Figure 13 shows a pair of tables A80 and A82 illustrating the virtual channels and traffic types and networks shown in Figures 6 to 9, used for one embodiment. As shown in Table A80, virtual channels may include bulk virtual channels, low-latency (LLT) virtual channels, real-time (RT) virtual channels, and virtual channels for non-DRAM messages (VCP). A bulk virtual channel may be the default virtual channel for memory access. A bulk virtual channel may receive a lower quality of service than, for example, LLT and RT virtual channels. An LLT virtual channel may be used for memory transactions where low latency is required for high-performance operation. An RT virtual channel may be used for memory transactions (e.g., video streams) that have latency and / or bandwidth requirements for proper operation. A VCP channel may be used to isolate traffic not directed to memory in order to prevent interference with memory transactions.

[0155] In one embodiment, bulk virtual channels and LLT virtual channels may be supported on all three networks (CPU, I / O, and relaxed order). RT virtual channels may be supported on the I / O network but not on the CPU or relaxed order network. Similarly, VCP virtual channels may be supported on the I / O network but not on the CPU or relaxed order network. In one embodiment, VCP virtual channels may be supported on the CPU and relaxed order networks only for transactions targeting network switches on that network (e.g., for configuration) and therefore may not be used during normal operation. Thus, as shown in Table A80, different networks can support different numbers of virtual channels.

[0156] Table A82 shows various traffic types and which network carries that traffic type. Traffic types may include coherent memory traffic, non-coherent memory traffic, real-time (RT) memory traffic, and VCP (non-memory) traffic. Both CPU and I / O networks can carry coherent traffic. In one embodiment, coherent memory traffic supplied by processor clusters A22A-A22B may be carried over the CPU network, and the I / O network may carry coherent memory traffic supplied by I / O clusters A24A-A24D. Non-coherent memory traffic may be carried over the relaxed sequence network, and RT and VCP traffic may be carried over the I / O network.

[0157] Figure 14 is a flowchart illustrating one embodiment of how to initiate a transaction on a network. In one embodiment, an agent can generate a transaction to be transmitted (block A90). The transaction is transmitted over one of several physically and logically independent networks. A first network among the several physically and logically independent networks is constructed according to a first topology, and a second network among the several physically and logically independent networks is constructed according to a second topology different from the first topology. One of the several physically and logically independent networks is selected to transmit the transaction based on the type of transaction (block A92). For example, processor clusters A22A-A22B can transmit coherent memory traffic over the CPU network and PIO traffic over the I / O network. In one embodiment, an agent can select a virtual channel from several virtual channels supported on the selected network among several physically and logically independent networks based on one or more attributes of the transaction other than its type (block A94). For example, the CPU may select an LLT virtual channel for a subset of memory transactions (e.g., the oldest memory transaction that is a cache miss, or the number of cache misses up to a threshold number after a bulk channel has been selected). The GPU may select between an LLT virtual channel and a bulk virtual channel based on the urgency of the data required. Video devices may use an RT virtual channel as needed (e.g., a display controller may issue frame data readouts over an RT virtual channel). A VCP virtual channel may be selected for transactions that are not memory transactions. An agent can transmit transaction packets over the selected network and virtual channel. In one embodiment, transaction packets in different virtual channels may take different paths through the network.In one embodiment, transaction packets may take different routes based on the type of transaction packet (e.g., request versus response). In one embodiment, different routes may be supported for both different virtual channels and different types of transactions. In other embodiments, one or more additional attributes of transaction packets may be adopted to determine the route those packets take through the network. Alternatively, network switches forming the network may route different packets based on virtual channel, type, or any other attribute. Different routes may refer to traversing at least one segment between network switches that are not traversed on other routes, even if transaction packets using different routes are traveling from the same source to the same destination. Using different routes can provide load balancing and / or reduced latency for transactions in the network.

[0158] In one embodiment, the system comprises a plurality of processor clusters, a plurality of memory controllers, a plurality of graphics processing units, a plurality of agents, and a plurality of network switches coupled to the plurality of processor clusters, a plurality of graphics processing units, a plurality of memory controllers, and a plurality of agents. A given processor cluster includes one or more processors. The memory controllers are configured to control access to memory devices. A first subset of the plurality of network switches is interconnected to form a central processing unit (CPU) network between the plurality of processor clusters and a plurality of memory controllers. A second subset of the plurality of network switches is interconnected to form an input / output (I / O) network between the plurality of processor clusters, a plurality of agents, and a plurality of memory controllers. A third subset of the plurality of network switches is interconnected to form a relaxed ordering network between the plurality of graphics processing units, a selected agent from among the plurality of agents, and a plurality of memory controllers. The CPU network, the I / O network, and the relaxed ordering network are independent of each other. The CPU network and the I / O network are coherent. The relaxed ordering network is non-coherent and has reduced ordering constraints compared to the CPU network and the I / O network. In one embodiment, at least one of the CPU network, I / O network, and relaxation sequence network has some physical channels that are different from some physical channels on another one of the CPU network, I / O network, and relaxation sequence network. In one embodiment, the CPU network is a ring network. In one embodiment, the I / O network is a ring network. In one embodiment, the relaxation sequence network is a mesh network. In one embodiment, a first agent of a plurality of agents includes an I / O cluster that includes a plurality of peripheral devices. In one embodiment, the I / O cluster further includes a bridge that is coupled to the plurality of peripheral devices and further coupled to a first network switch in a second subset.In one embodiment, the system further comprises a CPU circuit configured to translate communications from a given agent into communications for a given network among a network interface network, an I / O network, and a relaxed sequence network, wherein the network interface circuit is coupled to one of a plurality of network switches in a given network.

[0159] In one embodiment, a system-on-a-chip (SOC) includes a semiconductor die on which a circuit is formed. The circuit comprises a plurality of agents and a plurality of network switches coupled to the plurality of agents. The plurality of network switches are interconnected to form a plurality of physically and logically independent networks. A first network of the plurality of physically and logically independent networks is constructed according to a first topology, and a second network of the plurality of physically and logically independent networks is constructed according to a second topology different from the first topology. In one embodiment, the first topology is a ring topology. In one embodiment, the second topology is a mesh topology. In one embodiment, coherence is performed on the first network. In one embodiment, the second network is a relaxed sequence network. In one embodiment, at least one of the plurality of physically and logically independent networks implements a first number of physical channels, and at least one other of the plurality of physically and logically independent networks implements a second number of physical channels, where the first number is different from the second number. In one embodiment, the first network includes one or more first virtual channels, and the second network includes one or more second virtual channels. At least one of the one or more first virtual channels is different from one or more second virtual channels. In one embodiment, the SOC further includes a network interface circuit configured to translate communications from a given agent among a plurality of agents to communications for a given network among a plurality of physically and logically independent networks. The network interface circuit is coupled to one of a plurality of network switches in the given network. In one embodiment, the first agent among the plurality of agents is coupled to at least two of a plurality of physically and logically independent networks. The first agent is configured to generate the transactions to be transmitted.The first agent is configured to select, based on the type of transaction, one of at least two of a plurality of physically and logically independent networks to which the transaction should be transmitted. In one embodiment, one of the at least two networks is an I / O network through which the I / O transaction is transmitted.

[0160] In one embodiment, the method involves generating a transaction in an agent connected to a plurality of physically and logically independent networks, wherein a first network among the plurality of physically and logically independent networks is constructed according to a first topology, and a second network among the plurality of physically and logically independent networks is constructed according to a second topology different from the first topology, and selecting one of the plurality of physically and logically independent networks to which the transaction should be transmitted based on the type of the transaction. In one embodiment, the method further includes selecting a virtual channel from a plurality of virtual channels supported by one of the plurality of physically and logically independent networks based on one or more attributes of the transaction other than the type. interrupt

[0161] Figures 15 to 26 illustrate various embodiments of a scalable interrupt structure. For example, in a system including two or more integrated circuit dies, a given integrated circuit die may include a local interrupt distribution circuit for distributing interrupts among processor cores in a given integrated circuit die. At least one of the two or more integrated circuit dies may include a global interrupt distribution circuit, and the local and global interrupt distribution circuits implement a multilevel interrupt distribution scheme. In one embodiment, the global interrupt distribution circuit is configured to sequentially transmit interrupt requests to local interrupt distribution circuits, and the local interrupt distribution circuits are configured to sequentially transmit interrupt requests to local interrupt destinations before responding to interrupt requests from the global interrupt distribution circuit.

[0162] A computing system generally includes one or more processors that function as a central processing unit (CPU), along with one or more peripheral devices that implement various hardware functions. The CPU runs control software (e.g., an operating system) that controls the operation of the various peripheral devices. The CPU can also run applications that provide user functions within the system. Furthermore, the CPU can run software that interacts with the peripheral devices and performs various services on their behalf. Other processors not used as CPUs within the system (e.g., processors integrated into some peripheral devices) can also run such software for the peripheral devices.

[0163] Peripheral devices can use interrupts to have the processor execute software on their behalf. Generally, peripheral devices issue interrupts by asserting an interrupt signal to an interrupt controller, which typically controls the interrupts that go to the processor. The interrupt causes the processor to stop executing its current software task and save the task's state so that it can be resumed later. The processor can load the state associated with the interrupt and begin executing an interrupt service routine. The interrupt service routine may be driver code for the peripheral device, or execution may be transferred to driver code as needed. Generally, driver code is code provided for a peripheral device that is executed by the processor to control and / or configure the peripheral device.

[0164] The latency from interrupt assertion to interrupt service can be critical to system performance and even functionality. Furthermore, efficiently determining which CPU should service an interrupt and distributing it with minimal perturbation to the rest of the system can be crucial for maintaining both system performance and low power consumption. As the number of processors in a system increases, efficiently and effectively scaling interrupt distribution becomes even more critical.

[0165] Next, referring to Figure 15, a block diagram of one embodiment of a part of system B10 is shown, which includes an interrupt controller B20 coupled to multiple cluster interrupt controllers B24A to B24n. Each of the multiple cluster interrupt controllers B24A to B24n is coupled to each of the multiple processors B30 (e.g., a processor cluster). Interrupt controller B20 is coupled to multiple interrupt sources B32.

[0166] When at least one interrupt is received by interrupt controller B20, interrupt controller B20 may be configured to attempt to deliver the interrupt (for example, to processor B30 to service the interrupt by executing software to log the interrupt for further service by an interrupt service routine, and / or to provide the processing requested by the interrupt via the interrupt service routine). In system B10, interrupt controller B20 may attempt to deliver the interrupt via cluster interrupt controllers B24A-B24n. Each cluster controller B24A-B24n is associated with a processor cluster and may attempt to deliver the interrupt to processor B30 in each of the multiple processors that form the cluster.

[0167] More specifically, the interrupt controller B20 can be configured to attempt to distribute interrupts in multiple iterations via cluster interrupt controllers B24A-B24n. The interface between the interrupt controller B20 and each of the interrupt controllers B24A-B24n may include a request / acknowledgment (Ack) / non-acknowledgment (Nack) structure. For example, requests may be identified by soft, hard, and enforce iterations in the illustrated embodiment. The first iteration ("soft" iteration) may be signaled by asserting a soft request. The next iteration ("hard" iteration) may be signaled by asserting a hard request. The last iteration ("enforce" iteration) may be signaled by asserting a enforce request. A given cluster interrupt controller B24A~B24n can respond to soft and hard iterations with an Ack response (indicating that processor B30 in the processor cluster associated with the given cluster interrupt controller B24A~B24n has accepted the interrupt and will process at least one interrupt) or a Nack response (indicating that processor B30 in the processor cluster has rejected the interrupt). A forced iteration does not necessarily have to use an Ack / Nack response; rather, it may continue to request interrupts until an interrupt is served, as will be explained in more detail below.

[0168] Cluster interrupt controllers B24A to B24n can attempt to deliver interrupts to a given processor B30, also using the request / Ack / Nack structure with the processor B30. Based on requests from cluster interrupt controllers B24A to B24n, a given processor B30 may be configured to determine whether it can interrupt the current instruction execution within a given period. If the given processor B30 can commit the interrupt within that period, it may be configured to assert an Ack response. If the given processor B30 cannot commit the interrupt, it may be configured to assert a Nack response. The cluster interrupt controllers B24A to B24n can be configured to assert an Ack response to interrupt controller B20 if at least one processor asserts an Ack response to cluster interrupt controllers B24A to B24n, and can be configured to assert a Nack response if processor B30 asserts a Nack response in a given iteration.

[0169] In one embodiment, using a request / ack / nack structure can provide rapid indication of whether an interrupt has been accepted by the request receiver (e.g., depending on the interface, cluster interrupt controllers B24A-B24n or processor B30). This indication may be faster than a timeout, for example, in one embodiment. In addition, in one embodiment, the hierarchical structure of cluster interrupt controllers B24A-B24n and interrupt controller B20 may be more scalable for a larger number of processors in system B10 (e.g., multiple processor clusters).

[0170] An iteration across cluster interrupt controllers B24A to B24n may involve an attempt to deliver interrupts through at least a subset of cluster interrupt controllers B24A to B24n, up to all of them. The iteration can proceed in any desired manner. For example, in one embodiment, interrupt controller B20 may be configured to serially assert interrupt requests to each cluster interrupt controller B24A to B24n, which are terminated by an Ack response from one of the cluster interrupt controllers B24A to B24n (and, in one embodiment, the absence of additional pending interrupts), or by a Nack response from all of the cluster interrupt controllers B24A to B24n. That is, the interrupt controller can select one of the cluster interrupt controllers B24A to B24n and assert an interrupt request to the selected cluster interrupt controller B24A to B24n (for example, by asserting a soft request or a hard request depending on which iteration is being performed). The selected cluster interrupt controllers B24A-B24n may respond with an Ack response, thereby terminating the iteration. Alternatively, if the selected cluster interrupt controllers B24A-B24n assert a Nack response, the interrupt controller may be configured to select another cluster interrupt controller B24A-B24n, and may assert a soft or hard request to the selected cluster interrupt controllers B24A-B24n. The selection and asserting may continue until an Ack response is received or each of the cluster interrupt controllers B24A-B24n is selected and asserts a Nack response. Other embodiments may perform the iteration across the cluster interrupt controllers B24A-B24n in other ways.For example, interrupt controller B20 may be configured to simultaneously assert interrupt requests to a subset of two or more cluster interrupt controllers B24A-B24n, and to continue with the other subsets if each cluster interrupt controller B24A-B24n in the subset provides an Ack response to the interrupt request. Such an implementation may cause a spurious interrupt if two or more cluster interrupt controllers B24A-B24n in the subset provide an Ack response, and therefore the code executed in response to the interrupt may be designed to handle the occurrence of the spurious interrupt.

[0171] The initial iteration may be a soft iteration, as described above. In a soft iteration, a given cluster interrupt controller B24A~B24n can attempt to deliver interrupts to a subset of processors B30 associated with the given cluster interrupt controllers B24A~B24n. This subset may consist of processors B30 that are powered on, and a given cluster interrupt controller B24A~B24n cannot attempt to deliver interrupts to processors B30 that are powered off (or sleeping). In other words, powered-off processors are not included in the subset to which cluster interrupt controllers B24A~B24n attempt to deliver interrupts. Therefore, powered-off processors B30 may remain powered off in a soft iteration.

[0172] Based on the Nack responses from each cluster interrupt controller B24A-B24n during a soft iteration, the interrupt controller B20 can perform a hard iteration. In a hard iteration, a powered-off processor B30 in a given processor cluster may be powered on by individual cluster interrupt controllers B24A-B24n, and individual interrupt controllers B24A-B24n may attempt to deliver interrupts to each processor B30 in the processor cluster. More specifically, in one embodiment, if a processor B30 is powered on to perform a hard iteration, that processor B30 may be readily available for interrupts and may frequently produce Ack responses.

[0173] If a hard iteration terminates while one or more interrupts are still pending, or if a timeout occurs before the soft and hard iterations are completed, the interrupt controller can initiate a forced iteration by asserting a force signal. In one embodiment, the forced iteration may be performed in parallel with the cluster interrupt controllers B24A-B24n, and a Nack response may not be permitted. In one embodiment, the forced iteration may remain in progress until there are no more pending interrupts.

[0174] A given cluster interrupt controller B24A to B24n can attempt to deliver interrupts in any desired manner. For example, a given cluster interrupt controller B24A to B24n can serially assert an interrupt request to each processor B30 in the processor cluster, which is terminated by an Ack response from one of the processors B30, or by a Nack response from each of the processors B30 to which the given cluster interrupt controller B24A to B24n attempts to deliver the interrupt. That is, a given cluster interrupt controller B24A to B4n can select one of the processors B30 and assert an interrupt request to the selected processor B30 (for example, by asserting the request to the selected processor B30). The selected processor B30 can respond with an Ack response, thereby terminating the attempt. On the other hand, if the selected processor B30 asserts a Nack response, the given cluster interrupt controllers B24A to B24n can be configured to select another processor B30 and assert an interrupt request to the selected processor B30. The selection and assertion can continue until an Ack response is received or each of the processors B30 is selected and asserts a Nack response (except for processors that are powered off in a soft iteration). In other embodiments, interrupt requests can be asserted to multiple processors B30 simultaneously or in parallel to processor B30, which may result in spurious interrupts as described above. The given cluster interrupt controllers B24A to B24n can respond to interrupt controller B20 with an Ack response based on having received an Ack response from one of the processors B30, or can respond to interrupt controller B20 with a Nack response if each of the processors B30 responds with a Nack response.

[0175] In one embodiment, the order in which interrupt controller B20 asserts interrupt requests to cluster interrupt controllers B24A~B24n may be programmable. More specifically, in one embodiment, the order may vary based on the interrupt source (for example, interrupts from one interrupt source B32 may result in one order, and interrupts from another interrupt source B32 may result in a different order). For example, in one embodiment, multiple processors B30 in one cluster may be different from multiple processors B30 in another cluster. One processor cluster may have processors optimized for performance but which may be more power-efficient, while another processor cluster may have processors optimized for power efficiency. Interrupts from sources requiring relatively little processing may favor clusters with power-efficient processors, while interrupts from sources requiring considerable processing may favor clusters with higher-performance processors.

[0176] The interrupt source B32 can be any hardware circuit configured to assert an interrupt to cause the processor B30 to execute an interrupt service routine. For example, in one embodiment, various peripheral components (peripherals) may be interrupt sources. Examples of various peripherals are described below with respect to Figure 16. The interrupt is asynchronous with the code being executed by the processor B30 when the processor B30 receives the interrupt. Generally, the processor B30 may be configured to handle interrupts by pausing the execution of the current code, saving the processor context to allow execution to resume after servicing the interrupt, and branching to a predetermined address to begin execution of the interrupt code. The code at the predetermined address can read the state from the interrupt controller to determine which interrupt source B32 asserted the interrupt and the corresponding interrupt service routine to be executed based on the interrupt. The code can queue the interrupt service routine for execution (which may be scheduled by the operating system) and provide the data expected by the interrupt service routine. The code can then return execution to the previously executed code (for example, by reloading the processor context and resuming execution with the instruction that was stopped).

[0177] An interrupt can be transmitted from the interrupt source B32 to the interrupt controller B20 in any desired manner. For example, a dedicated interrupt wire may be provided between the interrupt source and the interrupt controller B20. A given interrupt source B32 can transmit an interrupt to the interrupt controller B20 by asserting a signal on its dedicated wire. Alternatively, a message signal interrupt can be used, in which the message is transmitted via an interconnect used for other communications within the system B10. The message may be, for example, in the form of a write to a specified address. The write data may be a message that identifies the interrupt. A combination of dedicated wires from some interrupt sources B32 and message signal interrupts from other interrupt sources B32 can be used.

[0178] The interrupt controller B20 can receive interrupts and record them as pending interrupts within the interrupt controller B20. Interrupts from various interrupt sources B32 can be prioritized by the interrupt controller B20 according to various programmable priorities configured by the operating system or other control codes.

[0179] Referring now to Figure 16, a block diagram of one embodiment of system B10 implemented as a system-on-a-chip (SOC) B10 coupled to memory B12 is shown. In one embodiment, SOC B10 may be an instance of SOC 10 shown in Figure 1. As implied by the name, the components of SOC B10 may be integrated on a single semiconductor substrate as an integrated circuit “chip”. In some embodiments, the components may be implemented on two or more separate chips in the system. However, for the purposes of this specification, SOC B10 is used as an example. In the illustrated embodiment, the components of SOC B10 include a plurality of processor clusters B14A to B14n, an interrupt controller B20, one or more peripheral components B18 (more simply “peripherals”), a memory controller B22, and a communication fabric B27. Components B14A to B14n, B18, B20, and B22 can all be coupled to the communication fabric B27. The memory controller B22 may be coupled to memory B12 during use. In some embodiments, there may be two or more memory controllers coupled to corresponding memories. The memory address space can be mapped across the memory controllers in any desired manner. In the illustrated embodiments, processor clusters B14A to B14n may include each of several processors (P) B30 and each of the cluster interrupt controllers (IC) B24A to B24n, as shown in Figure 16. The processors B30 can form the central processing unit (CPU(s)) of the SOC B10. In one embodiment, one or more processor clusters B14A to B14n may not be used as CPUs.

[0180] Peripheral device B18 may, in one embodiment, include a peripheral device that is an example of an interrupt source BB32. Thus, one or more peripheral devices B18 may have a dedicated wire to the interrupt controller B20 to transmit interrupts to the interrupt controller B20. Other peripheral devices B18 may use message signal interrupts transmitted via the communication fabric B27. In some embodiments, one or more off-SOC devices (not shown in Figure 16) may also be interrupt sources. The dotted line from the interrupt controller B20 to the off-chip indicates the possibility of an off-SOC interrupt source.

[0181] The hard / soft / forced Ack / Nack interface between cluster ICs B24A-B24n shown in Figure 15 is shown in Figure 16 via an arrow between cluster ICs B24A-B24n and interrupt controller B20. Similarly, the Req Ack / Nack interface between processor B30 and cluster ICs B24A-B24n in Figure 1 is shown by an arrow between cluster ICs B24A-B24n and processor B30 in each of the clusters B14A-B14n.

[0182] As described above, processor clusters B14A to B14n may include one or more processors B30 that can function as the CPU of SOC B10. The system's CPU includes one or more processors that run the system's primary control software, such as the operating system. Generally, the software run by the CPU during use can control other components of the system to achieve desired functions of the system. The processor may also run other software, such as application programs. Application programs may provide user functions and may rely on the operating system for low-level device control, scheduling, memory management, etc. Therefore, the processor may also be referred to as an application processor.

[0183] Generally, a processor may include any circuitry and / or microcode configured to execute instructions defined within the instruction set architecture implemented by the processor. A processor may encompass a processor core implemented as a system-on-a-chip (SOC B10) or other level of integration on an integrated circuit along with other components. A processor may further encompass separate microprocessors, processor cores and / or microprocessors integrated within a multi-chip module implementation, processors implemented as multiple integrated circuits, and so on.

[0184] The memory controller B22 may generally include circuitry for receiving memory operations from other components of the SOC B10 and accessing memory B12 to complete memory operations. The memory controller B22 may be configured to access any type of memory B12. For example, memory B12 may be double data rate (DRAM) such as static random access memory (SRAM), dynamic RAM (DDR, DDR2, DDR3, DDR4, etc.), and synchronous DRAM (SDRAM), including DRAM. Low-power / mobile versions of DDR DRAM (e.g., LPDDR, mDDR, etc.) may be supported. The memory controller B22 may include queues for memory operations for ordering (and possibly reordering) operations and presenting operations to memory B12. The memory controller B22 may further include data buffers for storing write data waiting to be written to memory and read data waiting for the return of memory operations to the source. In some embodiments, the memory controller B22 may include a memory cache for storing recently accessed memory data. In SOC implementations, for example, a memory cache can reduce power consumption in the SOC by avoiding re-access of data from memory B12 when it is expected to be accessed again soon. In some cases, the memory cache may also be referred to as a system cache, unlike private caches that serve only specific components such as L2 caches or processor caches. Furthermore, in some embodiments, the system cache does not need to be located within the memory controller B22.

[0185] Peripheral device B18 may be any set of additional hardware functions included in SOC B10. For example, peripheral device B18 may include video peripherals such as an image signal processor, GPU, video encoder / decoder, scaler, rotator, blender, and display controller configured to process image acquisition data from a camera or other image sensor. Peripheral devices may include audio peripherals such as a microphone, speaker, interface to the microphone and speaker, audio processor, digital signal processor, and mixer. Peripheral devices may include interface controllers for various interfaces outside SOC B10, including interfaces such as Universal Serial Bus (USB), Peripheral Components Interconnect (PCI) including PCI Express (PCIe), and serial and parallel ports. Interconnects to external devices are indicated by dashed arrows in Figure 15 that extend outside SOC B10. Peripheral devices may also include network peripherals such as a Media Access Controller (MAC). Any set of hardware may be included.

[0186] Communication Fabric B27 may be any communication interconnect and protocol for communication between components of SOC B10. Communication Fabric B27 may be bus-based, including tiered buses with shared bus configurations, crossbar configurations, and bridges. Communication Fabric B27 may also be packet-based and may be tiered with bridges, crossbars, point-to-point, or other interconnects.

[0187] It should be noted that the number of components in SOC B10 (and the number of sub-components for those shown in Figure 16, such as processor B30 in each processor cluster B14A-B14n) may vary from embodiment to embodiment. In addition, the number of processor B30 in one processor cluster B14A-B14n may differ from the number of processor B30 in another processor cluster B14A-B14n. Each component / sub-component may be more or less than the number shown in Figure 16.

[0188] Figure 17 is a block diagram showing one embodiment of an interrupt controller that may be implemented by interrupt controller B20 in one embodiment. In the illustrated embodiment, the states include idle state B40, soft state BB42, hard state B44, forced state B46, and standby drain state B48.

[0189] In idle state B40, there may be no pending interrupts. Generally, the state machine can return to idle state B40 from any of the other states, as shown in Figure 17, whenever there are no pending interrupts. When at least one interrupt is received, the interrupt controller B20 can transition to soft state B42. The interrupt controller B20 can also initialize a timeout counter to begin counting a timeout interval that can transition the state machine to forced state B46. The timeout counter may be initialized to zero, incremented, and compared to a timeout value to detect a timeout. Alternatively, the timeout counter may be initialized to a timeout value and decremented until it reaches zero. Incrementing / decrementing may occur per clock cycle of the clock for the interrupt controller B20, or according to another clock (e.g., a fixed-frequency clock from a piezoelectric oscillator).

[0190] In soft state B42, interrupt controller B20 may be configured to initiate a soft iteration attempting to deliver interrupts. If one of the cluster interrupt controllers B24A to B24n transmits an Ack response during the soft iteration and there is at least one pending interrupt, interrupt controller B20 may transition to standby drain state B48. A given processor can provide standby drain state B48 because, although it can make interrupts, it can actually take multiple interrupts from the interrupt controller and queue them for their respective interrupt service routines. In various embodiments, the processor may continue draining interrupts until all interrupts have been read from interrupt controller B20, or it may read interrupts up to a certain maximum number and return to processing, or it may read interrupts until a timer expires. If the aforementioned timer times out and there are still pending interrupts, interrupt controller B20 may be configured to transition to forced state B46 and initiate a forced iteration to deliver interrupts. If the processor stops draining interrupts and there is at least one pending interrupt or a new interrupt is pending, the interrupt controller B20 may be configured to return to a soft state B42 and continue the soft iteration.

[0191] If a soft iteration is completed with a Nack response from each cluster interrupt controller B24A-B24n (and at least one interrupt remains pending), interrupt controller B20 may be configured to transition to hard state B44 and initiate a hard iteration. If cluster interrupt controllers B24A-B24n provide an Ack response during the hard iteration and there is at least one pending interrupt, interrupt controller B20 may transition to standby drain state B48 as described above. If a hard iteration is completed with a Nack response from each cluster interrupt controller B24A-B24n and there is at least one pending interrupt, interrupt controller B20 may be configured to transition to forced state B46 and initiate a forced iteration. Interrupt controller B20 may remain in forced state B46 until there are no more pending interrupts.

[0192] Figure 18 is a flowchart illustrating the operation of one embodiment of interrupt controller B20 when performing a soft or hard iteration (for example, when in state B42 or B44 in Figure 17). For ease of understanding, the blocks are shown in a specific order, but other orders may be used. The blocks can be implemented in parallel by combinational logic circuits within interrupt controller B20. The flowcharts of the blocks, combinations of blocks, and / or the whole can be pipelined over multiple clock cycles. Interrupt controller B20 may be configured to implement the operation shown in Figure 18.

[0193] The interrupt controller can be configured to select cluster interrupt controllers B24A-B24n (block B50). Any mechanism can be used to select cluster interrupt controllers B24A-B24n from multiple interrupt controllers B24A-B24n. For example, a programmable order of clusters of interrupt controllers B24A-B24n can indicate which cluster of interrupt controllers B24A-B24n is selected. In one embodiment, the order can be based on the interrupt source of a given interrupt (e.g., there may be multiple available orders, and a particular order may be selected based on the interrupt source). Such an implementation can allow different interrupt sources to prefer a given type of processor (e.g., performance-optimized or efficiency-optimized) by first attempting to deliver interrupts to the desired type of processor cluster before moving to a different type of processor cluster. In another embodiment, a delivery algorithm that is not the most recent may be used to distribute interrupts across different processor clusters by selecting the most recent cluster interrupt controller B24A-B24n (e.g., the cluster interrupt controller B24A-B24n that produced the least recent Ack response to the interrupt). In another embodiment, the most recently delivered algorithm may be used to select a cluster interrupt controller (e.g., cluster interrupt controllers B24A-B24n that most recently generated an Ack response for the interrupt) in order to take advantage of the possibility that the interrupt code or state is still cached in the processor cluster. Any mechanism or combination of mechanisms can be used.

[0194] The interrupt controller B20 can be configured to transmit interrupt requests (hard or soft, depending on the current iteration) to the selected cluster interrupt controllers B24A-B24n (block B52). For example, the interrupt controller B20 can assert a hard or soft interrupt request signal to the selected cluster interrupt controllers B24A-B24n. If the selected cluster interrupt controllers B24A-B24n provide an Ack response to the interrupt request (decision block B54, "yes" branch), the interrupt controller B20 can be configured to transition to a waiting drain state B48 to allow processor B30 in processor cluster B14A-B14n associated with the selected cluster interrupt controllers B24A-B24n to service one or more pending interrupts (block B56). If the selected cluster interrupt controller provides a Nack response (decision block B58, branch to "yes") and there is at least one cluster interrupt controller B24A-B24n that is not selected in the current iteration (decision block B60, branch to "yes"), interrupt controller B20 may be configured to select the next cluster interrupt controller B24A-B24n according to the implemented selection mechanism (block B62), return to block B52, and assert an interrupt request to the selected cluster interrupt controller B24A-B24n. Thus, in this embodiment, interrupt controller B20 may be configured to attempt to serially distribute the interrupt controller to the multiple cluster interrupt controllers B24A-B24n during iterations across multiple cluster interrupt controllers B24A-B24n.If the selected cluster interrupt controllers B24A-B24n provide a Nack response (decision block B58, branch to "yes"), and there are no more cluster interrupt controllers B24A-B24n left to be selected (e.g., all cluster interrupt controllers B24A-B24n have been selected), cluster interrupt controller B20 may be configured to transition to the next state in the state machine (e.g., to hard state B44 if the current iteration is a soft iteration, or to forced state B46 if the current iteration is a hard iteration) (block B64). If no response has yet been received for the interrupt request (decision blocks B54 and B58, branch to "no"), interrupt controller B20 may be configured to continue waiting for a response.

[0195] As described above, there may be a timeout mechanism that can be initialized when the interrupt distribution process starts. If a timeout occurs during any of the states, in one embodiment, the interrupt controller B20 may be configured to transition to the forced state B46. Alternatively, timer expiration may only be considered in the standby drain state B48.

[0196] Figure 19 is a flowchart illustrating the operation of one embodiment of cluster interrupt controllers B24A-B24n based on interrupt requests from interrupt controller B20. For ease of understanding, the blocks are shown in a specific order, but other orders may be used. The blocks can be implemented in parallel within the combinational logic circuits of the cluster interrupt controllers B24A-B24n. The flowcharts of the blocks, combinations of blocks, and / or the whole can be pipelined over multiple clock cycles. The cluster interrupt controllers B24A-B24n can be configured to implement the operation shown in Figure 19.

[0197] If the interrupt request is a hard request or a forced request (decision block B70, "yes" branch), the cluster interrupt controllers B24A-B24n may be configured to power on any powered-down (e.g., sleeping) processor B30 (block B72). If the interrupt request is a forced interrupt request (decision block B74, "yes" branch), the cluster interrupt controllers B24A-B24n may be configured to interrupt all parallel processors B30 (block B76). Ack / Nack may not apply in the case of forced interrupts, and therefore the cluster interrupt controllers B24A-B24n may continue to assert the interrupt request until at least one processor takes the interrupt. Alternatively, the cluster interrupt controllers B24A-B24n may be configured to receive an Ack response from a processor indicating that an interrupt will be made, terminate the forced interrupt, and transmit the Ack response to the interrupt controller B20.

[0198] If the interrupt request is a hard request (decision block B74, branch to "no") or a soft request (decision block B70, branch to "no"), the cluster interrupt controller may be configured to select the powered-on processor B30 (block B78). The interrupt controller B20 may use any selection mechanism similar to the above-described mechanism for selecting the cluster interrupt controllers B24A-B24n (e.g., programmable order, least recently interrupted, most recently interrupted, etc.). In one embodiment, the order may be based on the processor ID assigned to the processors in the cluster. The cluster interrupt controllers B24A-B24n may be configured to assert the interrupt request to the selected processor B30 and transmit the request to processor B30 (block B80). If the selected processor B30 provides an Ack response (decision block B82, branch to "yes"), the cluster interrupt controllers B24A-B24n may be configured to provide an Ack response to the interrupt controller B20 (block B84) and terminate the attempt to distribute the interrupt within the processor cluster. If the selected processor 30 provides a Nack response (decision block B86, branch "yes") and there is at least one powered-on processor B30 that has not yet been selected (decision block B88, branch "yes"), the cluster interrupt controllers B24A-B24n may be configured to select the next powered-on processor (block B90) (for example, according to the selection mechanism described above) and assert an interrupt request to the selected processor B30 (block B80). Thus, the cluster interrupt controllers B24A-B24n can attempt to serially deliver the interrupt to the processor B30 in the processor cluster. If there are no further powered-on processors to select (decision block B88, branch "no"), the cluster interrupt controllers B24A-B24n may be configured to provide a Nack response to the interrupt controller B20 (block B92).If the selected processor B30 has not yet provided a response (decision blocks B82 and B86, branch to "no"), the cluster interrupt controllers B24A to B24n may be configured to wait for a response.

[0199] In one embodiment, if processor B30 is powered on from a power-off state during a hard iteration, processor B30 may be readily available for interrupts because it has not yet been assigned a task by the operating system or other control software. The operating system may be configured to unmask interrupts in processor B30 that has been powered on from a power-off state as soon as possible after initializing the processor. The cluster interrupt controllers B24A-B24n may select the most recently powered-on processor first in the selection order to improve the likelihood that a processor will provide an Ack response to an interrupt.

[0200] Figure 20 is a more detailed block diagram of one embodiment of processor B30. In the illustrated embodiment, processor B30 includes a fetch and decode unit B100 (including an instruction cache or I cache, B102), a Map-Dispatch-Rename (MDR) unit B106 (including a processor interrupt acknowledgment (Int Ack) control circuit B126 and a reorder buffer B108), one or more reservation stations B110, one or more execution units B112, a register file B114, a data cache (D cache) B104, a load / store unit (LSU) B118, a reservation station (RS) B116 for the load / store unit, and a core interface unit (CIF) B122. The fetch and decode unit B100 is coupled to the MDR unit B106, and the MDR unit B106 is coupled to the reservation station B110, the reservation station B116, and the LSU B118. The reservation station B110 is coupled to the execution unit B28. The register file B114 is coupled to the execution unit B112 and LSU B118. LSU B118 is also coupled to the D cache B104, which in turn is coupled to the CIF B122 and register file B114. LSU B118 includes the store queue B120 (STQ B120) and the load queue (LDQ B124). CIF B122 is coupled to the processor Int Ack control circuit BB126 to propagate asserted interrupt requests (Int Req) to processor B30 and to propagate Ack / Nack responses from the processor Int Ack control circuit B126 to the interrupt request source (e.g., cluster interrupt controllers B24A-B24n).

[0201] The processor Int Ack control circuit B126 may be configured to determine whether processor B30 is voluntarily accepting an interrupt request transmitted to processor B30, and based on that determination, it can provide Ack and Nack indications to CIF B122. If processor B30 provides an Ack response, processor B30 is committed to performing the interrupt (and starting execution of the interrupt code to identify the interrupt and the interrupt source) within the specified period. That is, the processor Int Ack control circuit B126 may be configured to generate an acknowledgment (Ack) response to the received interrupt request based on the determination that the reorder buffer B108 has retired the instruction operation to the interruptable point and LSU B118 has completed the load / store operation to the interruptable point within the specified period. If at least one of the reorder buffer B108 and LSU B118 has been determined not to reach (or may not reach) the interruptable point within the specified period, the processor Int Ack control circuit B126 may be configured to generate a non-acknowledgment (Nack) response to the interrupt request. For example, the specific period may be about 5 microseconds in one embodiment, but may be longer or shorter in other embodiments.

[0202] In one embodiment, the processor Int Ack control circuit B126 may be configured to examine the contents of the reorder buffer 108 in order to make an initial Ack / Nack decision. That is, there may be one or more cases in which the processor Int Ack control circuit B126 can determine that a Nack response will be generated based on the state in the MDR unit B106. For example, if the reorder buffer B108 contains one or more instruction operations that have not yet been executed and have a potential execution latency greater than a certain threshold, the processor Int Ack control circuit B126 may be configured to determine that a Nack response will be generated. Some instruction operations may have variable execution latency, which may be data-dependent, memory latency-dependent, etc., and therefore the execution latency is referred to as "potential". Thus, potential execution latency may be the longest execution latency that can occur, even if it does not always occur. In other cases, potential execution latency may be the longest execution latency that occurs beyond a certain probability, etc. Examples of such instructions may include certain cryptographic acceleration instructions, certain types of floating-point instructions, or vector instructions, etc. If an instruction is not interruptible, it can be considered to have potentially long latency. That is, a non-interruptible instruction must complete execution after it has started.

[0203] Another condition that may be considered when generating an Ack / Nack response is the interrupt masking state in processor 30. When an interrupt is masked, processor B30 is prevented from performing the interrupt. A Nack response can be generated when the processor Int Ack control circuit B126 detects that an interrupt is masked within the processor (this state may be maintained in MDR unit B106 in one embodiment). More specifically, in one embodiment, the interrupt mask may have an architectural current state corresponding to the most recently retired instruction, and one or more speculative updates to the interrupt mask may also be queued. In one embodiment, a Nack response can be generated if the architectural current state is that the interrupt is masked. In another embodiment, a Nack response can be generated if the architectural current state is that the interrupt is masked, or if any of the speculative states indicate that the interrupt is masked.

[0204] Other cases may also be considered Nack response cases in the processor Int Ack control circuit B126. For example, a Nack response may be generated if there are pending redirects related to exception handling in the reorder buffer (e.g., no microarchitectural redirects such as branch prediction misses). Certain debug modes (e.g., single-step mode) and high-priority internal interrupts may also be considered Nack response cases.

[0205] If the processor Int Ack control circuit B126 does not detect a Nack response based on examining the processor state in the reorder buffer B108 and MDR unit B106, the processor Int Ack control circuit B126 can interface with LSU B118 to determine if there are any incomplete high-latency load / store instructions issued (e.g., to CIF B122 or outside of processor B30) and coupled to the reorder buffer and load / store units. For example, loads and stores into device space (e.g., loads and stores mapped to peripherals instead of memory) can potentially have high latency. If LSU B118 responds that there are high-latency load / store instructions (e.g., potentially greater than the threshold used internally by the MDR unit B106, which may be different from or the same as the threshold described above), the processor Int Ack control circuit B126 can determine that the response is a Nack. Other potentially high-latency operations may include, for example, synchronous barrier operations.

[0206] In one embodiment, if the decision is not a Nack response for the above case, LSU B118 may provide reorder buffer B108 with a pointer that identifies the oldest load / store instruction that LSU B118 has committed to completion (e.g., initiated from LDQ B124 or STQ B120, or otherwise non-speculative in LSU B118). This pointer may be referred to as a "true load / store (LS) non-speculative (NS) pointer". MDR B106 / reorder buffer B108 may attempt to interrupt the LS NS pointer, and if this is not possible within a specified period, the processor Int Ack control circuit B126 may determine that a Nack response is generated. Otherwise, an Ack response may be generated.

[0207] The fetch and decode unit B100 may be configured to fetch instructions for execution by processor B30 and decode instructions into operations for execution. More specifically, the fetch and decode unit B100 may be configured to cache previously fetched instructions from memory (via CIF B122) in I cache B102, and may be configured to fetch speculative paths of instructions for processor B30. The fetch and decode unit B100 may implement various predictive structures to predict fetch paths. For example, a next fetch predictor can be used to predict a fetch address based on previously executed instructions. Various types of branch predictors may be used to verify the next fetch prediction, or, if a next fetch predictor is not used, to predict the next fetch address. The fetch and decode unit B100 may be configured to decode instructions into instruction operations. In some embodiments, a given instruction may be decoded into one or more instruction operations, depending on the complexity of the instruction. In some embodiments, particularly complex instructions may be microcoded. In such embodiments, the microcode routine for an instruction may be coded in the instruction operation. In other embodiments, each instruction in the instruction set architecture implemented by the processor B30 may be decoded into a single instruction operation, and thus the instruction operation may be essentially synonymous with an instruction (although its form may be modified by the decoder). The term “instruction operation” is sometimes more concisely referred to as “op” in this specification.

[0208] The MDR unit B106 may be configured to map operations (ops) to speculative resources (e.g., physical registers) to enable out-of-order and / or speculative execution, and may dispatch operations to reservation stations B110 and B116. Operations may be mapped from architectural registers used in the corresponding instructions to physical registers in register file B114. That is, register file B114 may implement a set of physical registers that may be greater than the number of architectural registers specified by the instruction set architecture implemented by processor B30. The MDR unit B106 can manage the mapping of architectural registers to physical registers. In one embodiment, separate physical registers may exist for different operand types (e.g., integer, media, floating-point, etc.). In other embodiments, physical registers may be shared across operand types. The MDR unit B106 can also be responsible for tracking speculative execution and retiring operations or flushing operations that have been misinferred. A reorder buffer B108 may be used to track the program order of operations and manage retirement / flush. In other words, the reorder buffer B108 can be configured to track multiple instruction operations corresponding to instructions that are fetched by the processor but not retired by the processor.

[0209] When the source operands for the operation are ready, the operation can be scheduled to run. In the illustrated embodiment, decentralized scheduling is used for each of the execution units B28 and LSU B118, for example, at reservation stations B116 and B110. Other embodiments may implement a centralized scheduler as needed.

[0210] LSU B118 can be configured to perform load / store memory operations. Generally, memory operations (memory operations) can be instruction operations that specify access to memory (however, memory access may be completed within a cache such as the D cache B104). Load memory operations can specify the transfer of data from a memory location to a register, and store memory operations can specify the transfer of data from a register to a memory location. Load memory operations may be referred to as load memory operations, load operations, or load, and store memory operations may be referred to as store memory operations, store operations, or store. In one embodiment, a store operation can be performed as a store address operation and a store data operation. A store address operation may be defined to generate the address of a store, probe the cache for initial hit / miss determination, and update the store queue with the address and cache information. Thus, a store address operation may have an address operand as its source operand. A store data operation may be defined to deliver store data to the store queue. Thus, a store data operation may have a store data operand as its source operand, but may not have an address operand as its source operand. In many cases, the address operand of a store may be available before the store data operand, and therefore the address may be determined and made available earlier than the store data. In some embodiments, for example, if the store data operand is provided before one or more of the store address operands, the store data op may be executed before the corresponding store address op. In some embodiments, the store op may be executed as a store address and a store data op, but other embodiments may not implement store address / store data partitioning. The remainder of this disclosure often uses the store address op (and store data op) as examples, but implementations that do not use store address / store data optimization are also intended. The address generated through the execution of the store address op may be referred to as the address corresponding to the store op.

[0211] Load / store operations may be received at a reservation station B116, which may be configured to monitor the source operands of an operation to determine when they are available and then issue the operation to the load or store pipeline, respectively. Some source operands may be available when the operation is received at the reservation station B116, which may be indicated in the data received by the reservation station B116 from the MDR unit B106 for the corresponding operation. Other operands may become available through the execution of an operation by another execution unit B112, or even through the execution of a previous load operation. Operands may be collected by the reservation station B116, or they may be read from the register file B114 when issued from the reservation station B116, as shown in Figure 20.

[0212] In one embodiment, the reservation station B116 may be configured to issue load / store operations in no particular order (from the original order in the code sequence being executed by the processor B30, referred to as the "program order") as operands become available. To ensure that there is space in LDQ B124 or STQ B120 for older operations that are bypassed by newer operations in the reservation station B116, the MDR unit B106 may include circuitry that pre-allocates entries in LDQ B124 or STQ B120 for operations transmitted to the load / store unit B118. If there are no available LDQ entries for a load being processed in the MDR unit B106, the MDR unit B106 may stall dispatching load operations and subsequent operations in program order until one or more LDQ entries become available. Similarly, if there are no available STQ entries for a store, the MDR unit B106 may stall operation dispatching until one or more STQ entries become available. In another embodiment, the reservation station B116 can issue operations in program order, and LRQ B46 / STQ B120 assignments can occur when issued from the reservation station B116.

[0213] LDQ B124 can track loads from the initial execution by LSU B118 to retirement. LDQ B124 can ensure that memory ordering rules are not violated (between loads executed in no particular order, and between loads and stores). If a memory ordering violation is detected, LDQ B124 can signal a redirection for the corresponding load. The redirection causes processor B30 to flush the load and subsequent operations in program order and refetch the corresponding instructions. Speculative state for loads and subsequent operations may be discarded, and operations may be refetched by the fetch and decryption unit B100 and reprocessed for execution again.

[0214] When a load / store address op is issued by reservation station B116, LSU B118 can be configured to generate the address accessed by the load / store, and can be configured to translate the address from the valid or virtual address created from the address operand of the load / store address op to the physical address actually used to address the memory. LSU B118 can be configured to generate access to D cache B104. For load operations that hit D cache B104, data may be speculatively transferred from D cache B104 to the destination operand of the load operation (e.g., a register in register file B114), unless the address hits a preceding operation in STQ B120 (i.e., an earlier store in program order) or the load is replayed. The data may also be speculatively scheduled and transferred to a dependent op in execution unit B28. In this case, execution unit B28 may bypass the transferred data instead of the data output from register file B114. If store data is available for transfer upon an STQ hit, the data output by STQ B120 may be transferred instead of cache data. Cache misses and STQ hits that prevent data transfer may be reasons for replay, and load data may not be transferred in those cases. Cache hit / miss status from D cache B104 may be logged to STQ B120 or LDQ B124 for subsequent processing.

[0215] LSU B118 can implement multiple load pipelines. For example, in one embodiment, three load pipelines ("pipes") may be implemented, but in other embodiments, more or fewer pipelines may be implemented. Each pipeline can independently and in parallel with other loads to execute different loads. That is, RS B116 can issue any number of loads up to the number of load pipes in the same clock cycle. LSU B118 can also implement one or more store pipes, and in particular, multiple store pipes. However, the number of store pipes does not have to be equal to the number of load pipes. In one embodiment, for example, two store pipes can be used. The reservation station B116 can issue store address ops and store data ops to the store pipes independently and in parallel. The store pipes can be coupled to STQ B120, which can be configured to hold executed but uncommitted store operations.

[0216] CIF B122 can take on the role of communicating with the rest of the system, including processor B30, on behalf of processor B30. For example, CIF B122 may be configured to request data for D cache B104 misses and I cache B102 misses. Once data is returned, CIF B122 can signal a cache fill to the corresponding cache. In the case of a D cache fill, CIF B122 can also notify LSU B118. LDQ B124 can attempt to schedule a reclaimed load waiting for a cache fill, so that the reclaimed load can transfer the fill data when the fill data is provided to D cache B104 (referred to as a fill transfer operation). If a reclaimed load is not successfully reclaimed during a fill, the reclaimed load can then be scheduled and reclaimed via D cache B104 as a cache hit. CIF B122 can also write back corrected cache lines that were rejected by D cache B104, merge store data for non-cacheable stores, and so on.

[0217] Execution unit B112 may include any type of execution unit in various embodiments. For example, execution unit B112 may include integer, floating-point, and / or vector execution units. An integer execution unit may be configured to perform integer operations. Generally, an integer operation is an operation that performs operations defined on integer operands (e.g., arithmetic, logical, shift / rotation, etc.). Integers may be numerical values ​​where each value corresponds to a mathematical integer. An integer execution unit may include branching hardware for handling branch operations, or there may be a separate branch execution unit.

[0218] A floating-point execution unit may be configured to perform floating-point operations. Generally, a floating-point operation may be an operation defined to perform operations on a floating-point operand. A floating-point operand is an operand expressed as the base raised to the power of an exponent multiplied by a mantissa (or mantissa part). The exponent, the sign of the operand, and the mantissa / mantissa part may be explicitly expressed in the operand, and the base may be implicit (for example, base 2 in one embodiment).

[0219] A vector execution unit can be configured to perform vector operations. Vector operations can be used, for example, to process media data (e.g., image data such as pixels, audio data, etc.). Media processing can be characterized by performing the same operation on a substantial amount of data, each of which is a relatively small value (e.g., 8 bits or 16 bits compared to 32 bits to 64 bits for integers). Thus, a vector operation may involve single-instruction multiple data (SIMD) or vector operations on operands representing multiple media data.

[0220] Therefore, each execution unit B112 may have hardware configured to perform defined actions for operations defined to be handled by a particular execution unit. Execution units may be independent of each other in general, in the sense that each execution unit may be configured to operate based on operations issued to it, without depending on other execution units. Alternatively, each execution unit may be an independent pipe for executing operations. Different execution units may have different execution latencies (e.g., different pipe lengths). Furthermore, different execution units may have different latencies for the pipeline stage in which bypass occurs, and therefore the clock cycles in which speculative scheduling of dependent operations occurs based on the load operations may vary based on the type of operation and the execution unit B28 executing the operation.

[0221] It should be noted that any number and type of execution units B112 may be included in various embodiments, including embodiments having one execution unit and embodiments having multiple execution units.

[0222] A cache line may be a unit of allocation / deallocation within a cache. That is, data within a cache line can be allocated / deallocated within the cache as a unit. Cache lines may vary in size (e.g., 32 bytes, 64 bytes, 128 bytes, or larger or smaller cache lines). Different caches may have different cache line sizes. Cache I B102 and Cache D B104 may each have any desired capacity, cache line size, and configuration. In various embodiments, there may be more additional levels of cache between Cache D B104 / Cache I B102 and main memory.

[0223] At various points in time, load / store operations may be referred to as newer or older than other load / store operations. If the first operation follows the second operation in program order, the first operation may be earlier than the second operation. Similarly, if the first operation precedes the second operation in program order, the first operation may be older than the second operation.

[0224] Figure 21 is a block diagram of one embodiment of the reorder buffer B108. In the illustrated embodiment, the reorder buffer B108 includes multiple entries. Each entry may correspond to an instruction, an instruction operation, or a group of instruction operations in various embodiments. Various states related to instruction operations may be stored in the reorder buffer (e.g., target logic and physical registers for updating the register map on the architecture, exceptions or redirects detected during execution, etc.).

[0225] Multiple pointers are shown in Figure 21. The retired pointer B130 can point to the oldest unretired op in processor B30. That is, ops prior to the op at retired B130 have been retired from the reorder buffer B108, and the architecture state of processor B30 has been updated to reflect the execution of the retired op, etc. The resolved pointer B132 can point to the oldest op where the preceding branch instruction has been resolved as correctly predicted, and the preceding op that could cause an exception has been resolved so that it does not cause an exception. Ops between the retired pointer B130 and the resolved pointer B132 may be ops committed in the reorder buffer B108. That is, the execution of the instruction that generated the op is completed up to the resolved pointer B132 (if there is no external interrupt). The newest pointer B134 can point to the op most recently fetched and dispatched from the MDR unit B106. Ops between the resolved pointer B132 and the newest pointer B134 are speculative and may be flushed due to exceptions, branch prediction misses, etc.

[0226] The true LS NS pointer B136 is the true LS NS pointer described above. The true LS NS pointer can only be generated when an interrupt request is asserted and other tests for the Nack response are negative (e.g., the Ack response is indicated by those tests). The MDR unit B106 can attempt to return the resolved pointer B132 back to the true LS NS pointer B136. There may be committed operations in the reorder buffer B108 that cannot be flushed (e.g., once they are committed, they must be completed and retired). Some groups of instruction operations may not be interruptible (e.g., microcode routines, certain non-interruptible exceptions, etc.). In such cases, the processor Int Ack controller B126 may be configured to generate a Nack response. There may be operations or combinations of operations that are too complex to "undo" in processor B30, and the existence of such operations between the resolved pointer and the true LS NS pointer B136 may cause the processor Int Ack controller B126 to generate a Nack response. If the reorder buffer B108 successfully returns the resolution pointer to the true LS NS pointer B136, the processor Int Ack control circuit B126 may be configured to generate an Ack response.

[0227] Figure 22 is a flowchart illustrating the operation of one embodiment of the processor Int Ack control circuit B126 based on the reception of an interrupt request by processor B30. For ease of understanding, the blocks are shown in a specific order, but other orders may be used. The blocks can be implemented in parallel by combinational logic circuits within the processor Int Ack control circuit B126. The flowcharts of the blocks, combinations of blocks, and / or the whole can be pipelined over multiple clock cycles. The processor Int Ack control circuit B126 may be configured to implement the operation shown in Figure 22.

[0228] The processor Int Ack control circuit B126 may be configured to determine whether there is any Nack condition detected in the MDR unit B106 (determination block B140). For example, potentially long-latency operations such as incomplete or masked interrupts may be Nack conditions detected in the MDR unit B106. If so (determination block B140, "yes" branch), the processor Int Ack control circuit B126 may be configured to generate a Nack response (block B142). Otherwise (determination block B140, "no" branch), the processor Int Ack control circuit B126 may communicate with the LSU to request a Nack condition and / or a true LS NS pointer (block B144). If the LSU B118 detects a Nack condition (determination block B146, "yes" branch), the processor Int Ack control circuit B126 may be configured to generate a Nack response (block B142). If LSU B118 does not detect a Nack condition (decision block B146, "no" branch), the processor Int Ack control circuit B126 may be configured to receive a true LS NS pointer from LSU B118 (block B148) and may attempt to return the resolved pointer in the reorder buffer B108 to the true LS NS pointer (block B150). If the move is unsuccessful (e.g., there is at least one instruction operation between the true LS NS pointer and the resolved pointer that cannot be flushed) (decision block B152, "no" branch), the processor Int Ack control circuit B126 may be configured to generate a Nack response (block B142). Otherwise (decision block B152, "yes" branch), the processor Int Ack control circuit B126 may be configured to generate an Ack response (block B154). The processor Int Ack control circuit B126 may be configured to freeze the resolution pointer to a true LS NS pointer and retire the op until the retired pointer reaches the resolution pointer (block B156). The processor Int Ack control circuit B126 may then be configured to generate an interrupt (block B158).In other words, processor B30 can begin fetching interrupt code (for example, from a predetermined address related to the interrupt according to the instruction set architecture implemented by processor B30).

[0229] In another embodiment, SOC B10 may be one of several SOCs in the system. More specifically, in one embodiment, multiple instances of SOC B10 may be employed. Other embodiments may have an asymmetric SOC. Each SOC may be a separate integrated circuit chip (for example, mounted on a separate semiconductor substrate or "die"). The dies may be packaged and connected to each other via an interposer, package-on-package solution, etc. Alternatively, the dies may be packaged in a chip-on-chip package solution, a multi-chip module, etc.

[0230] Figure 23 is a block diagram showing one embodiment of a system including multiple instances of SOC B10. For example, SOC B10A, SOC B10B, etc. ~ SOC B10q may be combined together in the system. Each SOC B10A~B10q includes an instance of interrupt controller B20 (for example, interrupt controller B20A, interrupt controller B20B, and interrupt controller B20q in Figure 23). One interrupt controller, interrupt controller B20A in this example, can function as the primary interrupt controller of the system. The other interrupt controllers B20B~B20q can function as secondary interrupt controllers.

[0231] The interface between the primary interrupt controller B20A and the secondary controller B20B is shown in detail in Figure 23, and the interface between the primary interrupt controller B20A and other secondary interrupt controllers such as interrupt controller B20q may be similar. In the embodiment of Figure 23, the secondary controller B20B is configured to provide interrupt information as Int B160 that identifies interrupts issued from an interrupt source on SOC B10B (or an external device coupled to SOC B10B, not shown in Figure 23). The primary interrupt controller B20A is configured to signal hard repetitions, soft repetitions, and forced repetitions to the secondary interrupt controller B20B (reference code B162) and to receive Ack / Nack responses from interrupt controller B20B (reference code B164). The interface can be implemented in any way. For example, dedicated wires may be coupled between SOC B10A and SOC B10B to implement reference codes B160, B162, and / or B164. In another embodiment, messages may be exchanged between the primary interrupt controller B20A and the secondary interrupt controllers B20B-B20q via a general-purpose interface between SOCs B10A-B10q, which is also used for other communications. In one embodiment, programmed input / output (PIO) writes may be used as data, along with interrupt data, hard / soft / force requests, and Ack / Nack responses, respectively.

[0232] The primary interrupt controller B20A may be configured to collect interrupts from various interrupt sources that may be on SOC B10A, one of the other SOCs B10B-B10q which may be off-chip devices, or any combination thereof. The secondary interrupt controllers B20B-B20q may be configured to transmit interrupts to the primary interrupt controller B20A (Int in Figure 23) and to identify the interrupt sources to the primary interrupt controller B20A. The primary interrupt controller B20A may also be responsible for ensuring the delivery of interrupts. The secondary interrupt controllers B20B-B20q may be configured to receive instructions from the primary interrupt controller B20A, receive soft, hard, and forced repetition requests from the primary interrupt controller B20A, and perform repetitions via cluster interrupt controllers B24A-B24n which are implemented on the corresponding SOCs B10B-B10q. Based on the Ack / Nack responses from the cluster interrupt controllers B24A to B24n, the secondary interrupt controllers B20B to B20q can provide Ack / Nack responses. In one embodiment, the primary interrupt controller B20A can attempt to distribute interrupts serially through the secondary interrupt controllers B20B to B20q in soft and hard iterations, and can distribute them in parallel to the secondary interrupt controllers B20B to B20q in forced iterations.

[0233] In one embodiment, the primary interrupt controller B20A may be configured to perform a given iteration on a subset of cluster interrupt controllers integrated into the same SOC B10A as the primary interrupt controller B20A, before performing a given iteration on a subset of cluster interrupt controllers on other SOCs B10B-B10q (with the assistance of secondary interrupt controllers B20B-B20q). That is, the primary interrupt controller B20A can attempt to distribute interrupts serially through the cluster interrupt controllers on SOC B10A, and then communicate with the secondary interrupt controllers BB20B-B20q. Attempts to distribute through the secondary interrupt controllers B20B-B20q may also be performed serially. The order of attempts through the secondary interrupt controllers BB20-B20q can be determined in any desired way (e.g., programmable order, least recent accepted order, most recently accepted order, etc.), similar to the embodiments described above for cluster interrupt controllers and processors in a cluster. Therefore, the primary interrupt controller B20A and secondary interrupt controllers B20B-B20q can be largely isolated from the presence of multiple SOCs B10A-B10q. That is, SOCs B10A-B10q can be configured as a single system that is almost transparent to software execution on a single system. During system initialization, some embodiments may be programmed to configure the interrupt controllers B20A-B20q as described above, but otherwise, the interrupt controllers B20A-B20q can manage the distribution of interrupts across potentially multiple SOCs B10A-B10q, each on a separate semiconductor die, without software assistance or specific software visibility into the multi-die nature of the system. For example, delays due to inter-die communication may be minimized in the system. Thus, during post-initialization execution, a single system may appear as a single system to the software, and the multi-die nature of the system may be transparent to the software.

[0234] It should be noted that the primary interrupt controller B20A and the secondary interrupt controllers B20B-B20q can operate in a manner that is also referred to by those skilled in the art as "master" (i.e., primary) and "slave" (i.e., secondary). Although the terms primary and secondary are used herein, it is expressly intended that the terms "primary" and "secondary" are to be interpreted as encompassing these corresponding terms.

[0235] In one embodiment, each instance of SOC B10A-B10q may have both a primary and a secondary interrupt controller circuit implemented within its interrupt controller B20A-B20q. One interrupt controller (e.g., interrupt controller B20A) may be designated primary during system manufacturing (e.g., via a fuse on SOC B10A-B10q or via a pin strap on one or more pins of SOC B10A-B10q). Alternatively, primary and secondary designations may be made during the system initialization (or boot) configuration.

[0236] Figure 24 is a flowchart illustrating the operation of one embodiment of a primary interrupt controller B20A based on the reception of one or more interrupts from one or more interrupt sources. For ease of understanding, the blocks are shown in a specific order, but other orders may be used. The blocks can be implemented in parallel within the combinational logic circuits of the primary interrupt controller B20A. The flowcharts of the blocks, combinations of blocks, and / or the whole can be pipelined over multiple clock cycles. The primary interrupt controller B20A may be configured to implement the operation shown in Figure 24.

[0237] The first interrupt controller B20A can be configured to perform a soft iteration on the cluster interrupt controller integrated on the local SOC B10A (block B170). For example, the soft iteration may be similar to the flowchart of FIG. 18. When the local soft iteration results in an Ack response (decision block B172, "yes" branch), the interrupt can be delivered normally, and the first interrupt controller B20A can be configured to return to the idle state B40 (assuming that there are no more pending interrupts). When the local soft iteration results in a Nack response (decision block B172, "no" branch), the first interrupt controller B20A can be configured to select one of the other SOCs B10B to B10q using any desired order as described above (block B174). The first interrupt controller B20A can be configured to assert a soft iteration request to the secondary interrupt controllers B20B to B20q on the selected SOCs B10B to B10q (block B176). When the secondary interrupt controllers B20B to B20q provide an Ack response (decision block B178, "yes" branch), the interrupt can be delivered normally, and the first interrupt controller B20A can be configured to return to the idle state B40 (assuming that there are no more pending interrupts). When the secondary interrupt controllers B20B to B20q provide a Nack response (decision block B178, "no" branch) and there are further SOCs B10B to B10q that have not yet been selected in the soft iteration (decision block B180, "yes" branch), the first interrupt controller B20A can be configured to select the next SOCs B10B to B10q according to the implemented ordering mechanism (block B182), transmit a soft iteration request to the secondary interrupt controllers B20B to B20q on the selected SOCs (block B176), and be configured to continue the process. On the other hand, when each SOC B10B to B10q has been selected, the soft iteration can be completed because the serial attempt to deliver the interrupt via the secondary interrupt controllers B20B to B20q has been completed.

[0238] Based on the completion of soft iterations to secondary interrupt controllers B20B-B20q without successfully delivering interrupts (decision block B180, branch to "no"), primary interrupt controller B20A may be configured to perform hard iterations to local cluster interrupt controllers integrated on local SOC B10A (block B184). For example, the soft iterations may be similar to the flowchart in Figure 18. If the local hard iteration results in an Ack response (decision block B186, branch to "yes"), the interrupt may be successfully delivered, and primary interrupt controller B20A may be configured to return to idle state B40 (assuming there are no more pending interrupts). If the local hard iteration results in a Nack response (decision block B186, branch to "no"), primary interrupt controller B20A may be configured to select one of the other SOCs B10B-B10q using any desired order as described above (block B188). The primary interrupt controller B20A may be configured to assert a hard iteration request to the secondary interrupt controllers B20B~B20q on the selected SOC B10B~B10q (block B190). If the secondary interrupt controllers B20B~B20q provide an Ack response (decision block B192, branch to "yes"), the interrupt may be successfully delivered, and the primary interrupt controller B20A may be configured to return to idle state B40 (assuming there are no more pending interrupts). If secondary interrupt controllers B20B~B20q provide a Nack response (decision block B192, branch to "no"), and there are further SOCs B10B~B10q that have not yet been selected in the hard iteration (decision block B194, branch to "yes"), primary interrupt controller B20A may be configured to select the next SOC B10B~B10q according to the implemented ordering mechanism (block B196), transmit a hard iteration request to the secondary interrupt controllers B20B~B20q on the selected SOC (block B190), and continue processing.On the other hand, when each of the SOCs B10B to B10q is selected, since the serial trials for distributing interrupts via the secondary interrupt controllers B20B to B20q are completed, the hard iterations can be completed (decision block B194, branch of "No"). The primary interrupt controller B20A may be configured to proceed to forced iteration (block B198). The forced iteration can be performed locally, or can be performed in parallel or serially across the local SOC B10A and the other SOCs B10B to B10q.

[0239] As described above, there may be a timeout mechanism that can be initialized when the interrupt distribution process starts. If a timeout occurs during any state, in one embodiment, the interrupt controller B20 can be configured to move to forced iteration. Alternatively, the timer expiration can also be considered only in the standby drain state B48 as described above.

[0240] FIG. 25 is a flowchart showing the operation of an embodiment of the secondary interrupt controllers B20B to B20q. For ease of understanding, the blocks are shown in a specific order, but other orders may be used. The blocks can be performed in parallel within the combinational logic circuit in the secondary interrupt controllers B20B to B20q. The blocks, combinations of blocks, and / or the flowchart as a whole can be pipelined over multiple clock cycles. The secondary interrupt controllers B20B to B20q may be configured to implement the operations shown in FIG. 25.

[0241] When an interrupt source within the corresponding SOCs B10B to B10q (or coupled to the SOCs B10B to B10q) provides an interrupt to the secondary interrupt controllers B20B to B20q (decision block B200, branch of "Yes"), the secondary interrupt controllers B20B to B20q can be configured to transmit the interrupt to the primary interrupt controller B20A for processing together with other interrupts from other interrupt sources (block B202).

[0242] If the primary interrupt controller B20A transmits an iteration request (decision block B204, "yes" branch), the secondary interrupt controllers B20B-B20q may be configured to perform the requested iteration (hard, soft, or forced) to the cluster interrupt controllers in the local SOC B10B-B10q (block B206). For example, hard and soft iterations may be as shown in Figure 18, and forced iterations may be performed in parallel to the cluster interrupt controllers in the local SOC B10B-B10q. If the iteration results in an Ack response (decision block B208, "yes" branch), the secondary interrupt controllers B20B-B20q may be configured to transmit an Ack response to the primary interrupt controller B20A (block B210). If the iteration results in a Nack response (decision block B208, "no" branch), the secondary interrupt controllers B20B-B20q may be configured to transmit a Nack response to the primary interrupt controller B20A (block B212).

[0243] Figure 26 is a flowchart illustrating one embodiment of how interrupts are handled. For ease of understanding, the blocks are shown in a specific order, but other orders may be used. The blocks may be executed in parallel within the combinational logic circuits of the system described herein. The blocks, combinations of blocks, and / or the flowchart as a whole may be pipelined over multiple clock cycles. The system described herein may be configured to implement the operation illustrated in Figure 26.

[0244] Interrupt controller B20 can receive interrupts from an interrupt source (block B220). In embodiments having primary and secondary interrupt controllers B20A-B20q, interrupts may be received by any of the interrupt controllers B20A-B20q and provided to primary interrupt controller B20A as part of receiving interrupts from an interrupt source. Interrupt controller B20 may be configured to perform a first iteration (e.g., a soft iteration) that attempts serially to distribute interrupts to a plurality of cluster interrupt controllers (block B222). Individual cluster interrupt controllers of the plurality of cluster interrupt controllers are associated with individual processor clusters containing a plurality of processors. A given cluster interrupt controller of the plurality of cluster interrupt controllers may be configured in the first iteration to attempt to distribute interrupts to a subset of each plurality of processors that are powered on, but not to attempt to distribute interrupts to any subset of each plurality of processors that are not included in the subset. If an Ack response is received, the iteration may be terminated by interrupt controller B20 (decision block B224, "yes" branch and block B226). On the other hand (decision block B224, "no" branch), based on the non-acknowledgment (Nack) responses from multiple cluster interrupt controllers in the first iteration, the interrupt controller may be configured to perform a second iteration (e.g., a hard iteration) across the multiple cluster interrupt controllers (block B228). A given cluster interrupt controller may be configured in the second iteration to power on some of the multiple processors that are powered off and attempt to deliver interrupts to each of the multiple processors. If an Ack response is received, the iteration may be terminated by interrupt controller B20 (decision block B230, "yes" branch and block B232).On the other hand (decision block B230, "no" branch), based on the non-acknowledgment (Nack) responses from multiple cluster interrupt controllers in the second iteration, the interrupt controller may be configured to perform a third iteration (e.g., a forced iteration) across the multiple cluster interrupt controllers (block B234).

[0245] Based on this disclosure, a system may comprise a plurality of cluster interrupt controllers and an interrupt controller coupled to the plurality of cluster interrupt controllers. An individual cluster interrupt controller of the plurality of cluster interrupt controllers may be associated with an individual processor cluster comprising a plurality of processors. An interrupt controller may be configured to receive an interrupt from a first interrupt source, perform a first iteration across the plurality of cluster interrupt controllers to attempt to deliver the interrupt based on the interrupt, and perform a second iteration across the plurality of cluster interrupt controllers based on non-acknowledgment (Nack) responses from the plurality of cluster interrupt controllers in the first iteration. A given cluster interrupt controller of the plurality of cluster interrupt controllers may be configured in the first iteration to attempt to deliver an interrupt to a subset of a plurality of processors in an individual processor cluster that is powered on, without attempting to deliver the interrupt to any of the plurality of processors in an individual cluster that is not included in the subset. In the second iteration, a given cluster interrupt controller may be configured to power on any of the plurality of processors that are powered off and attempt to deliver an interrupt to each of the plurality of processors. In one embodiment, during an attempt to distribute an interrupt across multiple cluster interrupt controllers, the interrupt controller may be configured to assert a first interrupt request to a first cluster interrupt controller of the multiple cluster interrupt controllers, and based on the Nack response from the first cluster interrupt controller, the interrupt controller may be configured to assert a second interrupt request to a second cluster interrupt controller of the multiple cluster interrupt controllers. In one embodiment, based on a second Nack response from the second cluster interrupt controller, during an attempt to distribute an interrupt across multiple cluster interrupt controllers, the interrupt controller may be configured to assert a third interrupt request to a third cluster interrupt controller of the multiple cluster interrupt controllers.In one embodiment, during an attempt to distribute an interrupt through multiple cluster interrupt controllers, the interrupt controller may be configured to terminate the attempt based on an acknowledgment (Ack) response from a second cluster interrupt controller and the absence of additional pending interrupts. In one embodiment, during an attempt to distribute an interrupt through multiple cluster interrupt controllers, the interrupt controller may be configured to assert an interrupt request to a first cluster interrupt controller of the multiple cluster interrupt controllers, and the interrupt controller may be configured to terminate the attempt based on an acknowledgment (Ack) response from the first cluster interrupt controller and the absence of additional pending interrupts. In one embodiment, during an attempt to distribute an interrupt across multiple cluster interrupt controllers, the interrupt controller may be configured to serially assert an interrupt request to one or more cluster interrupt controllers of the multiple cluster interrupt controllers, which is terminated by an acknowledgment (Ack) response from a first cluster interrupt controller of one or more cluster interrupt controllers. In one embodiment, the interrupt controller may be configured to assert serially in a programmable order. In one embodiment, the interrupt controller may be configured to serially assert an interrupt request based on a first interrupt source. A second interrupt from a second interrupt source may result in serial assertions in a different order. In one embodiment, during an attempt to distribute interrupts through multiple cluster interrupt controllers, the interrupt controller may be configured to assert an interrupt request to a first cluster interrupt controller of the multiple cluster interrupt controllers, and the first cluster interrupt controller may be configured to serially assert a processor interrupt request to multiple processors in a separate processor cluster based on the interrupt request to the first cluster interrupt controller. In one embodiment, the first cluster interrupt controller is configured to terminate the serial assertion based on an acknowledgment (Ack) response from a first processor among the multiple processors.In one embodiment, the first cluster interrupt controller may be configured to transmit an Ack response to the interrupt controller based on an Ack response from the first processor. In one embodiment, the first cluster interrupt controller may be configured to provide a Nack response to the interrupt controller based on Nack responses from multiple processors in separate clusters during serial assertion of a processor interrupt. In one embodiment, the interrupt controller may be included on a first integrated circuit on a first substrate which includes a first subset of the multiple cluster interrupt controllers. A second subset of the multiple cluster interrupt controllers may be implemented on a second integrated circuit on a second separate semiconductor substrate. The interrupt controller may be configured to serially assert interrupt requests to the first subset before attempting to deliver them to the second subset. In one embodiment, the second integrated circuit may include a second interrupt controller which may be configured to communicate interrupt requests to the second interrupt controller in response to the first subset rejecting the interrupt. The second interrupt controller may be configured to attempt to deliver the interrupt to the second subset.

[0246] In one embodiment, the processor includes a reorder buffer, a load / store unit, and a control circuit coupled to the reorder buffer and the load / store unit. The reorder buffer may be configured to track a plurality of instruction operations corresponding to instructions fetched by the processor and not retired by the processor. The load / store unit may be configured to perform load / store operations. The control circuit may be configured to generate an acknowledgment (Ack) response to an interrupt request received by the processor based on a determination that the reorder buffer has retired an instruction operation to an interruptable point and the load / store unit has completed a load / store operation to the interruptable point within a specified period. The control circuit may be configured to generate a non-acknowledgment (Nack) response to an interrupt request based on a determination that at least one of the reorder buffer and the load / store unit does not reach an interruptable point within a specified period. In one embodiment, the determination may be a Nack response based on a reorder buffer having at least one instruction operation having a potential execution latency greater than a threshold. In one embodiment, the determination may be a Nack response based on a reorder buffer having at least one instruction operation that masks the interrupt. In one embodiment, the determination is a Nack response based on a load / store unit having at least one load / store operation on an unprocessed device address space.

[0247] In one embodiment, the method includes receiving an interrupt from a first interrupt source in an interrupt controller. The method may further include performing a first iteration in which it attempts serially to distribute the interrupt to a plurality of cluster interrupt controllers. In the first iteration, an individual cluster interrupt controller of a plurality of cluster interrupt controllers associated with an individual processor cluster having a plurality of processors may be configured to attempt to distribute an interrupt to a subset of a plurality of processors in an individual processor cluster that are powered on, without attempting to distribute the interrupt to a subset of a plurality of processors in an individual processor cluster that are not included in the subset. The method may further include the interrupt controller performing a second iteration to the plurality of cluster interrupt controllers based on non-acknowledgment (Nack) responses from the plurality of cluster interrupt controllers in the first iteration. In the second iteration, a given cluster interrupt controller may be configured to power on a subset of a plurality of processors that are powered off in an individual processor cluster and attempt to distribute an interrupt to the plurality of processors. In one embodiment, the attempt to serially distribute the interrupt to the plurality of cluster interrupt controllers is terminated based on an acknowledgment from one of the plurality of cluster interrupt controllers. Coherence

[0248] Next, referring to Figures 27 to 43, various embodiments of the cache coherency mechanism that may be implemented in embodiments of SOC10 are shown. In one embodiment, the coherency mechanism may include a plurality of directories configured to track the coherency state of a subset of the integrated memory address space. The plurality of directories are distributed throughout the system. In one embodiment, the plurality of directories are distributed across memory controllers. In one embodiment, a given memory controller of one or more memory controller circuits includes a directory configured to track a plurality of cache blocks corresponding to data in a portion of the system memory to which the given memory controller interfaces, the directory is configured to track which of the plurality of caches in the system caches a given cache block of the plurality of cache blocks, and the directory is accurate with respect to memory requests that have been ordered and processed in the directory, even if the memory request has not yet been completed in the system. In one embodiment, a given memory controller is configured to issue one or more coherency maintenance commands for a given cache block based on a memory request for that cache block, the one or more coherency maintenance commands include the cache state for the given cache block in a corresponding cache among a plurality of caches, and the corresponding cache is configured to delay processing the given coherency maintenance command based on the fact that the cache state in the corresponding cache does not match the cache state in the given coherency maintenance command. In one embodiment, a first cache is configured to store a given cache block in a primary shared state, and a second cache is configured to store a given cache block in a secondary shared state, and a given memory controller is configured to cause the first cache to transfer the given cache block to the requester based on a memory request and the primary shared state in the first cache.In one embodiment, a given memory controller is configured to issue a first coherency maintenance command and one of a second coherency maintenance command to a first cache among a plurality of caches based on the type of a first memory request, the first cache is configured to transfer a first cache block to the requester that issued the first memory request based on the first coherency maintenance command, and the first cache is configured to return the first cache block to the given memory controller based on the second coherency maintenance command.

[0249] This provides a scalable cache coherency protocol for systems including multiple coherent agents coupled to one or more memory controllers. A coherent agent may generally include a cache for caching memory data, or, otherwise, any circuitry that can acquire ownership of one or more cache blocks and potentially modify the cache blocks locally. The coherent agents participate in the cache coherency protocol to ensure that modifications made by one coherent agent are visible to other agents later reading the same data, and that modifications made by two or more coherent agents in a specific order (determined by an ordering point in the system, such as a memory controller of the memory storing the cache blocks) are observed in that order by each of the coherent agents.

[0250] A cache coherency protocol can specify a set of messages or commands that can be transmitted between an agent and a memory controller (or a coherency controller within a memory controller) to complete a coherent transaction. Messages may include requests, snoops, snoop responses, and completions. A “request” is a message that initiates a transaction and specifies the requested cache block (e.g., having the address of the cache block) and the state in which the requester will accept the cache block (or a minimum state, which may in some cases provide a more permissive state). As used herein, a “snoop” or “snoop message” refers to a message transmitted to a coherent agent to request a state change within a cache block, and may also request that the cache block be provided by the coherent agent if the coherent agent has an exclusive copy of the cache block or otherwise bears the cache block. A snoop message may be an example of a coherency maintenance command, which may be any command transmitted to a particular coherence agent to cause a change in the coherent state of a cache line within that particular coherence agent. Another term that is an example of a coherency maintenance command is a probe. Coherency maintenance commands do not refer to broadcast commands sent to all coherency agents, such as those often used in shared bus systems. The term "snoop" is used as an example below, but it should be understood that this term generally refers to coherency maintenance commands. A "completion" or "snoop response" may be a message from a coherent agent indicating that a state change has occurred and, if applicable, providing a copy of the cache block. In some cases, completion may be provided by the request source of a particular request.

[0251] The “State” or “Cache State” can generally refer to a value indicating whether a copy of a cache block is valid in the cache, and can also indicate other attributes of the cache block. For example, the State can indicate whether the cache block has been modified with respect to its copy in memory. The State can also indicate the level of ownership of the cache block (e.g., whether the agent with the cache is allowed to modify the cache block, whether the agent is responsible for providing the cache block, or returning the cache block to the memory controller if it is removed from the cache). The State may also indicate the possibility of the cache block existing in other coherent agents (for example, the “Shared” State may indicate that a copy of the cache block may be stored in one or more other cacheable agents).

[0252] The diverse embodiments of the cache coherency protocol can include a variety of features. For example, each memory controller(s) may implement a coherency controller and a directory for cache blocks corresponding to the memory controlled by that memory controller. The directory can track the state of cache blocks in multiple cacheable agents, allowing the coherency controller to determine which cacheable agents should snoop to change the state of a cache block and, if necessary, provide a copy of the cache block. That is, the snoop does not need to be broadcast to all cacheable agents based on a request received by the cache controller; rather, the snoop can be transmitted to agents that have a copy of the cache block affected by the request. Once a snoop is generated, after it is processed and the data has been provided to the source of the request, the directory can be updated to reflect the state of the cache block in each coherent agent. Thus, the directory can be accurate for subsequent requests processed for the same cache block. Snooping can be minimized, and traffic on the interconnect between coherent agents and memory controllers is reduced, for example, compared to broadcast solutions. In one embodiment, a "3-hop" protocol can be supported in which one of the caching coherent agents provides a copy of the cached block to the source of the request, or, if there is no caching agent, the memory controller provides the copy. Thus, the data is provided in three "hops" (or messages transmitted through the interface): a request from the source to the memory controller, a snoop to a coherent agent responding to the request, and completion by the cached block of data from the coherent agent to the source of the request. If there is no cached copy, there may be two hops: a request from the source to the memory controller and completion of the data from the memory controller to the source.Additional messages may be present (for example, completion from other agents indicating that the requested state change has been made, when there are multiple snoops for a request), but the data itself may be delivered in three hops. In contrast, many cache coherency protocols are four-hop protocols in which a coherent agent responds to a snoop by returning a cache block to a memory controller, which then forwards the cache block to the source. In one embodiment, a four-hop flow may be supported by the protocol in addition to the three-hop flow.

[0253] In one embodiment, a request for a cache block may be processed by a coherence controller, and when a snoop (and / or completion from the memory controller if no cached copy exists) is generated, the directory may be updated. Another request for the same cache block can then be served. Thus, requests for the same cache block may not be serialized, as in some other cache coherence protocols. Since messages related to subsequent requests may arrive at a given coherent agent before messages related to preceding requests ("subsequent" and "preceding" refer to requests ordered in the coherence controller within the memory controller), various race conditions can arise when there are multiple pending requests for a cache block. To enable the agent to sort requests, messages (e.g., snoop and completion) may include the expected cache state at the receiving agent, as indicated by the directory when the request is processed. Thus, if the receiving agent does not have a cache block in the state indicated in the message, the receiving agent may delay processing the message until the cache state changes to the expected state. The change to the expected state may occur via messages related to the previous request. Further explanation of the competitive conditions and how to use the expected cash state to resolve them is provided below with respect to Figures 29-30 and 32-34.

[0254] In one embodiment, the cache state may include a primary shared state and a secondary shared state. The primary shared state can be applied to a coherent agent responsible for transmitting a copy of the cache block to the requesting agent. The secondary shared agent may not even need to snoop during the processing of a given request (e.g., reading a cache block that is permitted to be returned in a shared state). Further details regarding the primary and secondary shared states are described with reference to Figures 40 and 42.

[0255] In one embodiment, at least two types of snooping, namely snoop-forward and snoop-back, can be supported. A snoop-forward message may be used to cause a coherent agent to forward a cache block to a requesting agent, and a snoop-back message may be used to cause a coherent agent to return a cache block to a memory controller. In one embodiment, a snoop-invalid message may also be supported (and may also include forward and back variants to specify the destination of completion). A snoop-invalid message causes a caching coherent agent to invalidate a cache block. Supporting snoop-forward and snoop-back flows can, for example, provide both cacheable (snoop-forward) and non-caching (snoop-back) behavior. Since a caching agent can store a cache block and potentially use the data within it, snoop-forward can be used to minimize the number of messages when a cache block is provided to the caching agent. On the other hand, a non-coherent agent may not store the entire cache block, and therefore copy-back to memory can ensure that the complete cache block is captured in the memory controller. Therefore, variations or types of snoop-forward and snoop-back messages may be selected based on the capabilities of the requesting agent (e.g., based on the requesting agent's identification information) and / or the type of request (e.g., cacheable or non-cacheable). Further details regarding snoop-forward and snoop-back messages are provided below with respect to Figures 37, 38, and 40. Various other features are shown in the remaining figures and described in more detail below.

[0256] Figure 27 is a block diagram of an embodiment of a system including a system-on-a-chip (SOC) C10 coupled to one or more memories such as memories C12A to C12m. SOC C10 may be, for example, an instance of SOC10 shown in Figure 1. SOC C10 may include a plurality of coherent agents (CAs) C14A to C14n. Coherent agents may include one or more processors (Ps) C16 coupled to one or more caches (e.g., cache C18). SOC C10 may include one or more non-coherent agents (NCAs) C20A to C20p. SOC C10 may include one or more memory controllers C22A to C22m, each coupled to a separate memory C12A to C12m during use. Each memory controller C22A to C22m may include a coherency controller circuit C24 (more simply, a "coherency controller" or "CC") coupled to directory C26. The memory controllers C22A-C22m, non-coherent agents C20A-C20p, and coherent agents C14A-C14n may be coupled to interconnect C28 for communication between the various components C22A-C22m, C20A-C20p, and C14A-C14n. As indicated by their names, the components of SOC C10 may, in one embodiment, be integrated on a single integrated circuit "chip". In other embodiments, the various components may be outside of SOC C10 on other chips or possibly individual components. Any number of integrated or individual components may be used. In one embodiment, a subset of coherent agents C14A-C14n and memory controllers C22A-C22m may be implemented on one of several integrated circuit chips coupled together to form the components shown in SOC C10 in Figure 27.

[0257] The coherency controller C24 can implement the memory controller portion of the cache coherency protocol. Generally, the coherency controller C24 may be configured to receive requests from the interconnect C28 (e.g., through one or more queues not shown in the memory controllers C22A-C22m) targeting cache blocks mapped to memories C12A-C12m to which memory controllers C22A-C22m are coupled. The directory may contain multiple entries, each entry being able to track the coherency state of individual cache blocks in the system. The coherency state may include, for example, the cache state of cache blocks in various coherent agents C14A-C14N (e.g., in cache C18, or in other caches such as the cache in processor C16 not shown). Therefore, based on the directory entry for the cache block corresponding to a given request and the type of a given request, the coherency controller C24 may be configured to determine which coherent agents C14A-C14n should receive the snoop and the type of snoop (e.g., snoop disable, snoop share, change to share, change to ownership, change to disable, etc.). The coherency controller C24 can also independently determine whether a snoop forward or snoop back is transmitted. The coherent agents C14A-C14n can receive the snoop, process the snoop, update the cache block state within the coherent agents C14A-C14n, and (if specified by the snoop) provide a copy of the cache block to the requesting coherent agent C14A-C14n or memory controller C22A-C22m that transmitted the snoop. Further details are provided below.

[0258] As described above, coherent agents C14A to C14n may include one or more processors C16. Processor C16 can function as the central processing unit (CPU) of the SOC C10. The system's CPU includes one or more processors that run the system's primary control software, such as the operating system. Generally, the software run by the CPU during use can control other components of the system to achieve desired functions of the system. The processor may also run other software, such as application programs. Application programs may provide user functions and may rely on the operating system for low-level device control, scheduling, memory management, etc. Therefore, the processor may also be referred to as an application processor. Coherent agents C14A to C14n may further include other hardware such as a cache C18 and / or interfaces to other components of the system (e.g., an interface to the interconnect C28). Other coherent agents may include processors that are not CPUs. Furthermore, other coherent agents may not include a processor (for example, fixed-function circuits such as a display controller or other peripheral circuits, or fixed-function circuits with processor support via one or more embedded processors may be coherent agents).

[0259] Generally, a processor may include any circuitry and / or microcode configured to execute instructions defined within an instruction set architecture implemented by the processor. The processor can include a processor core implemented on an integrated circuit having other components as a system-on-chip (SOC C10) or other level of integration. The processor can further include an individual microprocessor, a processor core and / or microprocessor integrated in a multi-chip module implementation, a processor implemented as multiple integrated circuits, and the like. The number of processors C16 within a given coherent agent C14A~C14n can be different from the number of processors C16 within another coherent agent C14A~C14n. Generally, one or more processors may be included. Further, the processors C16 may have different microarchitecture implementations, performance, and power characteristics, etc. In some cases, the processors may also differ in terms of the instruction set architecture they implement, their functions (e.g., CPU, graphics processing unit (GPU) processor, microcontroller, digital signal processor, image signal processor, etc.), and the like.

[0260] The cache C18 can have any capacity and configuration, such as set associative, direct mapped, or fully associative. The cache block size can be any desired size (e.g., 32 bytes, 64 bytes, 128 bytes, etc.). A cache block can be a unit of allocation and deallocation in the cache C18. Further, a cache block can be an address space in which coherence is maintained in this embodiment (e.g., a segment of the memory unit with aligned coherence fine grain size). A cache block may sometimes be referred to as a cache line.

[0261] In addition to the coherency controller C24 and directory C26, the memory controllers C22A-C22m may generally include circuitry for receiving memory operations from other components of the SOC C10 and accessing memory C12A-C12m to complete the memory operations. The memory controllers C22A-C22m may be configured to access any type of memory C12A-C12m. For example, memory C12A-C12m may be static random access memory (SRAM), double data rate (DRAM) such as synchronous DRAM (SDRAM) including dynamic RAM (DDR, DDR2, DDR3, DDR4, etc.), non-volatile memory, graphics DRAM such as graphics DDR DRAM (GDDR), and high-bandwidth memory (HBM). Low-power / mobile versions of DDR DRAM (e.g., LPDDR, mDDR, etc.) may be supported. The memory controllers C22A-C22m may include queues for memory operations to order (and possibly reorder) operations and present operations to memories C12A-C12m. The memory controllers C22A-C22m may further include data buffers for storing write data awaiting writing to memory and read data awaiting a reply to the source of a memory operation (if data is not provided by the snoop). In some embodiments, the memory controllers C22A-C22m may include a memory cache for storing recently accessed memory data. In SOC implementations, for example, the memory cache can reduce power consumption in the SOC by avoiding re-access of data from memories C12A-C12m when it is expected to be accessed again soon. In some cases, the memory cache may be referred to as a system cache, in contrast to private caches such as cache C18, or caches in processor C16 that serve only specific components. Furthermore, in some embodiments, the system cache does not need to be located within the memory controllers C22A-C22m.

[0262] Non-coherent agents C20A-C20p may generally include various additional hardware functions (e.g., “Peripherals”) included in the SOC C10. For example, Peripherals may include video peripherals such as image signal processors configured to process image acquisition data from cameras or other image sensors, GPUs, video encoders / decoders, scalers, rotators, blenders, etc. Peripherals may include audio peripherals such as microphones, speakers, interfaces to microphones and speakers, audio processors, digital signal processors, mixers, etc. Peripherals may include interface controllers for various interfaces outside the SOC C10, including interfaces such as Universal Serial Bus (USB), Peripheral Components Interconnect (PCI) including PCI Express (PCIe), serial and parallel ports. Peripherals may include networking peripherals such as Media Access Controllers (MACs). Any set of hardware may be included. In one embodiment, non-coherent agents C20A-C20p may also include a bridge to a set of peripherals.

[0263] Interconnect C28 can be any communication interconnect and protocol for communication between components of SOC C10. Interconnect C28 can be bus-based, including hierarchical buses with shared bus configurations, crossbar configurations, and bridges. Interconnect C28 may be packet-based or circuit-switched, and may be hierarchical with bridges, crossbars, point-to-point, or other interconnects. In one embodiment, Interconnect C28 may include multiple independent communication fabrics.

[0264] In general, the number of each component C22A-C22m, C20A-C20p, and C14A-C14n may vary from embodiment to embodiment, and any number may be used. As indicated by the postfixes "m", "p", and "n", the number of one type of component may differ from the number of another type of component. However, the number of a given type may be the same as the number of other types. Furthermore, although the system in Figure 27 is shown with multiple memory controllers C22A-C22m, embodiments having one memory controller C22A-C22m are also conceivable and can implement the cache coherency protocol described herein.

[0265] Referring to Figure 28, a block diagram is shown illustrating a plurality of coherent agents C12A-C12D and a memory controller C22A that perform a coherent transaction for a cacheable read-exclusive request (CRdEx) according to one embodiment of the Scalable Cache Coherency Protocol. A read-exclusive request may be a request for an exclusive copy of a cache block, and therefore, any other copy that coherent agents C14A-C14D invalidate, and the requester has only one valid copy when the transaction is complete. Memories C12A-C12m that have memory locations allocated to the cache block have data in the locations allocated to the cache block within memories C12A-C12m, but that data becomes "invalid" if the requester modifies that data. A read-exclusive request can be used, for example, to give the requester the ability to modify a cache block without transmitting an additional request in the Cache Coherency Protocol. If an exclusive copy is not required, other requests may be used (for example, if a writable copy is not necessarily required by the requester, a read-shared request CRdSh may be used). The "C" in the "CRdEx" label may stand for "cacheable". Other transactions may be issued by non-coherent agents (e.g., agents C20A-C20p in Figure 27), and such transactions may be labeled "NC" (e.g., NCRd). Further discussion of request types and other messages in transactions is provided below with respect to Figure 40 for one embodiment, and further discussion of cache states is provided below with respect to Figure 39 for one embodiment.

[0266] In the example in Figure 28, coherent agent C14A can initiate a transaction by transmitting a read-exclusive request to memory controller C22A (which controls the memory location assigned to the address in the read-exclusive request). Memory controller C22A (more specifically, coherence controller C24 within memory controller C22A) can read entries in directory C26 and determine that coherent agent C14D has a cache block in primary shared state (P) and is therefore a coherent agent that should provide the cache block to the requesting coherent agent C14D. Coherence controller C24 can generate a snoop-forward (SnpFwd[st]) message to coherent agent C14D and can issue a snoop-forward message to coherent agent C14D. Coherence controller C24 can include an identifier of the current state to the coherent agent that received the snoop, according to directory C26. For example, in this case, according to directory C26, the current state in coherent agent C14D is "P". Based on the snoop, coherent agent C14D can access the cache storing the cache block and generate a fill completion (Fill in Figure 28) with the data corresponding to the cache block. Coherent agent C14D can then transmit the fill completion to coherent agent C14A. Thus, the system implements a "three-hop" protocol for delivering data to the requester, namely CRdEx, SnpFwd[st], and Fill. As indicated by "[st]" in the SnpFwd[st] message, the snoop-forward message may also be coded with the state of the cache block to which the coherent agent should transition after processing the snoop. In various embodiments, different variations of the message may exist, or the state may be carried as a field in the message.In the example in Figure 28, because the request is a read-exclusive request, the new state of the cache block in the coherent agent may be invalid. Other requests may allow the new shared state.

[0267] Furthermore, the coherency controller C24 can determine from the directory entries of the cache blocks that coherent agents C14B-C14C have cache blocks that are in a secondary shared state (S). Therefore, it can issue snoops to each coherent agent that (i) has a cached copy of the cache block and (ii) whose block state changes based on transactions. Since coherent agent C14A has acquired an exclusive copy, the shared copy is invalidated, and therefore the coherency controller C24 can generate a snoop invalidation (SnpInvFw) message for coherent agents C14B-C14C and issue snoops to coherent agents C14B-C14C. The snoop invalidation message includes an identifier indicating that the current state of coherent agents C14B-C14C is shared. Coherent agents C14B-C14C can process the snoop invalidation request and provide coherent agent C14A with an acknowledgment (Ack) completion. Note that in the illustrated protocol, messages from the snooping agent to the coherency controller C24 are not implemented in this embodiment. The coherency controller C24 can update directory entries based on the issuance of snoops and process subsequent transactions. Therefore, as previously stated, transactions to the same cache block cannot be serialized in this embodiment. The coherency controller C24 can enable additional transactions to be initiated to the same cache block and can identify which snoop belongs to which transaction based on the current state indication within the snoop (for example, the next transaction to the same cache block will find the cache state corresponding to the previous completed transaction). In the illustrated embodiment, the snoop invalidation message is the SnpInvFw message because completion is sent to the initiating coherent agent C14A as part of the 3-hop protocol.In one embodiment, a 4-hop protocol is also supported for a specific agent. In such an embodiment, the SnpInvBk message can be used to indicate that the snooping agent has returned completion to the coherency controller C24.

[0268] Therefore, the cache state identifier in the snoop allows the coherent agent to resolve competition between messages forming different transactions to the same cache block. That is, messages may be received in an order different from the order in which the corresponding requests were processed by the coherency controller. The order in which coherency controller C24 processes requests to the same cache block via directory C26 can define the order of requests. That is, coherency controller C24 can be the ordering point for transactions received in a given memory controller C22A-C22m. On the other hand, message serialization can be managed within coherent agents C14A-C14n based on the current cache state corresponding to each message and the cache state within coherent agents C14A-C14n. A given coherent agent can access cache blocks within the coherent agent based on the snoop and can be configured to compare the cache state specified in the snoop with the cache state currently in the cache. If the states do not match, the snoop belongs to a transaction that is ordered after another transaction that changes the cache state in the agent to the state specified in the snoop. Therefore, the snooping agent may be configured to delay snooping based on the fact that the first state does not match the second state until the second state changes to the first state in response to a different communication relating to a different request than the first request. For example, the state may change based on fill completions received by the snooping agent from different transactions, etc.

[0269] In one embodiment, Snoop may include a completion count (Cnt) indicating the number of completions corresponding to a transaction, so that the requester can determine when all completions related to the transaction have been received. The coherency controller C24 can determine the completion count based on the state indicated in the directory entry of the cache block. The completion count may be, for example, the number of completions minus 1 (e.g., 2 in the example in Figure 28, since there are 3 completions). This implementation allows the completion count to be used as the initialization of the transaction's completion counter when the first completion of the transaction is received by the requesting agent (e.g., it has already been decremented to reflect the reception of completions that carry the completion count). Once the count is initialized, further completions of the transaction allow the requesting agent to update the completion counter (e.g., decrement the counter). In other embodiments, an actual completion count may be provided, or it may be decremented by the requester to initialize the completion count. In general, the completion count can be any value that identifies the number of completions that the requester should observe before the transaction is fully completed, that is, the requesting agent can complete the request based on the completion counter.

[0270] Figures 29 and 30 illustrate exemplary race conditions that can occur in transactions to the same cache block, and the use of the current cache state of a given agent (also referred to as the “expected cache state”) as reflected in the directory when the transaction is processed by the memory controller, and the current cache state of a given agent (e.g., reflected in the given agent’s cache(s) or buffers that can temporarily store cache data). In Figures 29 and 30, coherent agents are enumerated as CA0 and CA1, and memory controllers associated with a cache block are shown as MC. The vertical lines 30, 32, and 34 of CA0, CA1, and MC indicate the source (base of the arrow) and destination (tip of the arrow) of various messages corresponding to a transaction. In Figures 29 and 30, time progresses from top to bottom. A memory controller may be associated with a cache block if the memory to which the memory controller is associated contains a memory location allocated to the address of the cache block.

[0271] Figure 29 illustrates a competition between a fill completion for one transaction and snooping for different transactions on the same cache block. In the example in Figure 29, CA0 initiates a read-exclusive transaction with a CRdEx request to the MC (arrow 36). CA1 also initiates a read-exclusive transaction with a CRdEx request (arrow 38). The CA0 transaction is processed first by the MC, establishing an ordered CA0 transaction before the CA1 request. In this example, the directory indicates that there is no cached copy of the cache block in the system, and therefore the MC fills in exclusive state in response to the CA0 request (FillE, arrow 40). The MC updates the directory entry for the cache block in CA0's exclusive state.

[0272] The MC selects a CRdEx relationship from CA1 for processing and detects that CA0 has a cache block in exclusive state. Therefore, the MC can generate a snoop-forward request (SnpFwdI) to CA0 requesting that CA0 invalidate the cache block in its cache(s) and provide that cache block to CA1. The snoop-forward request also includes the E-state identifier of CA0's cache block, as it reflects the cache state in CA0's directory. The MC can issue a snoop (arrow 42) and update the directory to indicate that CA1 has an exclusive copy and CA0 no longer has a valid copy.

[0273] Snoop and fill completion may arrive at CA0 in either time order. Messages may travel within different virtual channels, and / or other delays within the interconnect may allow messages to arrive in either order. In the illustrated example, snoop arrives at CA0 before fill completion. However, CA0 may delay processing snoop because the expected state in snoop (E) does not match the current state of the cache block in CA0 (I). Fill completion can then arrive at CA0. CA0 may write the cache block to the cache and set its state to exclusive (E). CA0 may also be allowed to perform at least one action on the cache block to support the forward progression of tasks within CA0, the action may change its state to modify (M). In the cache coherence protocol, directory C26 does not have to track the M state separately (for example, it may be treated as E), but it may match the E state as the expected state in snoop. CA0 can issue a FillComplete to CA1 in a modified state (FillM, arrow 44). Therefore, the race condition between snooping and fill completion of the two transactions is handled correctly.

[0274] In the example in Figure 29, the CRdEx request is issued by CA1 following the CRdEx request from CA0, but the CRdEx request may also be issued by CA1 before the CRdEx request from CA0, and the CRdEx request from CA0 may still be ordered by the MC before the CRdEx request from CA1, since the MC is the transaction ordering point.

[0275] Figure 30 illustrates a race between a snoop for one coherent transaction and a completion for another coherent transaction for the same cache block. In Figure 30, CA0 initiates a write-back transaction (CWB) to write the modified cache block into memory (arrow 46), although the cache block may actually be tracked as exclusive in the directory as described above. The CWB may, for example, cause CA0 to remove the cache block from its cache, if the cache block is in a modified state. CA1 initiates a read-shared transaction (CRdS) for the same cache block (arrow 48). The CA1 transaction is ordered before the CA0 transaction by the MC, which reads the directory entry for the cache block and determines that CA0 has the cache block in an exclusive state. The MC issues a snoop-forward request to CA0, requesting a change to a secondary shared state (SnpFwdS, arrow 50). The identifier in the snoop indicates the current cache state of exclusive (E) in CA0. The MC updates the directory entry to indicate that CA0 has the cache block in secondary shared state and CA1 has the copy in primary shared state (since the previous exclusive copy was provided to CA1).

[0276] The MC processes the CWB request from CA0 and reads the directory entry for the cache block again. The MC issues an Ack completion, along with the cache state identifier in the Ack completion, indicating that the current cache state is secondary share (S) within CA0 (arrow 52). Based on the fact that the expected state of secondary share does not match the corrected current state, CA0 may delay processing the Ack completion. Processing the Ack completion allows CA0 to discard the cache block and not have a copy of the cache block to provide to CA1 in response to a later arriving SnpFwdS request. When the SnpFwdS request is received, CA0 can provide CA1 with a fill completion (arrow 54) and put the cache block into primary share (P). CA0 can also change the state of the cache block within CA0 to secondary share (S). The state change matches the expected state of the Ack completion, and therefore CA0 can invalidate the cache block and complete the CWB transaction.

[0277] Figure 31 is a block diagram showing in more detail one embodiment of a part of one embodiment of coherent agent C14A. Other coherent agents C14B to C14n may be similar. In the illustrated embodiment, coherent agent C14A may include a request control circuit C60 and a request buffer C62. The request buffer C62 is coupled to the request control circuit C60, and both the request buffer C62 and the request control circuit C60 are coupled to the cache C18 and / or processor C16, and the interconnect C28.

[0278] The request buffer C62 may be configured to store multiple requests generated by the cache C18 / processor C16 for a coherent cache block. That is, the request buffer C62 can store requests that initiate transactions on the interconnect C28. Figure 31 shows one entry in the request buffer C62, but other entries may be similar. An entry may include a valid (V) field C63, a request (Req.) field C64, a count valid (CV) field C66, and a completed count (CompCnt) field C68. The valid field C63 may store a valid indication (e.g., a valid bit) indicating whether the entry is valid (e.g., storing an unprocessed request). The request field C64 may store data that defines the request (e.g., request type, cache block address, tag or other identifier for the transaction). The count valid field C66 may store a valid indication for the completed count field C68, indicating that the completed count field C68 has been initialized. When processing a completion received from interconnect C28 for a request, the request control circuit C68 can use the count valid field C66 to determine whether the request control circuit C68 should initialize the field with the completion count included in the completion (count field is not valid) or update the completion count, such as by decrementing it (count field is valid). The completion count field C68 can store the current completion count.

[0279] The request control circuit C60 can receive requests from the cache 18 / processor 16 and allocate request buffer entries in the request buffer C62 to the requests. The request control circuit C60 can track the requests in buffer C62, transmit the requests over the interconnect C28 (for example, according to any arbitration scheme), track received completions in the requests, complete the transactions, and transfer the cache blocks to the cache C18 / processor C16.

[0280] Referring now to Figure 32, a flowchart illustrating the operation of one embodiment of a coherency controller C24 within memory controllers C22A-C22m based on receiving requests to be processed. The operation in Figure 32 may be performed when a request is selected from among the received requests for service in memory controllers C22A-C22m via an arbitrary desired arbitration algorithm. For ease of understanding, the blocks are shown in a specific order, but other orders may be used. The blocks may be performed in parallel in combinational logic within the coherency controller C24. The blocks, combinations of blocks, and / or the entire flowchart may be pipelined over multiple clock cycles. The coherency controller C24 may be configured to implement the operation shown in Figure 32.

[0281] The coherency controller C24 may be configured to read a directory entry from directory C26 based on the address of the request. The coherency controller C24 may be configured to determine which snoop is generated based on the type of request (e.g., the state requested for the cache block by the requester) and the current state of the cache block in the various coherent agents C14A-C14n indicated in the directory entry (block C70). The coherency controller C24 may also generate the current state to be contained in each snoop based on the current state of the coherent agents C14A-C14n that receive the snoops indicated in the directory. The coherency controller C24 may be configured to insert the current state into the snoop (block C72). The coherency controller C24 may also be configured to generate a completion count and insert the completion count into each snoop (block C74). As mentioned above, in one embodiment the completion count may be the number of completions minus 1, or the total number of completions. The number of completions can be the number of snoops, or a fill completion from memory controllers C22A-C22m if memory controllers C22A-C22m provide a cache block. In most cases where there is a snoop for a cacheable request, one of the snooped coherent agents C14A-C14n can provide the cache block, and therefore the number of completions can be the number of snoops. However, if coherent agents C14A-C14n do not have a copy of the cache block (no snoop), for example, a memory controller can provide a fill completion. Coherence controller C24 can be configured to queue snoops for transmission to coherent agents C14A-C14n (block C76). If the snoop queuing is successful, coherence controller C24 can be configured to update the directory entry to reflect the completion of the request (block C78).For example, an update could involve changing the cache state tracked within the directory entry to match the cache state requested by Snoop, or changing the agent identifier that indicates which agent should provide a copy of the cache block to coherent agents C14A-C14n, which would then place the cache block in an exclusive, modified, owned, or primary shared state upon transaction completion.

[0282] Referring now to Figure 33, a flowchart illustrating the operation of one embodiment of the request control circuit C60 in coherent agents C14A-C14n based on the reception of completion of outstanding requests in the request buffer C62 is shown. For ease of understanding, the blocks are shown in a specific order, but other orders may be used. The blocks can be implemented in parallel in the combinational logic within the request control circuit C60. The blocks, combinations of blocks, and / or the entire flowchart can be pipelined over multiple clock cycles. The request control circuit C60 may be configured to implement the operation shown in Figure 33.

[0283] The request control circuit C60 may be configured to access the request buffer entry in the request buffer C62 associated with the request to which the received completion is associated. If the count valid field C66 indicates that the completion count is valid (decision block C80, "yes" branch), the request control circuit C60 may be configured to decrement the count in the request count field C68 (block C82). If the count is zero (decision block C84, "yes" branch), the request is complete and the request control circuit C60 may be configured to forward the completion indication (and, if applicable, the received cache block) to the cache C18 and / or processor C16 that generated the request (block C86). Completion may update the state of the cache block. If the new state of the updated cache block matches the expected state in the pending snoop (decision block C88, "yes" branch), the request control circuit C60 may be configured to process the pending snoop (block C90). For example, the request control circuit C60 may be configured to pass a snoop to the cache C18 / processor C16 to generate a completion corresponding to the pending snoop (and to change the state of the cache block as indicated by the snoop).

[0284] The new state may coincide with the expected state if the new state is the same as the expected state. In addition, the new state may coincide with the expected state if the expected state is a state tracked by directory C26 for the new state. For example, in one embodiment, the modified state is tracked as an exclusive state in directory C26, and therefore the modified state coincides with the exclusive expected state. The new state can be modified, for example, if a cache block is exclusive and the state is provided in a fill completion transmitted by another coherent agent C14A~C14n that has modified the cache block locally.

[0285] If the count valid field C66 indicates that the completion count is valid (decision block C80), and the completion count is not zero after decrementing (decision block C84, branch to "no"), then the request is not completed and therefore remains pending in the request buffer C62 (and any pending snoops waiting for the request to complete may also remain pending). If the count valid field C66 indicates that the completion count is not valid (decision block C80, branch to "no"), then the request control circuit C60 may be configured to initialize the completion count field C68 with the completion count provided in completion (block C92). The request control circuit C60 may still be configured to check that the completion count is 0 (for example, if there is only one completion for the request, the completion count may be 0 in completion) (decision block C84), then processing may continue as described above.

[0286] Figure 34 is a flowchart illustrating the operation of one embodiment of coherent agents C14A-C14n based on Snoop reception. For ease of understanding, the blocks are shown in a specific order, but other orders may be used. The blocks can be implemented in parallel using combinational logic within coherent agents 14CA-C14n. Blocks, combinations of blocks, and / or the entire flowchart can be pipelined over multiple clock cycles. Coherent agents 14CA-C14n can be configured to implement the operation shown in Figure 34.

[0287] Coherent agents C14A to C14n can be configured to check the expected state in the snoop against the state in the cache C18 (decision block C100). If the expected state does not match the current state of the cache block (decision block C100, branch to "no"), the completion to change the current state of the cache block to the expected state remains pending. The completion corresponds to a transaction ordered before the transaction corresponding to the snoop. Therefore, coherent agents C14A to C14n can be configured to hold snoops, delaying the processing of the snoop until the current state changes to the expected state indicated in the snoop (block C102). In one embodiment, held snoops may be stored in a buffer dedicated to held snoops. Alternatively, held snoops may be absorbed into an entry in the request buffer C62 that stores conflicting requests, as will be described in more detail below with respect to Figure 36.

[0288] If the expected state matches the current state (decision block C100, branch to "yes"), the coherent agents C14A-C14n can be configured to process the state change based on the snoop (block C104). That is, the snoop can indicate the desired state change. The coherent agents C14A-C14n can be configured to generate a completion (e.g., a fill if the snoop is a snoop forward request, a copyback snoop response if the snoop is a snoop back request, or an affirmative response (forward or back depending on the snoop type) if the snoop is a state change request). The coherent agents can be configured to generate a completion with a completion count from the snoop (block C106) and to place the completion in a queue for transmission to the requesting coherent agents C14A-C14n (block CC108).

[0289] Using the cache coherency algorithm described herein, cache blocks can be transmitted from one coherent agent C14A-C14n to another coherent agent across a chain of competing requests with low message bandwidth overhead. For example, Figure 35 is a block diagram showing the transmission of a cache block between four coherent agents CA0-CA3. As in Figures 29 and 30, the coherent agents are listed as CA0-CA3, and the memory controller associated with the cache block is shown as MC. The vertical lines 110, 112, 114, 116, and 118 to CA0, CA1, CA2, CA3, and MC indicate the source (base of the arrow) and destination (tip of the arrow) of various messages corresponding to the transaction, respectively. Time progresses from top to bottom in Figure 35. At the point corresponding to the top of Figure 35, coherent agent CA3 has the cache block involved in the transaction in a modified state (tracked as exclusive within directory C26). All transactions in Figure 35 are for the same cache block.

[0290] Coherent agent CA0 initiates a read-exclusive transaction with a CRdEx request to the memory controller (arrow 120). Coherent agents CA1 and CA2 also initiate read-exclusive transactions (arrows 122 and 124, respectively). As indicated by the tips of arrows 120, 122, and 124 in line 118, the memory controller MC orders the transactions as CA0, then CA1, and finally CA2. In the exclusive state, the directory state of the transaction from CA0 is CA3, so the snoop forward and invalidation (SnpFwdI) are transmitted as exclusive in the current cache state (arrow 126). Coherent agent CA3 receives the snoop and forwards the FillM completion along with the data to coherent agent CA0 (arrow 128). Similarly, the directory state of a transaction from CA1 is coherent agent CA0, which is in an exclusive state (from the preceding transaction to CA0), and therefore the memory controller MC issues a SnpFwdI to coherent agent CA0 with a current cache state of E (arrow 130), and the directory state of a transaction from CA2 is coherent agent CA1 with a current cache state of E (arrow 132). When coherent agent CA0 has an opportunity to perform at least one memory operation on a cache block, coherent agent CA0 responds to coherent agent CA1 with a FillM complete (arrow 134). Similarly, when coherent agent CA1 has an opportunity to perform at least one memory operation on a cache block, coherent agent CA1 responds to its snoop by returning a FillM complete to coherent agent CA2 (arrow 136). The order and timing of various messages can change (for example, as shown in the competition conditions in Figures 29 and 30), but generally, cache blocks can move from agent to agent with one extra message (FillM complete) once the competing requests are resolved.

[0291] In one embodiment, due to the aforementioned competitive state, a snoop may be received before the fill completion it is supposed to snoop (detected by the snoop carrying the expected cache state). Furthermore, a snoop may be received before the Ack completion is collected and can process the fill completion. The Ack completion is due to the snoop and therefore depends on the progress of the virtual channel carrying the snoop. Thus, a competing snoop (a delayed wait in the expected cache state) may fill back pressure on the internal buffer and fabric, which can cause a deadlock. In one embodiment, coherent agents C14A~C14n may be configured to absorb one snoop forward and one snoop invalid into an pending request in the request buffer, rather than allocating separate entries. A non-competing snoop, or a competing snoop that has reached a point where it can be processed without further interconnect dependency, can then flow around the competing snoop and avoid a deadlock. When a snoop forward occurs, the responsibility for the forwarding is transferred to the target, so one snoop forward and one snoop invalidation may be sufficient. Therefore, no further snoop forwards will occur until the requesting party completes its current request and issues another new request after the previous snoop forward has completed. When a snoop invalidation occurs, the requesting party will not receive another invalidation until it has invalidated according to the directory, processed the previous invalidation, requested the cache block again, and obtained a new copy.

[0292] Therefore, coherent agents C14A to C14n may be configured to help ensure forward progress and / or prevent deadlocks by detecting the snoop received by the coherent agent and directing any pending requests ordered prior to the snoop into a cache block held by the coherent agent. The coherent agent may be configured to absorb a second snoop into a pending request (e.g., into a request buffer entry that stores the request). The coherent agent may process the absorbed snoop after completing the pending request. For example, if the absorbed snoop is a snoop-forward request, the coherent agent may, after completing the pending request, transfer the cache block to another coherent agent indicated in the snoop-forward snoop (and may also change the cache state to the state indicated by the snoop-forward request). If the absorbed snoop is a snoop-invalidation request, the coherent agent may update the cache state to invalid and transmit an acknowledgment completion after completing the pending request. Absorbing snoops into competing requests can be implemented, for example, by including additional storage for data describing the absorbed snoop in each request buffer entry.

[0293] Figure 36 is a flowchart illustrating the operation of one embodiment of coherent agents C14A-C14n based on snoop reception. For ease of understanding, the blocks are shown in a specific order, but other orders may be used. The blocks can be implemented in parallel using combinational logic within coherent agents C14A-C14n. Blocks, combinations of blocks, and / or the entire flowchart can be pipelined over multiple clock cycles. Coherent agents C14A-C14n can be configured to implement the operation shown in Figure 36. For example, the operation shown in Figure 36 may be part of the detection of a snoop with an expected cache state that does not match the expected cache state and is therefore pending (decision blocks C100 and C102 in Figure 34).

[0294] Coherent agents C14A-C14n may be configured to compare the address of a snoop that is held due to a lack of consistent cache state with the address of an unprocessed request (or pending request) in request buffer C62. If an address conflict is detected (decision block C140, "yes" branch), request buffer C62 may absorb the snoop into a buffer entry assigned to the pending request where the address conflict was detected (block C142). If there is no address conflict with a pending request (decision block C140, "no" branch), coherent agents C14A-C14n may be configured to allocate a separate buffer location for the snoop (e.g., in request buffer C62 or another buffer in coherent agents C14A-C14n) and may be configured to store data describing the snoop in the buffer entry (block C144).

[0295] As described above, in one embodiment, the cache coherency protocol can support both cacheable and non-cacheable requests while maintaining the coherence of the associated data. Non-cacheable requests may be issued, for example, by non-coherent agents C20A to C20p, which do not have the ability to coherently store cache blocks. In one embodiment, coherent agents C14A to C14n may also issue non-cacheable requests, which do not have to cache the data provided in response to such requests. Therefore, for example, if the data requested by a given non-coherent agent C20A~C20p is in a modified cache block in one of the coherent agents C14A~C14n, and the modified cache block is forwarded to the given non-coherent agent C20A~C20p expecting it to be saved by the given non-coherent agent C20A~C20p, then a snoop-forward request for a non-cacheable request is inappropriate.

[0296] To support coherent non-caching transactions, one embodiment of the scalable cache coherency protocol may include multiple types of snoops. For example, in one embodiment, snoops may include snoop-forward requests and snoop-back requests. As described above, a snoop-forward request can cause a cache block to be forwarded to the requesting agent. A snoop-back request, on the other hand, can cause a cache block to be returned to the memory controller. In one embodiment, a snoop-invalidation request may be supported to invalidate a cache block (along with forward and back versions to indicate completion).

[0297] More specifically, the memory controllers C22A-C22m (and, even more specifically, the coherency controller C24 within the memory controllers C22A-C22m) that receive the request may be configured to read entries from directory C26 corresponding to the cache block identified by the address in the request. The memory controllers C22A-C22m may also be configured to issue snoops to a given agent of coherent agents C14A-C14m that has a cached copy of the cache block according to the entries. The snoop indicates that the given agent should transmit the cache block to the source of the request based on the first request being of a first type (e.g., a cacheable request). The snoop indicates that the given agent should transmit the first cache block to the memory controller based on the first request being of a second type (e.g., a non-cacheable request). The memory controllers C22A-C22n may be configured to respond with completion to the source of the request based on having received the cache block from the given agent. Furthermore, as with other coherent requests, memory controllers C22A~C22n may be configured to update entries in directory C26 to reflect the completion of non-cacheable requests, based on issuing multiple snoops about the non-cacheable requests.

[0298] Figure 37 is a block diagram illustrating an example of a coherently managed non-cacheable transaction in one embodiment. Figure 37 may be an example of a four-hop protocol for passing snooped data to the requester via a memory controller. The non-coherent agent is listed as NCA0, the coherent agent as CA1, and the memory controller associated with the cache block as MC. The vertical lines 150, 152, and 154 of NCA0, CA1, and MC indicate the source (base of the arrow) and destination (tip of the arrow) of various messages corresponding to the transaction. Time progresses from top to bottom in Figure 37.

[0299] At the point corresponding to the upper part of Figure 37, coherent agent CA1 has the cache block in an exclusive state (E). NCA0 issues a non-cacheable read request (NCRd) to MC (arrow 156). MC determines from directory 26 that CA1 has the cache block containing the data requested by NCRd in an exclusive state and generates a snoopback request (SnpBkI(E)) to CA1 (arrow 158). CA1 provides MC with a copyback snoop response (CpBkSR) containing the cache block of data (arrow 160). If the data has been modified, MC can update memory with the data and provide NCA0 with the data for the non-cacheable read request in a non-cacheable read response (NCRdRsp) (arrow 162), completing the request. In one embodiment, there may be two or more types of NCRd requests, namely, requests to invalidate the cache block in the snooped coherent agent and requests to allow the snooped coherent agent to retain the cache block. The above explanation indicates invalidation. In other cases, the snooped agent can retain the cache block in the same state.

[0300] A non-cacheable write request may also be performed by using a snoopback request to retrieve a cache block and modifying the cache block with the non-cacheable write data before writing the cache block to memory. A non-cacheable write response can still be provided to notify the non-cacheable agent (NCA0 in Figure 37) that the write is complete.

[0301] Figure 38 is a flowchart illustrating the operation of one embodiment of the memory controllers C22A-C22m (more specifically, the coherency controller 24 within the memory controllers C22A-C22m in one embodiment) in response to a request, showing both cacheable and non-cacheable operations. The operation illustrated in Figure 38 may be, for example, a more detailed diagram of some of the operations shown in Figure 32. For ease of understanding, the blocks are shown in a specific order, but other orders may be used. The blocks can be implemented in parallel in the combinational logic within the coherency controller C24. The blocks, combinations of blocks, and / or the entire flowchart can be pipelined over multiple clock cycles. The coherency controller C24 may be configured to implement the operation shown in Figure 38.

[0302] The coherency controller C24 may be configured to read a directory based on an address in the request. If the request is a directory hit (decision block C170, branch to "yes"), the cached block resides in one or more caches among the coherent agents C14A-C14n. If the request is not cacheable (decision block C172, branch to "yes"), the coherency controller C24 may be configured to issue a snoop-back request to the coherent agents C14A-C14n responsible for providing a copy of the cached block (and, if applicable, snoop an invalidation request to the shared agent (back variant) - block C174). The coherency controller C24 may be configured to update the directory to reflect that the snooping is complete (e.g., invalidate the cached block in the coherent agents C14A-C14n - block C176). The coherency controller C24 can be configured to await copyback snoop responses (decision block C178, branch to "yes"), as well as any Ack snoop responses from shared coherent agents C14A-C14n, and can be configured to generate a non-cacheable completion for the requesting agent (NCRdRsp or NCWrRsp as needed) (block C180). The data may also be written to memory by memory controllers C22A-C22m if the cache blocks are modified.

[0303] If the request is cacheable (decision block C172, branch to "no"), the coherency controller C24 may be configured to generate snoop-forward requests (block C182) to coherent agents C14A-C14n responsible for forwarding the cached block, and, if necessary, other snoops to other caching coherent agents C14A-C14n. The coherency controller C24 may update directory C24 to reflect the completion of the transaction (block C184).

[0304] If the request is not a hit in directory C26 (decision block C170, branch to "no"), there is no cached copy of the cache block in coherent agents C14A-C14n. In this case, a snoop does not need to be generated, and memory controllers C22A-C22m may be configured to generate a fill complete (for cacheable requests) or a non-cacheable complete (for non-cacheable requests) to provide data or complete the request (block C186). For cacheable requests, coherence controller C24 can update directory C26 to create an entry for the cache block and initialize requesting coherent agents C14A-C14n to have a copy of the cached state cache block requested by coherent agents C14A-C14n (block C188).

[0305] Figure 39 is Table C190, which shows exemplary cache states that may be implemented in one embodiment of coherent agents C14A-C14n. Other embodiments may employ different cache states, subsets of the indicated cache states and other cache states, supersets of the indicated cache states and other cache states, and so on. The modified state (M) or “dirty exclusive” state may be a state in coherent agents C14A-C14n that has only a cached copy of a cache block (the copies are exclusive), and the data in the cached copy is modified with respect to the corresponding data in memory (for example, at least one byte of the data is different from the corresponding byte in memory). Modified data is sometimes referred to as dirty data. The owned state (O) or “dirty shared” state may be a state in coherent agents C14A-C14n that has a modified copy of a cache block, but may share the copy with at least one other coherent agent C14A-C14n (provided that the other coherent agent C14A-C14n has subsequently removed the shared cache block). Other coherent agents C14A-C14n place the cache block into a secondary shared state. The exclusive (E) state, or "clean exclusive" state, may be a state within coherent agents C14A-C14n where there is only a cached copy of the cache block, but the cached copy has the same data as the corresponding data in memory. The exclusive no data (EnD) state, or "clean exclusive, no data" state, may be a state within coherent agents C14A-C14n similar to the exclusive (E) state, except that the cache block of data has not been delivered to the coherent agents. Such a state may be used when coherent agents C14A-C14n should modify each byte in the cache block, and therefore there is no benefit or coherence reason to supply previous data in the cache block. The EnD state may also be an optimization to reduce traffic on interconnect C28 and may not be implemented in other embodiments.The primary shared (P) state, or "clean shared primary" state, may be a state within coherent agents C14A-C14n that has a shared copy of a cache block but is also responsible for forwarding the cache block to another coherent agent based on a snoop-forward request. The secondary shared (S) state, or "clean shared secondary" state, may be a state within coherent agents C14A-C14n that has a shared copy of a cache block but is not responsible for providing the cache block if another coherent agent C14A-C14n has a cache block in the primary shared state. In some embodiments, if coherent agents C14A-C14n do not have a cache block in the primary shared state, the coherency controller C24 may select a secondary shared agent to provide the cache block (and may also send a snoop-forward request to the selected coherent agent). In other embodiments, if there are no coherent agents C14A-C14n in a primary shared state, the coherency controller C24 can cause the memory controllers C22A-C22m to provide the cache blocks to the requester. The invalid state (I) may be a state within coherent agents C14A-C14n that does not have a cached copy of a cache block. Coherent agents C14A-C14n in the invalid state may not have previously requested a copy, or they may have any copy and have invalidated it based on snooping or on excluding the cache block to cache a different cache block.

[0306] Figure 40 is Table C192, showing various messages that may be used in one embodiment of the Scalable Cache Coherence Protocol. Other embodiments may include alternative messages, subsets of the illustrated messages and additional messages, supersets of the illustrated messages and additional messages, and so on. Messages can carry transaction identifiers that link messages from the same transaction (e.g., initial request, snoop, completion). Initial requests and snoops can carry addresses of cache blocks affected by the transaction. Some other messages can also carry addresses. In some embodiments, all messages can carry addresses.

[0307] A cacheable read transaction can be initiated with a cacheable read request message (CRd). Different versions of CRd requests may exist to request different cache states. For example, a CRdEx may request an exclusive state, and a CRdS may request a secondary shared state, and so on. The cache state actually provided in response to a cacheable read request may be at least as permissive as the requested state, and may be more permissive. For example, a CRdEx may receive a cached block in an exclusive or modified state. A CRdS may receive a block in a primary shared state, an exclusive state, an owned state, or an modified state. In one embodiment, a timely CRd request can be implemented to provide the most permissive state possible (that does not invalidate other copies of the cached block) (e.g., exclusive if no other coherent agent has a cached copy, owned or primary shared if other coherent agents have cached copies, etc.).

[0308] A Change to Exclusive (CtoE) message may be used by a coherent agent that has a copy of a cache block in a state that does not allow modification (e.g., owned, primary shared, secondary shared), and the coherent agent is attempting to modify the cache block (e.g., the coherent agent requires exclusive access to change the cache block to modified). In one embodiment, a conditional CtoE message may be used for a store conditional instruction. A store conditional instruction is part of a load reservation / store conditional pair in which a load obtains a copy of a cache block and sets a reservation for that cache block. Coherent agents C14A-C14n can monitor other agents' access to the cache block and can conditionally perform a store based on whether the cache block has not been modified by another coherent agent C14A-C14n between the load and the store (successfully store if the cache block has not been modified, and not store if the cache block has been modified). Further details are provided below.

[0309] In one embodiment, when coherent agents C14A~C14n modify an entire cache block, a cache read exclusive data-only (CRdE-Donly) message can be used. If the cache block has not been modified by another coherent agent C14A~C14n, the requesting coherent agents C14A~C14n can use the EnD cache state to modify all bytes of the block without transferring the previous data in the cache block to the agent. Once the cache block is modified, the modified cache block can be transferred to the requesting coherent agents C14A~C14n, which can use the M cache state.

[0310] Non-cacheable transactions can be initiated using non-cacheable read and non-cacheable write (NCRd and NCWr) messages.

[0311] Snoop-forward and snoop-back (SnpFwd and SnpBk, respectively) may be used for snooping as previously described. There may be messages requesting various states (e.g., invalid or shared) within the receiving coherent agents C14A-C14n after processing the snoop. There may also be snoop-forward messages for CRdE-Donly requests, which request a transfer if the cache block has been modified, but not otherwise, and are invalidated at the receiver. In one embodiment, there may also be invalidation-only snoop-forward and snoop-back requests, shown as SnpInvFw and SnpInvBk in Table C192 (e.g., snoops that cause the receiver to invalidate without returning data, respectively, and prompt the requester or memory controller to acknowledge).

[0312] Completion messages may include fill messages and acknowledgment messages. A fill message can specify the state of the cache block that the requester should take upon completion. Cacheable write-back (CWB) messages may be used to transmit cache blocks to memory controllers C22A-C22m (for example, based on removing the cache block from the cache). Copy-back snoop responses (CpBkSR) may be used to transmit cache blocks to memory controllers C22A-C22m (for example, based on a snoop-back message). Non-cacheable write completion (NCWrRsp) and non-cacheable read completion (NCRdRsp) may be used to complete non-cacheable requests.

[0313] Figure 41 is a flowchart illustrating the operation of one embodiment of the coherency controller C24 based on the reception of a conditional exclusive modification (CtoECond) message. For example, Figure 41 may be a more detailed description of a portion of block C70 in Figure 32 in one embodiment. The blocks are shown in a specific order for ease of understanding, but other orders may be used. The blocks can be implemented in parallel in the combinational logic within the coherency controller C24. The blocks, combinations of blocks, and / or the entire flowchart can be pipelined over multiple clock cycles. The coherency controller C24 may be configured to implement the operation shown in Figure 41.

[0314] The CtoECond message may be issued by coherent agents C14A~14n ("source") based on the execution of a conditional store instruction. If the source has lost a copy of the cached block before the store condition instruction (e.g., the copy is no longer valid), the store condition instruction may fail locally at the source. If the source still has a valid copy (e.g., secondary or primary shared state, or owned state), when the conditional store instruction is executed, there is still a possibility that another transaction may be ordered before an exclusive modification message from the source that invalidates the cached copy at the source. The same transaction that invalidates the cached copy will also cause the store conditional instruction to fail at the source. To avoid the invalidation of the cached block and the transfer of the cached block to the source that causes the store conditional instruction to fail, the CtoECond message is provided and may be used by the source.

[0315] A CtoECond message can be defined to have at least two possible outcomes when ordered by the coherency controller C24. When the CtoECond message has been ordered and processed, if the source still has a valid copy of the cache block, as shown in directory C26, the CtoECond can, like an unconditional CtoE message, issue a snoop and proceed to obtain exclusive state of the cache block. If the source does not have a valid copy of the cache block, the coherency controller C24 can return an Ack completion to the source with an indication that the CtoE transaction failed. The source can then terminate the CtoE transaction based on the Ack completion.

[0316] As shown in Figure 41, the coherency controller C24 may be configured to read the directory entry for the address (block C194). If the source holds a valid copy of the cache block (e.g., in a shared state) (decision block C196, branch to "yes"), the coherency controller C24 may be configured to generate a snoop (e.g., a snoop to invalidate the cache block so that the source can change it to an exclusive state) based on the cache state in the directory entry (block C198). If the source does not hold a valid copy of the cache block (decision block C196, branch to "no"), the cache controller C24 may be configured to transmit an acknowledgment completion to the source indicating a failure of the CtoECond message (block C200). Thus, the CtoE transaction may be terminated.

[0317] Referring now to Figure 42, a flowchart illustrating the operation of one embodiment of the coherency controller C24 (for example, in one embodiment, at least a portion of block C70 in Figure 32) that reads a directory entry and determines a snoop. For ease of understanding, the blocks are shown in a specific order, but other orders may be used. The blocks can be implemented in parallel in the combinational logic within the coherency controller C24. The blocks, combinations of blocks, and / or the entire flowchart can be pipelined over multiple clock cycles. The coherency controller C24 may be configured to implement the operation shown in Figure 42.

[0318] As shown in Figure 42, the coherency controller C24 may be configured to read a directory entry for the address of the request (block C202). Based on the cache state in the directory entry, the coherency controller C24 may be configured to generate a snoop. For example, based on the cache state in at least one of the primary sharing agents (decision block C204, branch to "yes"), the coherency controller C24 may be configured to transmit a SnpFwd snoop to the primary sharing agent indicating that the primary sharing agent should transmit the cached block to the requesting agent. For other agents (e.g., in a secondary sharing state), the coherency controller C24 may be configured to generate an invalidation-only snoop (SnpInv) indicating that the other agent will not transmit the cached block to the requesting agent (block C206). In some cases (e.g., a CRdS request for a shared copy of a cached block), the other agents do not need to receive a snoop because they do not need to change their state. An agent can have a cache state that is at least primary shared if that cache state is at least as permissive as primary shared (for example, primary shared, owned, exclusive, or modified in the embodiment of Figure 39).

[0319] If no agents have a cache state that is at least primary shared (decision block C204, branch to "no"), the coherency controller C24 can be configured to determine whether one or more agents have a cache block in a secondary shared state (decision block C208). If so (decision block C208, branch to "yes"), the coherency controller C24 can be configured to select one of the agents that has a secondary shared state and can transmit a SnpFwd request instruction to the requesting agent for the selected agent to transfer to the cache block. The coherency controller C24 may be configured to generate a SnpInv request to the other agents in a secondary shared state, indicating that the other agents will not transfer the cache block to the requesting agent (block C210). If, as described above, the other agents do not need to change their state, the SnpInv message may not be generated and transmitted.

[0320] If there are no agents with secondarily shared cache state (decision block C208, branch to "no"), the coherency controller C24 may be configured to generate fill complete and cause the memory controller to read the cache block for transmission to the requesting agent (block C212).

[0321] Figure 43 is a flowchart illustrating the operation of one embodiment of the coherency controller C24 (for example, in one embodiment, at least a portion of block C70 in Figure 32) for reading directory entries and determining a snoop in response to a CRdE-Donly request. The blocks are shown in a specific order for ease of understanding, but other orders may be used. The blocks can be implemented in parallel in the combinational logic within the coherency controller C24. Blocks, combinations of blocks, and / or the entire flowchart can be pipelined over multiple clock cycles. The coherency controller C24 may be configured to implement the operation shown in Figure 43.

[0322] As described above, a CRdE-Donly request can be used by coherent agents C14A~C14n that modify all bytes in a cache block. Therefore, the coherency controller C24 can cause other agents to invalidate the cache block. If an agent has a modified cache block, it can supply the modified cache block to the requesting agent. Otherwise, the agent does not have to supply the cache block.

[0323] The coherency controller C24 can be configured to read a directory entry for the address of the request (block C220). Based on the cache state in the directory entry, the coherency controller C24 can be configured to generate a snoop. More specifically, if a given agent may have a modified copy of a cache block (e.g., a given agent has a cache block in exclusive or primary state) (block C222, branch "yes"), the cache controller C24 can generate a snoop-forward dirty-only (SnpFwdDonly) to the agent in order to transmit the cache block to the requesting agent (block C224). As described above, a SnpFwdDonly request can cause the receiving agent to transmit the cache block if the data has been modified, and not otherwise. In either case, the receiving agent can invalidate the cache block. The receiving agent can transmit a Fill complete and provide the modified cache block if the data has been modified. Otherwise, the receiving agent can transmit an Ack complete. If no agent has a modified copy (decision block C222, branch to "no"), the coherency controller C24 may be configured to generate a snoop invalidation (SnpInv) for each agent that has a cached copy of the cache block (block C226). In another embodiment, the coherency controller C24 may not request data transfer even if the cache block is modified, because the requester modifies the entire cache block. That is, the coherency controller C24 can cause agents with a modified copy to invalidate the data without transferring the data.

[0324] Based on this disclosure, a system may include a plurality of coherent agents, a given agent among the plurality of coherent agents including one or more caches for caching memory data. The system may further include a memory controller coupled to one or more memory devices, the memory controller including a directory configured to track which of the plurality of coherent agents is caching copies of a plurality of cache blocks in the memory device, and the state of the cached copies in the plurality of coherent agents. Based on a first request by a first agent among the plurality of coherent agents for a first cache block, the memory controller may be configured to read an entry from the directory corresponding to the first cache block, issue a snoop to a second agent among the plurality of coherent agents that has a cached copy of the first cache block according to the entry, and include in the snoop an identifier of the first state of the first cache block in the second agent. Based on the snoop, the second agent may be configured to compare the first state with the second state of the first cache block in the second agent and to delay processing the snoop on the basis that the first state does not match the second state until the second state changes to the first state in response to a different communication related to a different request than the first request. In one embodiment, the memory controller may be configured to determine a completion count indicating the number of completions that the first agent will receive for a first request, the determination being based on the state from the entry, and to include the completion count in a plurality of snoops issued based on the first request, including the snoop issued to the second agent. The first agent may be configured to initialize a completion counter with the completion count based on receiving an initial completion from one of the plurality of coherent agents, update the completion counter based on receiving a subsequent completion from another of the plurality of coherent agents, and complete the first request based on the completion counter.In one embodiment, the memory controller may be configured to update the state in the directory entries to reflect the completion of a first request, based on issuing a number of snoops based on the first request. In one embodiment, the first agent may be configured to detect a second snoop received by the first agent for a first cache block, and the first agent may be configured to absorb the second snoop into the first request. In one embodiment, the first agent may be configured to process the second snoop after completing the first request. In one embodiment, the first agent may be configured to transfer the first cache block to a third agent indicated in the second snoop after completing the first request. In one embodiment, the third agent may be configured to generate a conditional change request to an exclusive state based on a conditional store instruction to a second cache block that is in an active state in the third agent. The memory controller may be configured to determine whether a third agent holds a valid copy of the second cache block based on a second entry in a directory associated with the second cache block, and the memory controller may be configured to transmit a completion indicating failure to the third agent and terminate the conditional change to an exclusive request based on the determination that the third agent no longer holds a valid copy of the second cache block. In one embodiment, the memory controller may be configured to issue one or more snoops to other coherent agents among a plurality of coherent agents indicated by the second entry, based on the determination that the third agent holds a valid copy of the second cache block. In one embodiment, the snoop indicates that the second agent should transmit the first cache block to the first agent based on the first state being primary shared, and the snoop indicates that the second agent should not transmit the first cache block based on the first state being secondary shared.In one embodiment, Snoop indicates that the second agent transmits the first cache block even if the first state is secondary sharing.

[0325] In another embodiment, the system includes a plurality of coherent agents, a given agent among the plurality of coherent agents including one or more caches for caching memory data. The system further includes a memory controller coupled to one or more memory devices. The memory controller may include a directory configured to track which of the plurality of coherent agents caches copies of a plurality of cache blocks in the memory device, and the state of the cached copies among the plurality of coherent agents. Based on a first request by a first agent among the plurality of coherent agents for a first cache block, the memory controller may be configured to read an entry from the directory corresponding to the first cache block and issue a snoop to a second agent among the plurality of coherent agents that has a cached copy of the first cache block according to the entry. The snoop may indicate that the second agent should transmit the first cache block to the first agent, based on an entry indicating that the second agent has the first cache block in at least primary shared state. A snoop indicates that a second agent will not transmit the first cache block to the first agent based on the fact that a different agent has the first cache block in at least primary shared state. In one embodiment, if a different agent is in primary shared state, the first agent is in secondary shared state with respect to the first cache block. In one embodiment, a snoop indicates that the first agent will invalidate the first cache block based on the fact that a different agent has the first cache block in at least primary shared state. In one embodiment, the memory controller is configured not to issue a snoop to the second agent based on the fact that a different agent has the first cache block in primary shared state and the first request is a request for a shared copy of the first cache block.In one embodiment, the first request may be to obtain the exclusive state of a first cache block, and the first agent modifies the entire first cache block. Snoop may indicate that if the first cache block is in a modified state at the second agent, the second agent should transmit the first cache block. In one embodiment, Snoop indicates that if the first cache block is not in a modified state at the second agent, the second agent should invalidate the first cache block.

[0326] In another embodiment, the system includes a plurality of coherent agents, a given agent among the plurality of coherent agents including one or more caches for caching memory data. The system further includes a memory controller coupled to one or more memory devices. The memory controller may include a directory configured to track which of the plurality of coherent agents caches copies of a plurality of cache blocks in the memory device, and the state of the cached copies among the plurality of coherent agents. Based on a first request for a first cache block, the memory controller may be configured to read an entry from the directory corresponding to the first cache block and issue a snoop to a second agent among the plurality of coherent agents that has a cached copy of the first cache block according to the entry. The snoop may indicate that the second agent should transmit the first cache block to the source of the first request based on an attribute associated with the first request having a first value, or the snoop may indicate that the second agent should transmit the first cache block to the memory controller based on an attribute having a second value. In one embodiment, the attribute is the type of request, the first value is cacheable, and the second value is not cacheable. In another embodiment, the attribute is the source of the first request. In one embodiment, the memory controller may be configured to respond to the source of the first request based on having received the first cache block from a second agent. In one embodiment, the memory controller is configured to update the state in the directory entry to reflect the completion of the first request based on issuing a number of snoops based on the first request. IOA

[0327] Figures 44–48 illustrate various embodiments of input / output agents (IOAs) that may be employed in various embodiments of a SOC. An IOA can be inserted between a given peripheral device and an interconnect fabric. An IOA agent can be configured to enforce the interconnect fabric's coherency protocol with respect to a given peripheral device. In one embodiment, the IOA uses the coherency protocol to ensure the ordering of requests from a given peripheral device. In one embodiment, the IOA is configured to connect a network of two or more peripheral devices to the interconnect fabric.

[0328] In many cases, computer systems implement data / cache coherency protocols that guarantee a coherent view of data within the computer system. As a result, changes to shared data are typically propagated throughout the computer system in a timely manner to ensure a consistent view. Computer systems also typically include or interface with peripherals such as input / output (I / O) devices. However, these peripherals are not configured to understand or efficiently use the cache coherency protocols implemented by the computer system. For example, peripherals often use specific ordering rules (described further below) that are stricter than the cache coherency protocol for their transactions. Many peripherals also do not have a cache; i.e., they are not cacheable devices. As a result, it can take a considerable amount of time for a peripheral to receive a completion acknowledgment for their transactions because they have not completed in their local cache. This disclosure addresses, among other things, these technical problems concerning peripherals that do not have a cache and cannot properly use the cache coherency protocol.

[0329] This disclosure describes various techniques for implementing an I / O agent configured to bridge peripherals to a coherent fabric and implement a coherency mechanism for processing transactions associated with those I / O devices. In the various embodiments described below, a system-on-a-chip (SOC) includes memory, a memory controller, and an I / O agent coupled to the peripherals. The I / O agent is configured to receive read and write transaction requests from the peripherals targeting specified memory addresses where data may be stored in the SOC's cache lines. (Cache lines may also be referred to as cache blocks.) In the various embodiments, specific ordering rules of the peripherals impose that read / write transactions are completed sequentially (e.g., not in any order with respect to the order in which they are received). As a result, in one embodiment, the I / O agent is configured to complete read / write transactions before initiating the next read / write transaction that occurs, according to their execution order. However, in order to execute those transactions in a more efficient manner, in various embodiments, the I / O agent is configured to acquire exclusive ownership of the targeted cache lines so that the data of those cache lines is not cached in a valid state by other caching agents of the SOC (e.g., processor cores). Instead of waiting for the first transaction to complete before starting work on the second transaction, the I / O agent can preemptively acquire exclusive ownership of the cache line(s) targeted by the second transaction. As part of acquiring exclusive ownership, in various embodiments, the I / O agent receives the data of those cache lines and stores that data in the I / O agent's local cache.Once the first transaction is complete, the I / O agent can then send a request for data from those cache lines and complete a second transaction within its local cache without having to wait for the data to be returned. As will be explained in more detail below, the I / O agent can acquire exclusive read or exclusive write ownership depending on the type of transaction involved.

[0330] In some cases, an I / O agent may lose exclusive ownership of a cache line before executing the corresponding transaction. For example, an I / O agent may receive a snoop that causes it to relinquish exclusive ownership of the cache line, including invalidating the data stored in the I / O agent for the cache line. As used herein, “snoop” or “snoop request” refers to a message transmitted to a component to request a change in the state of a cache line (e.g., invalidating the data of the cache line stored in the component’s cache), and if the component has an exclusive copy of the cache line or otherwise assumes the cache line, the message may also request that the cache line be provided by the component. In various embodiments, if there are a threshold number of outstanding transactions directed to the cache line, the I / O agent may reacquire exclusive ownership of the cache line. For example, if there are three outstanding write transactions targeting a cache line, the I / O agent may reacquire exclusive ownership of that cache line. This can prevent unduly slow serialization of the remaining transactions targeting a particular cache line. In various embodiments, a larger or smaller number of unprocessed transactions may be used as the threshold.

[0331] These techniques may be advantageous over conventional methods because they allow the preservation of peripheral device ordering rules while partially or completely negating their negative effects by implementing coherency mechanisms. In particular, paradigms that execute transactions in a specific order according to ordering rules, where transactions are completed before work on subsequent transactions begins, can be unduly slow. For example, loading data from a cache line into the cache can take more than 500 clock cycles. Therefore, if subsequent transactions are not started until the previous transaction is completed, each transaction will take at least 500 clock cycles to complete, resulting in a high number of clock cycles being used to process a set of transactions. As disclosed in this disclosure, a high number of clock cycles for each transaction can be avoided by preemptively acquiring exclusive ownership of the relevant cache lines. For example, when an I / O agent is processing a set of transactions, the I / O agent may preemptively start caching data before the first transaction is completed. As a result, when the first transaction is completed, the data for the second transaction can be cached and made available, allowing the I / O agent to complete the second transaction immediately thereafter. Therefore, some parts of the transactions may not require, for example, more than 500 clock cycles to complete. Next, illustrative applications of these techniques will be described with reference to Figure 44.

[0332] Referring now to Figure 44, a block diagram of an exemplary system-on-a-chip (SOC) D100 is shown. In one embodiment, SOC D100 may be an embodiment of SOC 10 shown in Figure 1. As the name implies, the components of SOC D100 are integrated on a single semiconductor substrate as an integrated circuit “chip”. However, in some embodiments, the components are mounted on two or more separate chips in a computing system. In the illustrated embodiment, SOC D100 includes a caching agent D110, memory controllers D120A and D120B coupled to memories DD130A and 130B respectively, and an input / output (I / O) cluster D140. Components D110, D120, and D140 are coupled to each other via an interconnect D105. Furthermore, as shown in the figure, the caching agent D110 includes the processor D112 and the cache D114, and the I / O cluster D140 includes the I / O agent D142 and peripherals D144. In various embodiments, the SOC D100 may be implemented differently from that shown. For example, the SOC D100 may include a display controller, power management circuits, etc., and the memories D130A and D130B may be included on the SOC D100. As another example, the I / O cluster D140 may have multiple peripherals D144, one or more of which may be outside the SOC D100. Therefore, it should be noted that the number of components (and the number of sub-components) of the SOC D100 may vary between embodiments. The number of components / sub-components may be greater or less than the number shown in Figure 44.

[0333] In various embodiments, the caching agent D110 is any circuit that includes a cache for caching memory data, or optionally controls cache lines, and optionally updates the data on those cache lines locally. The caching agent D110 can participate in a cache coherency protocol to ensure that updates to data made by one caching agent D110 are visible to other caching agents D110 that subsequently read that data, and that updates made by two or more caching agents D110 in a specific order (determined at ordering points within the SOC D100, such as memory controllers D120A-B) are observed in that order by the caching agents D110. The caching agent D110 may include, for example, a processing unit (e.g., CPU, GPU, etc.), fixed-function circuits, and fixed-function circuits with processor assistance via (one or more) embedded processors. Since the I / O agent D142 includes a set of caches, the I / O agent D142 can be considered a type of caching agent D110. However, I / O agent D142 differs from other caching agents D110, at least in that it functions as a cacheable entity configured to cache data for other separate entities that do not have their own cache (e.g., peripherals such as displays and USB-connected devices). In addition, I / O agent D142 can temporarily cache a relatively small number of cache lines to improve peripheral memory access latency, but can proactively retire cache lines once a transaction is complete.

[0334] In the illustrated embodiment, the caching agent D110 is a processing unit having a processor D112 that can function as the CPU of the SOC D100. In various embodiments, the processor D112 includes arbitrary circuitry and / or microcode configured to execute instructions defined in the instruction set architecture implemented by the processor D112. The processor D112 may comprise one or more processor cores implemented on an integrated circuit together with other components of the SOC D100. Those individual processor cores of the processor D112 may share a common last-level cache (e.g., L2 cache) while each including its own cache (e.g., L0 cache and / or L1 cache) for storing data and program instructions. The processor D112 can execute the main control software of the system, such as an operating system. Generally, the software executed by the CPU controls other components of the system to achieve desired functions of the system. The processor D112 may further execute other software, such as application programs, and is therefore sometimes referred to as an application processor. The caching agent D110 may further include hardware configured to interface the caching agent D110 to other components of the SOC D100 (for example, an interface to the interconnect D105).

[0335] Cache D114 is, in various embodiments, a storage array containing entries configured to store data or program instructions. Thus, cache D114 can be a data cache, an instruction cache, or a shared instruction / data cache. Cache D114 can be an associative storage array (e.g., fully associative or set-associative, such as 4-way content) or a directly-mapped storage array, and can have any storage set-associative scheme. In various embodiments, a cache line (or alternatively, a "cache block") is a unit of allocation and deallocation within cache D114, and may be of any desired size (e.g., 32 bytes, 64 bytes, 128 bytes, etc.). During the operation of the caching agent DD110, information can be pulled from other components of the system into cache D114 and used by the processor core of processor D112. For example, as the processor core progresses along the execution path, it can fetch program instructions from memory D130A-B to cache D114, then fetch them from cache D114, and execute them. Also, during the operation of the caching agent D110, information can be written from cache D114 to memory (e.g., memory D130A-B) via the memory controllers D120A-B.

[0336] In various embodiments, the memory controller D120 includes circuitry configured to receive memory requests (e.g., load / storage requests, instruction fetch requests, etc.) from other components of the SOC D100 to perform memory operations, such as accessing data from memory D130. The memory controller D120 may be configured to access any type of memory D130. Memory D130 may be implemented using a variety of different physical memory media, such as hard disk storage, floppy disk storage, removable disk storage, flash memory, random access memory (RAM-SRAM, EDO RAM, SDRAM, DDR SDRAM, RAMBUS RAM, etc.), and read-only memory (PROM, EEPROM, etc.). However, the memory available to the SOC D100 is not limited to primary storage such as memory D130. Rather, the SOC D100 may further include other forms of storage, such as cache memory (e.g., L1 cache, L2 cache, etc.) within the caching agent D110. In some embodiments, the memory controller D120 includes a queue for storing and ordering memory operations to be presented to memory D130. The memory controller D120 may also include a data buffer for storing write data waiting to be written to memory D130 and read data waiting to be returned to a source of memory operations, such as a caching agent D110.

[0337] As will be explained in more detail with respect to Figure 45, the memory controller D120 may include various components for maintaining cache coherence within the SOC D100, including components for tracking the location of data in cache lines within the SOC D100. Thus, in various embodiments, requests for cache line data are routed through the memory controller D120, which can access data from other caching agents D110 and / or memories D130A-B. In addition to accessing data, the memory controller D120 can cause caching agents D110 and I / O agents D142, which store data in their local caches, to issue snoop requests. As a result, the memory controller D120 can ensure coherence within the system by causing those caching agents D110 and I / O agents D142 to invalidate and / or remove data from their caches. Therefore, in various embodiments, the memory controller D120 processes exclusive cache line ownership requests and grants exclusive ownership of the cache line components while using snoop requests to ensure that the data is not cached by other caching agents D110 and I / O agent D142.

[0338] The I / O cluster D140 includes, in various embodiments, one or more peripheral devices D144 (or simply peripherals D144) that may provide additional hardware functionality and I / O agents D142. Peripherals D144 may include, for example, video peripherals (e.g., GPU, blender, video encoder / decoder, scaler, display controller, etc.) and audio peripherals (e.g., microphone, speaker, interface to microphone and speaker, digital signal processor, audio processor, mixer, etc.). Peripherals D144 may include interface controllers for various external interfaces of the SOC D100 (e.g., Universal Serial Bus (USB), Peripheral Components Interconnect (PCI) and PCI Express (PCIe), serial and parallel ports, etc.). Interconnects to external components are indicated in Figure 44 by dashed arrows extending outside the SOC D100. Peripherals D144 may also include networking peripherals such as a Media Access Controller (MAC). Although not shown, in various embodiments, the SOC D100 includes a plurality of I / O clusters D140, each having a set of peripheral devices D144. For example, the SOC D100 may include a first I / O cluster D140 having an external display peripheral device D144, a second I / O cluster D140 having a USB peripheral device D144, and a third I / O cluster D140 having a video encoder peripheral device D144. Each of these I / O clusters D140 may include its own I / O agent D142.

[0339] In various embodiments, the I / O agent D142 includes circuitry configured to bridge its peripherals D144 to the interconnect D105 and implement a coherency mechanism for processing transactions associated with those peripherals D144. As will be described in more detail with respect to Figure 45, the I / O agent D142 can receive transaction requests from peripherals D144 to read and / or write data to cache lines associated with memories D130A-B. In response to those requests, in various embodiments, the I / O agent D142 communicates with the memory controller D120 to acquire exclusive ownership of the target cache lines. Thus, the memory controller D120 can grant exclusive ownership to the I / O agent D142, which may involve providing cache line data to the I / O agent D142 and sending snoop requests to other caching agents D110 and I / O agent D142. After acquiring exclusive ownership of a cache line, I / O agent D142 can begin completing transactions targeting the cache line. In response to transaction completion, I / O agent D142 can send an acknowledgment to requesting peripheral D144 that the transaction has been completed. In some embodiments, I / O agent D142 does not acquire exclusive ownership of relaxed ordered requests that do not need to be completed in a specified order.

[0340] Interconnect D105 is, in various embodiments, an arbitrary communication-based interconnect and / or protocol for communication between components of the SOC D100. For example, interconnect D105 can enable a processor D112 in a caching agent D110 to interact with a peripheral D144 in an I / O cluster D140. In various embodiments, interconnect D105 is bus-based, including a shared bus configuration, a crossbar configuration, and a hierarchical bus with bridges. Interconnect D105 can be packet-based and may be hierarchical with bridges, crossbars, point-to-point, or other interconnects.

[0341] Referring next to Figure 45, a block diagram of exemplary interacting elements is shown, including a caching agent D110, a memory controller D120, an I / O agent D142, and peripherals D144. In the illustrated embodiment, the memory controller D120 includes a coherency controller D210 and a directory D220. In some cases, the illustrated embodiment may be implemented differently from that shown. For example, there may be multiple caching agents D110, multiple memory controllers D120, and / or multiple I / O agents D142.

[0342] As stated, the memory controller D120 can maintain cache coherency within the SOC D100, including tracking the location of cache lines within the SOC D100. Thus, the coherency controller D210 is configured in various embodiments to implement the memory controller portion of the cache coherency protocol. The cache coherency protocol can specify messages or commands that can be transmitted between the caching agent D110, the I / O agent D142, and the memory controller D120 (or coherency controller D210) to complete a coherent transaction. These messages may include a transaction request D205, a snoop D225, and a snoop response D227 (or alternatively, "complete"). The transaction request D205 is, in various embodiments, a message that initiates a transaction, specifying the requested cache line / block (e.g., having the address of that cache line) and the state in which the requester will receive that cache line (or, in various cases, the minimum state in which a more acceptable state may be provided). Transaction request D205 may be a write transaction in which the requester attempts to write data to a cache line, or a read transaction in which the requester attempts to read data from a cache line. For example, transaction request D205 may specify an unrelaxed ordered dynamic random access memory (DRAM) request. In some embodiments, the coherency controller D210 is also configured to issue a memory request D222 to memory D130 to access data from memory D130 on behalf of a component of the SOC D100, and to receive a memory response D224 which may contain the requested data.

[0343] As shown in the figure, I / O agent D142 receives transaction requests D205 from peripheral device D144. I / O agent D142 can receive a series of write transaction requests D205, a series of read transaction requests D205, or a combination of read and write transaction requests D205 from a given peripheral device D144. For example, within a set time interval, I / O agent D142 may receive four read transaction requests D205 from peripheral device D144A and three write transaction requests D205 from peripheral device D144B. In various embodiments, the transaction requests D205 received from peripheral device D144 must be completed in a specific order (for example, they must be completed in the order in which they were received from peripheral device D144). Instead of waiting for transaction request D205 to complete before starting work on the next transaction request D205 in sequence, in various embodiments, the I / O agent D142 performs work on subsequent requests D205 by preemptively acquiring exclusive ownership of the target cache line. Thus, the I / O agent D142 can issue an exclusive ownership request D215 to the memory controller D120 (in particular, the coherency controller D210). In some examples, a set of transaction requests D205 may target cache lines managed by different memory controllers D120, and thus the I / O agent D142 can issue an exclusive ownership request D215 to the appropriate memory controller D120 based on those transaction requests D205. For read transaction requests D205, the I / O agent D142 can acquire exclusive read ownership. For write transaction requests D205, the I / O agent D142 can acquire exclusive write ownership.

[0344] The coherency controller D210 is a circuit configured to receive requests (e.g., exclusive ownership requests D215) from the interconnect D105 (e.g., via one or more queues included in the memory controller D120) that target cache lines mapped to memory D130 to which the memory controller D120 is coupled. The coherency controller D210 can process these requests and generate a response (e.g., exclusive ownership response D217) containing the data of the requested cache line, while also maintaining cache coherency within the SOC D100. To maintain cache coherency, the coherency controller D210 can use a directory D220. In various embodiments, the directory D220 is a storage array having a set of entries, each of which can track the coherency state of individual cache lines in the system. In some embodiments, the entries also track the location of data in the cache line. For example, an entry in directory D220 may indicate that data for a particular cache line is cached in the cache D114 of caching agent D110 in an active state. (While exclusive ownership is discussed, in some cases a cache line may be shared among multiple cache-enabled entities (e.g., caching agent D110) for read purposes, and therefore shared ownership may be provided.) To provide exclusive ownership of a cache line, coherency controller D210 may ensure that the cache line is not stored outside of memory D130 and memory controller D120 in an active state. Thus, based on the directory entry associated with the cache line targeted by the exclusive ownership request D215, in various embodiments, coherency controller D210 determines which component (e.g., caching agent D110, I / O agent D142, etc.) should receive snoop D225 and the type of snoop D225 (e.g., invalidate, change to ownership, etc.).For example, the memory controller D120 can determine that the caching agent 110 is storing data for a cache line requested by the I / O agent D142, and therefore can issue a snoop D225 to the caching agent D110, as shown in Figure 45. In some embodiments, the coherency controller D210 does not target any particular component, but instead broadcasts a snoop D225 that is observed by many of the components of the SOC D100.

[0345] In various embodiments, at least two types of snooping are supported: snoop-forward and snoop-back. Snoop-forward messages may be used to cause a component (e.g., cache agent D110) to transfer data from a cache line to a requesting component, and snoop-back messages may be used to cause a component to return data from a cache line to the memory controller D120. Supporting snoop-forward and snoop-back flows allows for both 3-hop (snoop-forward) and 4-hop (snoop-back) behavior. For example, snoop-forward may be used to minimize the number of messages when a cache line is provided to a component, since the component may store the cache line and potentially use the data within it. On the other hand, a non-cacheable component may not store the entire cache line, and therefore copy-back to memory can ensure that the full cache line data is captured by the memory controller D120. In various embodiments, the caching agent D110 receives snoop D225 from the memory controller D120, processes snoop D225 to update the cache line state (e.g., invalidate the cache line), and returns a copy of the cache line data (if specified by snoop D225) to the initial ownership claimant or the memory controller D120. The snoop response D227 (or "complete") is, in various embodiments, a message indicating that a state change has been made and, if applicable, provides a copy of the cache line data. When a snoop-forward mechanism is used, the data is provided to the requesting component in three hops via the interconnect D105: a request from the requesting component to the memory controller D120, a snoop from the memory controller D120 to the caching, and a snoop response from the caching component to the requesting component.When the snoopback mechanism is used, four hops may occur, i.e., a request and snoop, a snoop response to the memory controller D120 by the caching component, and data from the memory controller D120 to the requesting component, as in the case of a three-hop protocol.

[0346] In some embodiments, the coherency controller D210 may update directory D220 when snoop D225 is generated and transmitted, instead of when snoop response D227 is received. When the requested cache line is reclaimed by the memory controller D120, in various embodiments, the coherency controller D210 grants exclusive read (or write) ownership to the ownership requester (e.g., I / O agent D142) via exclusive ownership response D217. Exclusive ownership response D217 may contain the data of the requested cache line. In various embodiments, the coherency controller D210 updates directory D220 to indicate that the cache line has been granted ownership to the ownership requester.

[0347] For example, I / O agent D142 may receive a series of read transaction requests D205 from peripheral device D144A. In response to one of these requests, I / O agent D142 may send an exclusive read ownership request D215 to memory controller D120 for data associated with a particular cache line (or, if the cache line is managed by another memory controller D120, the exclusive read ownership request D215 is sent to the other memory controller D120). Coherency controller D210 may determine, based on entries in directory D220, that cache agent D110 is currently storing data associated with a particular cache line that is in an active state. Therefore, coherency controller D210 sends a snoop D225 to caching agent D110, causing caching agent D110 to relinquish ownership of that cache line and return a snoop response D227 that may contain the cache line data. After receiving the snoop response D227, the coherency controller D210 can generate an exclusive ownership response D217 and send it to the I / O agent D142, thereby providing the I / O agent D142 with the cache line data and exclusive ownership of the cache line.

[0348] After receiving exclusive ownership of a cache line, in various embodiments, I / O agent D142 waits until the corresponding transaction can be completed (according to the ordering rules). That is, it waits until the corresponding transaction becomes the highest-level transaction and an ordering dependency resolution for the transaction exists. For example, I / O agent D142 may receive transaction request D205 from peripheral device D144 to perform write transactions A-D. I / O agent D142 may acquire exclusive ownership of the cache line associated with transaction C, but transactions A and B may not be completed. As a result, I / O agent D142 waits until transactions A and B are completed before writing the relevant data to the cache line associated with transaction C. After completing a given transaction, in various embodiments, I / O agent D142 provides the transaction requester (e.g., peripheral device D144A) with a transaction response D207 indicating that the requested transaction has been performed. In various cases, I / O agent D142 can acquire exclusive read ownership of a cache line, execute a set of read transactions on the cache line, and then release exclusive read ownership of the cache line without performing any writes to the cache line while exclusive read ownership was held.

[0349] In some cases, I / O agent D142 may receive multiple transaction requests D205 targeting the same cache line (within a reasonably short period of time), and as a result, I / O agent D142 can perform bulk reads and writes. For example, two write transaction requests D205 received from peripheral device D144A may target the lower and upper portions of the cache line, respectively. Therefore, I / O agent D142 can acquire exclusive write ownership of the cache line and retain the data associated with the cache line until at least both write transactions are completed. Thus, in various embodiments, I / O agent D142 can transfer executive ownership between transactions targeting the same cache line. That is, I / O agent D142 does not need to send an ownership request D215 for each individual transaction request D205. In some cases, I / O agent D142 can transfer executive ownership from a read transaction to a write transaction (or vice versa), but in other cases, I / O agent D142 only transfers executive ownership between transactions of the same type (for example, from one read transaction to another read transaction).

[0350] In some cases, I / O agent D142 may lose exclusive ownership of a cache line before it executes the associated transaction on the cache line. For example, while waiting for a transaction to become senior enough for the transaction to be executed, I / O agent D142 may receive a snoop D225 from memory controller D120 as a result of another I / O agent D142 attempting to acquire exclusive ownership of the cache line. After relinquishing exclusive ownership of a cache line, in various embodiments, I / O agent D142 determines whether to reacquire ownership of the lost cache line. If the lost cache line is associated with one pending transaction, I / O agent D142 will often not reacquire exclusive ownership of the cache line, but in some cases, if the pending transaction is behind a set number of transactions (and therefore not about to become a higher transaction), I / O agent D142 may issue an exclusive ownership request D215 for the cache line. However, in various embodiments, if there are a threshold number of pending transactions directed to the cache line (e.g., two pending transactions), the I / O agent D142 reacquires exclusive ownership of the cache line.

[0351] Referring next to Figure 46A, a block diagram of exemplary elemen...

Claims

1. It is a system, Multiple processor cores, Multiple graphics processing units, Multiple peripheral devices different from the aforementioned processor core and graphics processing unit, One or more memory controller circuits configured to interface with system memory, The system comprises one or more memory controller circuits and an interconnect fabric configured to provide communication between the processor core, the graphics processing unit, and the peripheral devices. A system in which the processor core, the graphics processing unit, the peripheral devices, and the memory controller are configured to communicate via an integrated memory architecture.

2. The system according to claim 1, wherein the processor core, the graphics processing unit, and the peripheral devices are configured to access any address in the integrated address space defined by the integrated memory architecture.

3. The system according to claim 2, wherein the integrated address space is a virtual address space different from the physical address space provided by the system memory.

4. The system according to any one of claims 1 to 3, wherein the integrated memory architecture provides a common set of semantics for memory access by the processor core, the graphics processing unit, and the peripheral devices.

5. The system according to claim 4, wherein the semantics include memory ordering properties.

6. The system according to claim 4 or 5, wherein the semantics include service quality attributes.

7. The system according to any one of claims 4 to 6, wherein the semantics include memory coherency.

8. The system according to any one of claims 1 to 7, wherein the one or more memory controller circuits each include an interface to one or more memory devices that can be mapped to random access memory.

9. The system according to claim 8, wherein one or more memory devices include dynamic random access memory (DRAM).

10. The system according to any one of claims 1 to 9, further comprising one or more levels of cache between the processor core, the graphics processing unit, the peripheral device and the system memory.

11. The system according to claim 10, wherein each of the one or more memory controller circuits includes a memory cache inserted between the interconnect fabric and the system memory, and each of the memory caches is one of the one or more levels of caches.

12. The system according to any one of claims 1 to 11, wherein the interconnect fabric includes at least two networks having heterogeneous interconnect topologies.

13. The system according to any one of claims 1 to 12, wherein the interconnect fabric includes at least two networks having heterogeneous operating characteristics.

14. The system according to claim 12 or 13, wherein the at least two networks include a coherent network interconnecting the processor core and the one or more memory controller circuits.

15. The system according to any one of claims 12 to 14, wherein the at least two networks include relaxation sequence networks coupled to the graphics processing unit and the one or more memory controller circuits.

16. The system according to claim 15, wherein the peripheral devices include a subset of devices, the subset including one or more machine learning accelerator circuits or relaxation sequence bulk media devices, and the relaxation sequence network is further coupled to the subset of devices to one or more memory controller circuits.

17. The system according to any one of claims 12 to 16, wherein the at least two networks include an input / output network coupled to interconnect the peripheral devices and the one or more memory controller circuits.

18. The system according to claim 17, wherein the peripheral device includes one or more real-time devices.

19. The system according to any one of claims 12 to 18, wherein the at least two networks include a first network having one or more characteristics for reducing latency compared to a second network among the at least two networks.

20. The system according to claim 19, wherein one or more of the characteristics include a route shorter than the second network.

21. The system according to claim 19 or 20, wherein the one or more characteristics include the wiring in a metal layer closer to the surface of the substrate on which the system is mounted, rather than the wiring for the second network.

22. The system according to any one of claims 12 to 21, wherein the at least two networks include a first network having one or more characteristics for increasing bandwidth compared to a second network among the at least two networks.

23. The system according to claim 22, wherein one or more of the characteristics include a wider interconnect compared to the second network.

24. The system according to claim 22 or 23, wherein the one or more of the characteristics include wiring in a metal layer that is further from the surface of the substrate on which the system is mounted than the wiring for the second network.

25. The system according to any one of claims 12 to 24, wherein the interconnect topology employed by the at least two networks includes at least one of a star topology, a mesh topology, a ring topology, a tree topology, a fat tree topology, a hypercube topology, or a combination of one or more of the interconnect topologies.

26. The system according to any one of claims 12 to 25, wherein the operating characteristics employed by the at least two networks include at least one of strongly ordered memory coherence or relaxed ordered memory coherence.

27. The system according to any one of claims 12 to 26, wherein the at least two networks are physically and logically independent.

28. The system according to any one of claims 12 to 27, wherein the at least two networks are physically separated in a first operating mode, and the first network and the second network are virtual in a second operating mode and share a single physical network.

29. The system according to any one of claims 1 to 28, wherein the processor core, the graphics processing unit, the peripheral devices, and the interconnect fabric are distributed across two or more integrated circuit dies.

30. The system according to claim 29, wherein the integrated address space defined by the integrated memory architecture extends across the two or more integrated circuit dies in a manner transparent to software running on the processor core, the graphics processing unit, or the peripheral device.

31. The system according to claim 29 or 30, wherein the interconnect fabric is configured to route, the interconnect fabric extends across the two integrated circuit dies, and communication is routed transparently between the source and the destination with respect to the source and destination locations on the integrated circuit dies.

32. The system according to any one of claims 29 to 31, wherein the interconnect fabric extends across the two integrated circuit dies using hardware circuitry to automatically route communication between the source and the destination, regardless of whether the source and destination are on the same integrated circuit dies.

33. The system according to any one of claims 29 to 32, further comprising at least one interposer device configured to connect the bus of the interconnect fabric across two or more integrated circuit dies.

34. The system according to any one of claims 1 to 33, wherein a given integrated circuit die includes a local interrupt distribution circuit for distributing interrupts among processor cores within the given integrated circuit die.

35. The system according to claim 34, comprising two or more integrated circuit dies, each including a local interrupt distribution circuit, at least one of the two or more integrated circuit dies including a global interrupt distribution circuit, and the local interrupt distribution circuit and the global interrupt distribution circuit realizing a multilevel interrupt distribution scheme.

36. The system according to claim 35, wherein the global interrupt distribution circuit is configured to sequentially transmit interrupt requests to the local interrupt distribution circuit, and the local interrupt distribution circuit is configured to sequentially transmit the interrupt requests to the local interrupt destination before responding to the interrupt requests from the global interrupt distribution circuit.

37. The system according to any one of claims 1 to 36, wherein a given integrated circuit die includes a power manager circuit configured to manage the local power state of the given integrated circuit die.

38. The system according to claim 37, comprising two or more integrated circuit dies, each including a power manager circuit configured to manage the local power state of the integrated circuit die, wherein at least one of the two or more integrated circuit dies includes another power manager circuit configured to synchronize the power manager circuits.

39. The system according to any one of claims 1 to 38, wherein the peripheral device includes one or more of the following: an audio processing device, a video processing device, a machine learning accelerator circuit, a matrix arithmetic accelerator circuit, a camera processing circuit, a display pipeline circuit, a non-volatile memory controller, a peripheral component interconnect controller, a security processor, or a serial bus controller.

40. The interconnect fabric interconnects coherent agents, as described in any one of claims 1 to 39.

41. The system according to claim 40, wherein each of the processor cores corresponds to a coherent agent.

42. The system according to claim 40, wherein the cluster of processor cores corresponds to a coherent agent.

43. The system according to any one of claims 1 to 42, wherein one of the peripheral devices is a non-coherent agent.

44. The system according to claim 43, further comprising an input / output agent inserted between the given peripheral device and the interconnect fabric, wherein the input / output agent is configured to implement the coherency protocol of the interconnect fabric with respect to the given peripheral device.

45. The system according to claim 44, wherein the input / output agent uses the coherency protocol to ensure the ordering of requests from the given peripheral device.

46. The system according to claim 44 or 45, wherein the input / output agent is configured to couple a network of two or more peripheral devices to the interconnect fabric.

47. The system according to any one of claims 1 to 46, further comprising a hash circuit configured to distribute memory request traffic to system memory according to a selectively programmable hashing protocol.

48. The system according to claim 47, wherein the programming of at least one of the programmable hashing protocols distributes a set of memory requests evenly across a plurality of memory controllers in the system for the diverse range of memory requests in the set of memory requests.

49. The system according to claim 29, wherein the programming of at least one of the programmable hashing protocols distributes adjacent requests in memory space to physically separated memory interfaces at a specified granularity.

50. The system according to any one of claims 1 to 49, further comprising a plurality of directories configured to track the coherency state of a subset of the integrated memory address space specified by the integrated memory architecture, wherein the plurality of directories are distributed within the system.

51. The system according to claim 50, wherein the plurality of directories are distributed across the memory controller.

52. A given memory controller of one or more memory controller circuits includes a directory configured to track a plurality of cache blocks corresponding to data in a portion of the system memory to which the given memory controller interfaces, the directory being configured to track which of a plurality of caches in the system caches a given cache block among the plurality of cache blocks, and the directory is accurate with respect to memory requests that have been ordered and processed in the directory, even if the memory request has not yet been completed in the system, according to any one of claims 1 to 51.

53. The system according to claim 52, wherein the given memory controller is configured to issue one or more coherency maintenance commands for the given cache block based on a memory request for the given cache block, the one or more coherency maintenance commands include a cache state for the given cache block in a corresponding cache among the plurality of caches, and the corresponding cache is configured to delay processing of the given coherency maintenance command on the basis that the cache state in the corresponding cache does not match the cache state in the given coherency maintenance command.

54. The system according to claim 52 or 53, wherein a first cache is configured to store the given cache block in a primary shared state, a second cache is configured to store the given cache block in a secondary shared state, and a given memory controller is configured to cause the first cache to transfer the given cache block to the requester based on the memory request and the primary shared state in the first cache.

55. The given memory controller is configured to issue one of a first coherency maintenance command and a second coherency maintenance command to a first cache among the plurality of caches based on the type of a first memory request, and the first cache is configured to transfer a first cache block to the requester that issued the first memory request based on the first coherency maintenance command. The system according to any one of claims 52 to 54, wherein the first cache is configured to return the first cache block to the given memory controller based on the second coherency maintenance command.

56. It is an integrated circuit, Multiple processor cores, Multiple graphics processing units, Multiple peripheral devices different from the aforementioned processor core and graphics processing unit, One or more memory controller circuits configured to interface with system memory, The one or more memory controller circuits and the interconnect fabric configured to provide communication between the processor core, the graphics processing unit, and the peripheral devices, An integrated circuit comprising: an off-chip interconnect coupled to the interconnect fabric and configured to couple the interconnect fabric to a corresponding interconnect fabric on another instance of the integrated circuit, wherein the interconnect fabric and the off-chip interconnect provide an interface for transparently connecting one or more memory controller circuits, the processor core, the graphics processing unit, and the peripheral devices in either a single instance of the integrated circuit or two or more instances of the integrated circuit.

57. It is a method, A method comprising communicating between a plurality of processing cores, a plurality of graphics processing units, a plurality of peripheral devices different from the plurality of processor cores and the plurality of graphics processing units, and one or more memory controller circuits via an interconnect fabric in a system, through an integrated memory architecture.