Coherence protocol for high-performance interconnects
Patent Information
- Application Number
- DE112013005086
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2013-03-15
- Filing Date
- 2013-03-15
- Publication Date
- 2026-09-17
- Estimated Expiration
- 2033-03-15
AI Technical Summary
As computer systems evolve with increased complexity and higher processing power, the demand for efficient communication between components, particularly in high-performance computing environments, exceeds the capabilities of existing interconnect architectures, necessitating improved interconnect solutions to manage bandwidth and energy consumption effectively.
The implementation of a High Performance Interconnect (HPI) system with a layered protocol architecture, including a coherent coherency protocol, routing layer, link layer, and physical layer, to facilitate efficient data transfer and maintain data coherence across multiple processors and devices, utilizing point-to-point serial links and credit-based flow control.
Enhances communication efficiency and data coherence, optimizing performance and energy usage in complex computing systems by providing scalable and flexible interconnect solutions adaptable to various topologies and devices, including servers, smartphones, and embedded applications.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
AREA
[0001] The present disclosure relates generally to the field of computer development and in particular to software development involving the coordination of mutually dependent restricted systems. BACKGROUND
[0002] Advances in semiconductor processing and logic design have enabled an increase in the amount of logic that can reside on integrated circuit devices. As a consequence, computer system configurations have evolved from a single or multiple integrated circuits within a system to multiple cores, multiple hardware threads, and multiple logical processors residing on single integrated circuits, along with other interfaces integrated into such processors. A processor or integrated circuit typically comprises a single physical processor chip, which can include any number of cores, hardware threads, logical processors, interfaces, memories, controller hubs, and so on.
[0003] As a result of the increased ability to pack more processing power into smaller packages, smaller computing devices have become more popular. Smartphones, tablets, ultrathin notebooks, and other consumer devices have proliferated exponentially. However, these smaller devices depend on servers for both data storage and complex processing that exceeds their form factor. Consequently, demand in the high-performance computing market (i.e., the server space) has also increased. For example, modern servers typically contain not just a single multi-core processor, but multiple physical processors (also known as multiple sockets) to increase computing power. However, as processing power increases along with the number of devices in a computer system, communication between sockets and other devices becomes more critical.
[0004] Indeed, interconnects have evolved from more traditional multipoint interconnect buses, which primarily handled electrical communications, into complete interconnect architectures that enable high-speed communication. Unfortunately, the requirement for future processors to achieve even higher power consumption at even greater speeds places a corresponding demand on the capabilities of existing interconnect architectures. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] Fig. Figure 1 illustrates a simplified block diagram of a system comprising a point-to-point intermediate connection for connecting I / O devices in a computer system according to one embodiment;
[0006] Fig. Figure 2 illustrates a simplified block diagram of a shift log stack according to one embodiment;
[0007] Fig.Figure 3 illustrates an embodiment of a transaction descriptor;
[0008] Fig. Figure 4 illustrates an embodiment of a serial point-to-point link;
[0009] Fig. Figure 5 illustrates embodiments of potential high-performance interconnect (HPI) system configurations;
[0010] Fig. Figure 6 illustrates a shift log stack associated with HPI.
[0011] Fig. Figure 7 illustrates a flowchart of an exemplary coherence protocol conflict handling process;
[0012] Fig. Figure 8 illustrates a flowchart of another exemplary coherence protocol conflict handling;
[0013] Fig. Figure 9 illustrates a flowchart of another exemplary coherence protocol conflict handling;
[0014] Fig.Figure 10 illustrates an embodiment of a block diagram for a computer system that includes a multi-core processor.
[0015] Identical reference symbols and names in the different drawings denote the same elements. DETAILED DESCRIPTION
[0016] The following description sets forth numerous specific details, such as examples of specific processor types and system configurations, specific hardware structures, specific architectural and microarchitectural details, specific register configurations, specific instruction types, specific system components, specific processor pipeline stages, specific intermediate link layers, specific packet / transaction configurations, specific transaction names, specific protocol exchange operations, specific link widths, specific implementations, and specific operation, etc., to provide a comprehensive understanding of the present invention. However, it may be apparent to a person skilled in the art that these specific details are not necessarily required to put the subject matter of the present disclosure into practice.In other cases, a detailed description of known components and procedures, such as specific and alternative processor architectures, specific logic circuits or specific logic code for described algorithms, specific firmware code, low-level interconnection operation, specific logic configurations, specific manufacturing techniques and materials, specific compiler implementations, specific expression of algorithms in code, specific shutdown and gate control techniques or specific shutdown and gate control logic, and other specific operational details of a computer system, has been avoided in order to prevent unnecessary complication of the present disclosure.
[0017] Although the following embodiments may be described with reference to energy saving, energy efficiency, processing efficiency, etc., in specific integrated circuits, such as computer platforms or microprocessors, other embodiments are applicable to other types of integrated circuits and logic devices. Similar techniques and teachings of embodiments described herein may be applied to other types of circuits or semiconductor devices that also benefit from such features. For example, the disclosed embodiments are not applicable to server computer systems, desktop computer systems, laptops, or ultrabooks. TMThese techniques are not limited to mobile devices but can also be used in other devices, such as portable devices, smartphones, tablets, other thin notebooks, system-on-a-chip (SoC) components, and embedded applications. Some examples of portable devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and portable PCs. Similar high-performance intermediary techniques can be applied to increase performance (or even save power) in a low-power intermediary connection. Embedded applications typically include a microcontroller, a digital signal processor (DSP), a system-on-a-chip, a network PC, set-top boxes, network hubs, wide area network (WAN) switches, or any other system capable of performing the functions and operations described below.Furthermore, the devices, methods, and systems described herein are not limited to physical computer devices but may also relate to software optimizations for energy saving and efficiency. As may be evident from the following description, the embodiments of methods, devices, and systems described herein (whether relating to hardware, firmware, software, or a combination thereof) may be considered essential for a future in which environmentally friendly technology and performance considerations are in balance.
[0018] As computer systems advance, their components become more complex. The interconnect architecture used to couple and communicate between these components has also increased in complexity to ensure sufficient bandwidth for optimal component operation. Furthermore, different market segments require adaptation of various aspects of interconnect architectures to their specific needs. For example, servers require higher performance, while the mobile ecosystem is sometimes willing to sacrifice overall performance for energy savings. Nevertheless, a key objective of most fabrics is to deliver the highest possible performance with maximum energy efficiency. Moreover, a wide variety of interconnects can potentially benefit from the subject matter described herein.For example, among other examples, the PCIe (Peripheral Component Interconnect (PCI) Express) interconnect fabric architecture and the QPI (QuickPath Interconnect) fabric architecture can potentially be improved according to one or more of the principles described herein.
[0019] Fig. Figure 1 illustrates an embodiment of a fabric consisting of point-to-point links connecting a set of components. A system 100 includes a processor 105 and a system memory 110 , which is equipped with a controller hub 115 is coupled. The processor 105 The processor can be any processing element, such as a microprocessor, a host processor, an embedded processor, a coprocessor, or another processor. 105 is with the controller hub 115 by a frontside bus (FSB) 106coupled. In one embodiment, the FSB 106 a serial point-to-point link, as described below. In another embodiment, the link comprises 106 a serial, differential interface architecture that is compatible with another interface standard.
[0020] The system memory 110 includes any storage device, such as random access memory (RAM), non-volatile (NV) memory, or other memory that can be accessed by devices in the system 100 can be accessed. The system memory 110 is with the controller hub 115 through a memory interface 116 coupled. Examples of a memory interface include a dual data rate (DDR) memory interface, a dual-channel DDR memory interface, and a DRAM (dynamic RAM) memory interface.
[0021] In one embodiment, the controller hub 115 comprise a root hub, root complex, or root controller, such as in a PCIe intermediate connection hierarchy. Examples of controller hubs 115 They comprise a chipset, a memory controller hub (MCH), a northbridge, an interconnect controller hub (ICH), a southbridge, and a root controller / hub. The term chipset often refers to two physically separate controller hubs, such as a memory controller hub (MCH) coupled to an interconnect controller hub (ICH). It's worth noting that current systems often integrate the MCH into the processor. 105 integrated, while the controller 115It serves to communicate with I / O devices in a similar manner to that described below. In some embodiments, partner-to-partner routing is optionally provided by the root complex. 115 supports.
[0022] This is the controller hub 115 through a serial link 119 with a switch or a bridge 120 coupled. Input / output modules 117 and 121 , which are also known as interfaces / ports 117 and 121 These can be described as including / implementing a layer protocol stack to facilitate communication between the controller hub. 115 and the switch 120 to provide. In one embodiment, several devices are able to use the switch. 120 to be coupled.
[0023] The switch or bridge 120 routes packages / messages from a facility 125upstream, i.e., a hierarchy upwards to a root complex, to the controller hub 115 , and downstream, i.e., a hierarchy downwards away from a root controller, from the processor 105 or the system memory 110 for the establishment 125 The Switch 120 In one embodiment, it is referred to as a logical arrangement of multiple virtual PCI-to-PCI bridge devices. The device 125This includes any internal or external device or component intended to be coupled with an electronic system, such as an I / O device, a network interface controller (NIC), an add-in card, an audio processor, a network processor, a hard disk drive, a storage device, a CD / DVD-ROM drive, a monitor, a printer, a mouse, a keyboard, a router, a portable storage device, a FireWire device, a USB (Universal Serial Bus) device, a scanner, and other input / output devices. In PCIe terminology, such a device is often referred to as an endpoint. Although not specifically depicted, the device can 125 include a bridge (e.g., a PCIe-to-PCI / PCI-X bridge) to support legacy or other versions of devices or intermediate interconnect fabrics supported by such devices.
[0024] Also a graphics accelerator 130 can be done via a serial link 132 with the controller hub 115 be coupled. In one embodiment, the graphics accelerator 130 coupled with an MCH, which is coupled with an ICH. The switch 120 and accordingly the I / O equipment 125 are then coupled with the ICH. Additionally, I / O modules serve this purpose. 131 and 118 to implement a layer protocol stack and associated logic to communicate between the graphics accelerator 130 and the controller hub 115 to communicate. Similar to the preceding MCH discussion, a graphics controller or graphics accelerator can be used. 130 even into the processor 105 be integrated.
[0025] Turning towards Fig. Figure 2 illustrates an embodiment of a shift log stack. The shift log stack 200It can include any form of layered communication stack, such as a QPI stack, a PCIe stack, a next-generation high-performance computing interlink (HPI) stack, or any other layered stack. In one embodiment, the protocol stack can 200 a transaction layer 205 , a link layer 210 and a physical layer 220 include an interface, such as the interfaces 117 , 118 , 121 , 122 , 126 and 131 in Fig. 1. can be considered a communication protocol stack 200 This can be represented as a communication protocol stack. It can also be referred to as a module or interface that implements / comprises a protocol stack.
[0026] Packets can be used to communicate information between components. Packets can be used in the transaction layer.205 and the data link layer 210 Packets are formed to transmit information from the sending component to the receiving component. As the sent packets pass through the other layers, they are augmented with additional information used to handle the packets at those layers. On the receiving side, the reverse process takes place, and the packets are processed from their physical layer representation. 220 into the representation of the data link layer 210 and finally (for transaction layer packages) converted into the form required by the transaction layer 205 can be processed by the receiving device.
[0027] In one embodiment, the transaction layer can 205 provide an interface between a processing core of a facility and the intermediate link architecture, such as the data link layer. 210and the physical layer 220 In this respect, a major responsibility can be attributed to the transaction layer. 205 This includes the packaging and unpackaging of packets (i.e., transaction layer packets or TLPs). The transaction layer 205 It can also control credit-based flow control for TLPs. In some implementations, split transactions—that is, transactions with a time-separated request and response—can be used, allowing one link, among others, to carry other traffic while the destination facility gathers data for the response.
[0028] Credit-based flow control can be used to implement virtual channels and networks that utilize the Interconnect Fabric. For example, an institution might allocate an initial amount of credit to each of the receive buffers in the transaction layer. 205Display. An external device at the opposite end of the link, such as the controller hub. 115 in Fig. 1. The system can count the number of credits consumed by each TLP. A transaction can be sent as long as it does not exceed a credit limit. Upon receiving a response, a credit amount is restored. One example of an advantage among other potential benefits of such a credit scheme is that the latency of credit return does not impact performance, provided the credit limit is not reached.
[0029] In one embodiment, four transaction address spaces can include a configuration address space, a memory address space, an input / output address space, and a message address space. Memory space transactions comprise one or more read and write requests to transfer data to and from a memory location. In one embodiment, memory space transactions are capable of using two different address formats, such as a short address format like a 32-bit address or a long address format like a 64-bit address. Configuration space transactions can be used to access the configuration space from different devices connected to the intermediate link. Configuration space transactions can include read and write requests.Message space transactions (or simply messages) can also be defined to support in-band communication between intermediary connection agents. Therefore, the transaction layer can... 205 in an exemplary embodiment package header / payload 206 package.
[0030] Briefly referring to Fig. Figure 3 illustrates an exemplary embodiment of a transaction layer package descriptor. In one embodiment, the transaction descriptor can 300 It should be a mechanism for transmitting transaction information. In this respect, the transaction descriptor supports this. 300 The identification of transactions within a system. Other potential uses include tracking modifications to the default transaction order and associating transactions with channels. For example, the transaction descriptor 300 a global identifier field302 , an attribute field 304 and a channel identification field 306 include. In the illustrated example, the global identifier field is 302 represented in such a way that it is a local transaction identifier field 308 and a source identification field 310 includes. In one embodiment, the global transaction identifier 302 unambiguous for all outstanding requirements.
[0031] According to one implementation, the local transaction identifier field 308 A field generated by a requesting agent, and can be unique for all pending requests that require completion by that requesting agent. Additionally, in this example, the source identifier identifies... 310 The requester agent is uniquely identified within the intermediate connection hierarchy. Accordingly, the local transaction identifier field represents this. 308 together with the source ID 310Provides a global identification of a transaction within a hierarchy area.
[0032] The attribute field 304 It specifies the characteristics and relationships of the transaction. In this respect, the attribute field 304 potentially used to provide additional information that allows for modification of the standard handling of transactions. In one embodiment, the attribute field includes 304 a priority field 312 , a reserved field 314 , an order field 316 and a no-snooping area 318 . Here, the priority subfield can be used. 312 It can be modified by an initiator to assign a priority to the transaction. The reserved attribute field 314It will be reserved for future or vendor-defined use. Possible usage models that utilize priority and security attributes can be implemented using the reserved attribute field.
[0033] In this example, the order attribute field is used. 316 Used to provide optional information that transmits the order type, which can modify default order rules. According to one example implementation, an order attribute of "0" means that default order rules should be applied, while an order attribute of "1" means random ordering, where write operations can overtake write operations in the same direction and read completions can overtake write operations in the same direction. The snoop attribute field 318This is used to determine whether transactions are being monitored. As shown, the channel ID field identifies... 306 a channel with which a transaction is associated.
[0034] Back to the discussion of Fig. 2 can be a link layer 210 , which is also known as the data link layer 210 is described as an intermediate stage between the transaction layer 205 and the physical layer 220 function. In one embodiment, a responsibility of the data link layer is assigned. 210 The provision of a reliable mechanism for exchanging transaction layer packets (TLPs) between two components on a link. One side of the data link layer. 210 accepted by the transaction layer 205 Packaged TLPs apply a packet sequence identifier. 211, i.e., an identification number or package number, calculates an error detection code, i.e., CRC. 212 , and applies this and transmits the modified TLPs to the physical layer 220 for transmission via a physical device to an external device.
[0035] In one example, the physical layer comprises 220 a logical subblock 221 and an electrical sub-block 222 , to physically send a packet to an external facility. The logical subblock is... 221 for the “digital” functions of the physical layer 221 responsible. In this respect, the logical subblock can be a transmit section for preparing outgoing information for transmission by the physical subblock. 222 and a receiving section to identify and prepare received information before passing it on to the link layer 210include.
[0036] The physical block 222 It comprises a sender and a receiver. The sender is defined by the logical subblock. 221 The receiver is supplied with symbols, which the sender serializes and forwards to an external device. The receiver is supplied with serialized symbols from an external device and converts the received signals into a bitstream. The bitstream is deserialized and sent to the logical subblock. 221 supplied. In an exemplary embodiment, an 8b / 10b transmission code is used, whereby ten-bit symbols are sent / received. Special symbols are used to frame a packet. 223 used. In addition, in one example, the receiver also provides a symbol clock that is recovered from the incoming serial stream.
[0037] Although, as already mentioned, the transaction layer 205 , the link layer 210and the physical layer 220 While a layered protocol stack is discussed in relation to a specific implementation of a protocol stack (such as a PCIe protocol stack), it is not limited to such a stack. In fact, any layered protocol can be included or implemented and adopt the features discussed herein. As an example, a port or interface represented as a layered protocol may include: (1) a first layer for packaging packets, i.e., a transaction layer; (2) a second layer for sequencing packets, i.e., a link layer; and (3) a third layer for sending the packets, i.e., a physical layer. A high-performance interlink layer protocol, as described herein, is used as a specific example.
[0038] Next, with reference to Fig.Figure 4 illustrates an exemplary embodiment of a serial point-to-point fabric. A serial point-to-point link can include any transmission path for sending serial data. In the illustrated embodiment, a link can include two differentially controlled low-voltage signal pairs: a transmit pair. 406 / 411 and a receiving couple 412 / 407 Accordingly, an institution comprises 405 Transfer logic 406 , to send data to an institution 410 to send, and receive logic 407 , to obtain data from the institution 410 to receive. In other words, some implementations of a link have two transmit paths, i.e., paths. 416 , 417 , and two receiving paths, i.e. paths 418 and 419 , contain.
[0039] "Transmission path" refers to any path for sending data, such as a transmission line, a copper wire, an optical line, a wireless communication channel, an infrared communication link, or any other communication path. A connection between two devices, such as a facility 405 and facility 410 , is used as a link, such as link 415 A link can support one lane – where each lane represents a set of differential signal pairs (one pair for transmitting, one pair for receiving). To scale bandwidth, a link can aggregate multiple lanes, denoted by xN, where N is any supported link width, such as 1, 2, 4, 8, 12, 16, 32, 64, or wider.
[0040] A “differential pair” can refer to two transmission paths, such as the lines. 416 and 417, for sending differential signals. As an example, the line controls 417 from a logical high level to a logical low level, i.e., falling edge when the line 416 The signal switches from a low-voltage level to a high-voltage level, i.e., a rising edge. Differential signals potentially exhibit better electrical characteristics, such as improved signal integrity (i.e., reduced cross-coupling, voltage over / undershoot, and ringing), among other advantages. This allows for a better clock window, which in turn enables faster transmission frequencies.
[0041] In one embodiment, a new high-performance interconnect (HPI) is provided. HPI can comprise a cache-coherent, link-based, next-generation interconnect. As an example, HPI can be used in high-performance computing platforms, such as workstations or servers, including systems that typically use PCIe or another interconnect protocol to connect processors, accelerators, I / O devices, and the like. However, HPI is not limited to these applications. Instead, HPI can be used in any of the systems or platforms described herein. Furthermore, the individual concepts developed can be applied to other interconnects and platforms, such as PCIe, MIPI, QPI, and so on.
[0042] To support multiple devices, HPI, in one exemplary embodiment, can be instruction set architecture (ISA) agnostic (i.e., HPI can be implemented in several different devices). In another scenario, HPI can also be used to connect high-performance I / O devices, not just processors or accelerators. For example, a high-performance PCIe device can be coupled to HPI via a suitable translation bridge (i.e., HPI to PCIe). Furthermore, the HPI links can be used by many HPI-based devices, such as processors, in various configurations (e.g., star, ring, mesh, etc.). Fig. Figure 5 illustrates exemplary implementations of several potential multi-socket configurations. A two-socket configuration 505As shown, a configuration can include two HPI links; however, other implementations may use a single HPI link. For larger topologies, any configuration can be used as long as an identifier (ID) can be assigned, along with other additional or substitute features, and some form of virtual path exists. As shown, one example has a four-socket configuration. 510 an HPI link from each processor to another. But in the eight-socket implementation, which is in configuration 515As shown, not all sockets are directly connected via an HPI link. However, if a virtual path or channel exists between the processors, the configuration is supported. The supported processor range is 2 to 32 in a native domain. Higher numbers of processors can be achieved, for example, by using multiple domains or other intermediate connections between node controllers.
[0043] The HPI architecture comprises a definition of a layered protocol architecture, which in some examples includes protocol layers (coherent, incoherent, and optionally memory-based protocols), a routing layer, a link layer, and a physical layer with associated I / O logic. Furthermore, HPI can include extensions related to power management (such as power control units (PCUs)), design for test and debugging (DFT), fault handling, registers, security, and more. Fig. Figure 6 illustrates an embodiment of an exemplary HPI layer protocol stack. In some implementations, at least one of the elements shown in Fig. The six illustrated layers are optional. Each layer deals with its own level of granularity or information quantity (the protocol layer). 605a , b with packages 630 , the link layer 610a , b with flits635 and the physical layer 605a , b with Phits 640 It should be noted that in some implementations, a package may include subflits, a single flit, or multiple flits.
[0044] As a first example, the width of a phit encompasses 640 A one-to-one mapping of link width to bits (e.g., a 20-bit link width corresponds to a phit of 20 bits, etc.). Flits can be larger, such as 184, 192, or 200 bits. It's worth noting that when a phit 640 It is 20 bits wide, and the size of a flit 635 184 bits is a fraction of phits. 640 is necessary to make a flit 635 to send (among other examples, e.g., 9.2 phits at 20 bits to create a 184-bit flit). 635to send, or 9.6 at 20 bits to send a 192-bit flit). It should be noted that widths of the base link on the physical layer can vary. The number of lanes per direction can, for example, include 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, etc. In one embodiment, the link layer 610a HPI is able to embed multiple parts of different transactions into a single flit, and one or more headers (e.g., 1, 2, 3, 4) can be embedded in the flit. In one example, HPI divides the headers into corresponding slots to allow multiple messages in the flit, each destined for a different node.
[0045] The physical layer 605a In one embodiment, layer b can be responsible for the rapid transmission of information on the physical medium (electrical or optical, etc.). The physical link can be between two link layer instances, such as layer 605a and650b , from point to point. The link layer 610a , b can the physical layer 605a The link layer abstracts from the upper layers and provides the capability for the reliable transmission of data (and requests) and the control of flow between two directly connected instances. The link layer can also be responsible for virtualizing the physical channel into multiple virtual channels and message classes. The protocol layer 620a , b is on the link layer 610a , b instructed to process protocol messages before handing them off to the physical layer 605a , b to assign message classes and virtual channels to the corresponding transmission via the physical links. The link layer 610a , b can support multiple messages, such as request, snoop, response, write-back, incoherent data, etc.
[0046] The physical layer605a , b (or PHY) of the HPI can be above the electrical layer (i.e., the electrical connectors that join two components) and below the link layer 610a , b be implemented as in Fig. Figure 6 illustrates this. The physical layer and corresponding logic can reside on each agent and connect the link layers on two agents (A and B) that are separated from each other (e.g., on devices on either side of a link). The local and remote electrical layers are connected by physical media (e.g., wires, conductors, optical, etc.). The physical layer 605aIn one embodiment, b has two main phases: initialization and operation. During initialization, the connection is opaque to the link layer, and signaling can include a combination of timed states and handshake events. During operation, the connection is transparent to the link layer, and signaling occurs at a single speed, with all lanes functioning together as a single link. During the operation phase, the physical layer transports flits from agent A to agent B and from agent B to agent A. The connection, also referred to as the link, abstracts some physical aspects, including media, width, and speed, from the link layers during the exchange of flits and control / status of the current configuration (e.g., width) with the link layer. The initialization phase includes smaller phases, such as querying and configuration.The operational phase also includes smaller phases (e.g., link power management status).
[0047] In one embodiment, the link layer can 610a , b must be implemented in such a way that it provides reliable data transmission between two protocol or routing instances. The link layer can be the physical layer. 605a , b from the protocol layer 620a , b abstract, and it can be responsible for flow control between two protocol agents (A, B) and provide virtual channel services for the protocol layer (message classes) and routing layer (virtual networks). The interface between the protocol layer 620a , b and the link layer 610a, b can typically be at the packet level. In one embodiment, the smallest transmission unit at the link layer is called a flit, which is a specified number of bits, such as 192 bits, or is given another name. The link layer 610a , b is on the physical layer 605a , b instructed to use the transmission unit (phit) of the physical layer 605a , b into the transmission unit (flit) of the link layer 610a , b to frame. Furthermore, the link layer can 610a , b can be logically decomposed into two parts, a sender and a receiver. A sender / receiver pair on one instance can be connected to a sender / receiver pair on another instance. Flow control is often performed on both a flit and a packet basis. Error detection and correction are also potentially performed on a flit-level basis.
[0048] In one embodiment, the routing layer 615a, b provide a flexible and distributed method for routing HPI transactions from an origin to a destination. The scheme is flexible because routing algorithms for multiple topologies can be specified by programmable routing tables at each router (programming is done in one embodiment by firmware, software, or a combination thereof). The routing functionality can be distributed: routing can be performed through a series of routing steps, each defined by a lookup of a table at each origin, intermediate, or destination router. The lookup at an origin can be used to inject an HPI packet into the HPI fabric. The lookup at an intermediate router can be used to route an HPI packet from an input port to an output port. The lookup at a destination port can be used to predetermine the destination HPI protocol agent.It's worth noting that the routing layer can be thin in some implementations, as the routing tables and, consequently, the routing algorithms are not specifically defined by specification. This allows for flexibility and means that a variety of usage models, including flexible architectural platform topologies, are defined by the system implementation. The routing layer... 615a , b is on the link layer 610a, b is instructed to provide the use of up to three (or more) virtual networks (VNs) – in one example, two non-blocking VNs, VN0 and VN1, with different message classes defined in each virtual network. A shared virtual network (VNA, or adaptive virtual network) can be defined in the link layer, but this adaptive network cannot be directly represented in routing concepts because each message class and virtual network may have dedicated resources and guaranteed forwarding progress, among other features and examples.
[0049] In one embodiment, HPI can be a coherence protocol layer. 620aThe protocol includes, but is not limited to, a coherence protocol, which supports agents that cache rows of data from memory. An agent wishing to cache memory data can use the coherence protocol to read the row of data and load it into its cache. An agent wishing to modify a row of data in its cache can use the coherence protocol to acquire ownership of the row before modifying the data. After modifying a row, an agent can keep it in its cache according to protocol rules until it either writes the row back to memory or includes the row in a response to an external request. Finally, an agent can comply with external requests to cancel a row in its cache. The protocol ensures data coherence by prescribing rules that all caching agents can follow.It also provides the means for agents without caches to read and write memory data coherently.
[0050] Two conditions can be enforced to support transactions using the HPI coherence protocol. First, the protocol can, for example, maintain data consistency on a per-address basis among the data in agent caches and between that data and the data in memory. Informally, data consistency can refer to each valid row of data in an agent's cache representing the most recent value of the data, and data sent in a coherence protocol packet representing the most recent value of the data at the time of transmission. If no valid copy of data exists in caches or is transmitted, the protocol can ensure the most recent value of the data in memory. Second, the protocol can provide well-defined commitment points for requests.Commitment points for read operations can indicate when the data is usable; and for write operations, they can indicate when the written data is globally observable and will be loaded by subsequent read operations. The protocol can support these commitment points for both cacheable and uncacheable (UC) requests in the coherent memory space.
[0051] The HPI coherence protocol can also ensure the forwarding progress of coherence requests made by an agent for an address in the coherent memory space. Transactions can then be reliably fulfilled and retired for proper system operation. In some implementations, the HPI coherence protocol may lack awareness of retries to resolve resource allocation conflicts. Therefore, the protocol itself can be defined to avoid cyclic resource dependencies, and implementations can take care in their designs to avoid introducing dependencies that could lead to blocking. Furthermore, the protocol can indicate when designs are capable of providing fair access to protocol resources.
[0052] Logically, the HPI coherence protocol, in one embodiment, can comprise three elements: coherence (or buffer) agents, home agents, and the HPI link fabric, which connects the agents. Coherence agents and home agents can work together to achieve data consistency by exchanging messages over the link fabric. The link layer 610a , b and their associated description can provide the details of the interlink fabric, including how it follows the requirements of the coherence protocol discussed herein. (It should be noted that the division into coherence agents and home agents is for clarity. A design may, among other things, contain multiple agents of both types within a socket or even combine the behavior of agents into a single design unit.)
[0053] In one embodiment, home agents can be configured to protect physical memory. Each home agent can be responsible for a region of the coherent memory space. Regions can be non-overlapping, in that a single address is protected by a single home agent, and together the home agent regions in a system comprise the coherent memory space. For example, each address can be protected by at least one home agent. Therefore, in one embodiment, each address in a coherent memory space of an HPI system can be assigned to exactly one home agent.
[0054] In one embodiment, home agents in the HPI coherence protocol can be responsible for handling requests to the coherent memory area. For read (Rd) requests, home agents can create snoops (Snp), process their responses, send a data response, and send a completion response. For cancel (Inv) requests, home agents can create necessary snoops, process their responses, and send a completion response. For write requests, home agents can write the data to memory and send a completion response.
[0055] Home agents can deploy snoops in the HPI coherence protocol and process snoop responses from coherence agents. Home agents can also process forwarding requests, which are special snoop responses, from coherence agents for conflict resolution. When a home agent receives a forwarding request, it can send a forwarding response to the coherence agent that generated the forwarding request (i.e., the agent that detected a conflicting snooping request). Coherence agents can use the sequence of these forwarding responses and completion responses from the home agent to resolve conflicts.
[0056] A coherence agent can issue supported coherence protocol requests. Requests can be issued to an address in the coherent memory space. Data received for read (Rd) requests can be consistent, with the exception of RdCur. Data for RdCur requests may have been consistent when the data packet was created (although it may have expired during delivery).
[0057] Table 1 presents an exemplary, non-exhaustive list of potential supported requirements: TABLE 1 name semantics Cache status RdCode Requesting a cache row in F or S state. F or S RdData Requesting a cache line in E, F, or S state. F or S RdMigr Requesting a cache line in M, E, F or S state. M and (F or S) RdInv Requesting a cached row in the E state. If a row was previously cached in the M state, the row is written to memory before the E data is delivered. E RdInvOwn Requesting a cache line in M or E state. M RdCur Requesting a non-caching snapshot of a cache line. InvItoE Requesting exclusive ownership of a cache row without receiving data. M or E InvItoM Requesting exclusive ownership of a cache line without receiving any data and with the intention of writing it back shortly afterwards. M or E InvXtoI Emptying a cached line from all caches. The requesting agent must clear the line in its cache before issuing this request. WbMtoI Writing a cache line in M-state back to memory and canceling the line in the cache. M WbMtoS Writing a cache line in M-state back into memory and transferring the line to S-state. M and S WbMtoE Writing a cache line in M-state back into memory and transferring the line to E-state. M and E WbMtoIPtl Writing a cache line in M-state back into memory according to a byte activation mask and transitioning the line to I-state. M WbMtoEPtl Writing a cache line in M-state back into memory according to a byte activation mask, transitioning the line to E-state, and deleting the mask of the line in the cache. M and E EvctCln Notification to a home agent that a cache line in the E state has been canceled in the cache. E WbPushMtoI Sending a line in M-state to a home agent and canceling the line in the cache; the home agent can either write the line back to memory or send it to a local cache agent in M-state. M WbFlush Request that Home perform a write operation to implementation-specific addresses in its memory hierarchy. No data is sent with the request.
[0058] The HPI can support a coherence protocol that utilizes principles of the MESI protocol. Each cache row can be marked (e.g., encoded in the cache row) with one or more supported states. A state of "M" or "modified" indicates that the cache row value has been modified from the value in main memory. A row in the M state exists only in the current cache, and the corresponding cache agent can be instructed to write the modified data back to memory at a future time, for example, before allowing any other read operation of the (no longer valid) main memory state. A write-back can transition the row from the M state to the E state. The state of "E" or "exclusive" indicates that the cache row exists only in the current cache, but that its value is identical to the value in main memory.The cache line in the E state can transition to the S state at any time in response to a read request, or it can be switched to the M state by writing to the line. The "S" state (for shared) indicates that the cache line may be stored in other caches on the machine and contains a value that corresponds to the value in main memory. The line can be discarded (switched to the I state) at any time. The "I" state (for invalid) indicates that a cache line is invalid or unused. HPI can also support other states, such as a shared "F" state (forward), which indicates that the respective shared line value should be forwarded to other caches that should also share the line.
[0059] Table 2 includes exemplary information that can be included in some coherence protocol messages and includes, among other things, snooping, reading, and writing requirements: TABLE 2 Field use cmd Message command (or name or opcode). addr Address of a coherent cache line. destNID Node ID (NID) of a target (home or coherence) agent. reqNID ND of a requesting coherence agent. peerNID NID of a coherence agent that sent the (forwarding request) message. reqTID ID of the resource assigned by the requesting agent for the transaction, also known as RTID (or request transaction identifier). homeTID The ID of the resource assigned by the home agent to process the transaction, also known as the HTID (or home transaction identifier). data A cache row of data. mask Byte mask for qualifying the data.
[0060] Snooping messages can be generated by home agents and routed to coherence agents. A virtual snooping (SNP) channel can be used for snoops, and in one embodiment, they are the only messages that use the virtual SNP channel. Snoops can include the NID of the requesting agent and the RTID it assigned to the request if the snooping results are sent directly to the requesting agent in data. In one embodiment, snoops can also include the HTID assigned by the home agent to process the request. The coherence agent processing the snoop can include the HTID in the snoop response it sends back to the home agent. In some cases, snoops might not include the home agent's NID because it can be inferred from the address the target coherence agent contains when sending its response.Fanout snoops (those with the prefix "SnpF") may not contain a destination NID because the routing layer is responsible for generating the corresponding snoop messages to all partners in the fanout region. An example list of snoop channel messages is shown in Table 3. TABLE 3 command semantics fields SnpCode Snooping to obtain data in the F or S state. cmd, addr, destNID, reqNID, reqTID, homeTID SnpData Snooping to obtain data in the E, F, or S state. SnpMigr Snooping to obtain data in the M, E, F, or S state. SnpInv Snooping to clear the partner's cache, emptying any M copies into memory. SnpInvOw Snooping to obtain data in M or E state. SnpCur Snooping to obtain a non-storable snapshot of a cache line. SnpFCode Snooping to obtain data in F or S state; routing layer handles distribution to all fanout partners. cmd, addr, reqNID, reqTID, homeTID SnpFData Snooping to obtain data in E, F, or S state; routing layer handles distribution to all fanout partners. SnpFMigr Snooping to obtain data in M, E, F or S state; routing layer handles distribution to all fanout partners. SnpFInvOw Snooping to obtain data in M or E state; routing layer handles distribution to all fanout partners. SnpFInv Snooping to clear the partner's cache, emptying any M copies in memory; routing layer handles distribution to all fanout partners. SnpCur Snooping to obtain a non-cacheable snapshot of a cache line; routing layer handles distribution to all fanout partners.
[0061] The HPI can also support non-snooping requests that it can issue to an address, such as those implemented as non-coherent requests. Examples of such requests might include, among other potential examples, a non-snooping read operation to request a read-only row from memory, a non-snooping write operation to write a row to memory, and writing a row to memory according to a mask.
[0062] In one example, four general types of response messages can be defined in the HPI coherence protocol: data, completion, snooping, and forwarding. Certain data messages can transmit an additional completion indication, and certain snooping responses can transmit data. Response messages can use the virtual RSP channel, and the communication fabric can maintain a correct message delivery order among ordered completion and forwarding responses.
[0063] Table 4 includes a list of at least some potential response messages supported by an exemplary HPI coherence protocol: TABLE 4 name semantics fields Data_M Data is in M state. cmd, destNID, reqTID, data Data_E Data is in electronic form. Data_F Data is in F-state. Data_SI Depending on the requirements, data is either in the S state or non-storable "snapshot" data. Data_M Data is in the M state with an ordered completion response. Data_E Data is in the E state with an ordered completion response. Data_F Data is in the F state with an ordered completion response. Data_SI: Depending on the requirements, data can be in the S state or non-storable “snapshot” data with an ordered completion response. CmpU Completion message without any sequence requirements. cmd, destNID, reqTID CmpO Completion response, which should be organized with forwarding responses. RspI Cache is in the I state. cmd, destNID, homeTID RspS Cache is in S-state. RspFwd A copy of the cache line was sent to a requesting agent; the cache state has not changed. RspFwdI A copy of the cache line was sent to a requesting agent; the cache is now in I state. RspFwdS A copy of the cache line was sent to a requesting agent; the cache is entering S state. RspIWb The modified line is implicitly written back to memory; the cache has been switched to I-state. cmd, destNID, homeTID, data RspSWb The modified line is implicitly written back to memory; the cache has been transitioned to S state. RspFwdIWb Modified line is implicitly written back to memory, copy of cache line was sent to a requesting agent, cache was transitioned to I state. RspFwdSWb Modified line is implicitly written back to memory, copy of cache line was sent to a requesting agent, cache was transitioned to S state. RspCnflt The partner has an outstanding request for the same address, requests an orderly forwarding response, and has allocated a resource for forwarding. cmd, destNID, homeTID, peerNID
[0064] In one example, data responses can be directed to a requesting coherence agent. A home agent can send any of the data responses. A coherence agent can only send data responses that do not contain an ordered completion indicator. Furthermore, coherence agents can be restricted to sending data responses only as a result of processing a snooping request. Combined data and completion responses can always be of the ordered completion type and kept ordered by forwarding responses through the communication fabric.
[0065] The HPI coherence protocol can use a general unordered completion message and a coherence-specific ordered completion message. A home agent can send completion responses to coherent requests, and completion responses can typically be directed to a coherence agent. The ordered completion response can be kept ordered by forwarding responses through the communication fabric.
[0066] Sniffer responses can be sent by coherence agents, particularly in response to processing a sniffer request, and are addressed to the home agent handling the sniffer request. The `destNID` is typically a home agent (determined from the address in the sniffer request), and the included `TID` is for the home agent resource allocated to process the request. Sniffer responses with "Wb" in the command are for implicit write-back operations of modified cache rows, and they transmit the cache row data. (Implicit write-back operations can include those performed by a coherence agent in response to a request from another agent, while other requests are made explicitly by the coherence agent using its request resources.)
[0067] Coherence agents can generate a forwarding request when a snooping request conflicts with an unfulfilled request. Forwarding requests are addressed to the home agent that generated the snoop, which is determined from the address in the snoop response. Thus, the `destNID` is a home agent. The forwarding request can also include the TID for the home agent's resource assigned to process the original request and the NID of the coherence agent that generated the forwarding request.
[0068] The HPI coherence protocol supports a single forward response, FwdCnfltO. Home agents can send one forward response for each forward request received, to the coherence agent in the peerNID field of the forward request. Forward responses transmit the cache line address so the coherence agent can compare the message to the forward resource it assigned. A forward response message can transmit the NID of the requesting agent, but in some cases, it might not transmit the TID of the requesting agent. If a coherence agent wants to support cache-to-cache transfers for forward responses, it can store the TID of the requesting agent when processing the sniffer and send a forward request.To support conflict resolution, the communication fabric can maintain an order between the forwarding response and all previously ordered completions sent to the same target coherence agent.
[0069] In some systems, home agent resources are pre-allocated, with RTIDs representing resources within the home agents. Caching agents assign RTIDs from system-configured pools when they generate new coherence requests. Such schemes can limit the number of active requests each caching agent can have for a home agent to the number of RTIDs assigned to it by the system, effectively statically distributing home resources among caching agents. This can lead to inefficient resource allocation, and among other potential problems, properly sizing a home agent to support request throughput for large systems can become impractical. For example, such schemes might enforce RTID pool management on the caching agents.Furthermore, in some systems, a caching agent might not reuse the RTID until the home agent has fully processed the transaction. However, waiting for a home agent to complete all processing can unnecessarily throttle caching agents. Additionally, certain protocol flows can cause caching agents to hold RTIDs beyond the notification of home agent release, further throttling their performance.
[0070] In one implementation, home agents may be allowed to allocate their resources when requests arrive from cache agents. In such cases, home agent resource management can be kept separate from coherence agent logic. In some implementations, home resource management and coherence agent logic may be at least partially intertwined. In some cases, coherence agents may have more pending requests for a home agent than the home agent can handle simultaneously. For example, the HPI may allow requests to form a queue in the communication fabric.The HPI coherence protocol can further be configured to ensure that other messages can proceed around the blocked requests, thus preventing blocking caused by the home agent blocking incoming requests until resources become available, and ensuring that active transactions reach completion.
[0071] In one example, resource management can be supported by allowing an agent receiving a request to allocate resources for its processing, with the agent sending the request and allocating appropriate resources for all responses to the request. The RTID can represent the resource that a home agent allocates for a particular request, which is included in some protocol messages. The RTID (along with the RNID / RTID) in snoop requests and forwarding responses can be used, among other things, to support responses for a home agent as well as data forwarding to a requesting agent. Furthermore, HPI can support an agent's ability to send an orderly completion (CmpO) early, i.e., before the home agent has finished processing the request, when it is determined that it is safe for a requesting agent to reuse its RTID resource.The general handling of snoops with similar RNID / RTID can also be defined by the protocol.
[0072] In an illustrative example, if a tracker state for a particular request is "busy," a directory state can be used to determine when the home agent can send a response. For example, a directory state of "Invalid" might allow sending a response, except for RdCur requests, indicating that there are no pending snoop responses. A directory state of "Unknown" might require that all partner agents be snooped and all their responses collected before a response can be sent. The directory state of "Exclusive" might require that the owner be snooped and all responses collected before a response is sent, or that if the requesting agent is the owner, a response can be sent immediately. The directory state of "Shared" might specify that a canceling request (e.g.,The Home agent has snooped on all partner agents (RdInv* or Inv*) and collected all snoop responses. If a tracker state for a particular request is "WbBuffered," the Home agent can send a data response. If the request's tracker state is "DataSent" (indicating that the Home agent has already sent a data response) or "DataXfrd" (indicating that a partner has transferred a copy of the row), the Home agent can send the completion response.
[0073] In cases like those described above, a home agent can send data and completion responses before all sniffer responses have been collected. The HPI interface enables these "early" responses. By sending data and completions early, the home agent can collect all pending sniffer responses before releasing the resource it allocated for the request. The home agent can also continue blocking further standard requests to the same address until all sniffer responses have been collected, and then release the resource. A home agent sending a response message from a "Busy" or "WbBuffered" state can, among other things, send a sub-action table (which, for example, contains...)(contained in a set of protocol tables, which embody the formal specification of the HPI coherence protocol) use which message to send and a sub-action table to determine how to update the directory state. In some cases, early completion can be performed without pre-assignment by a home node.
[0074] In one embodiment, the HPI coherence protocol can either omit the use of pre-assigned home resources or ordered request channels, or both. In such implementations, certain messages on the HPI-RSP communication channel can be ordered. For example, "ordered completion" and "forward response" messages can be provided, which can be sent from the home agent to the coherence agent. Home agents can send an ordered completion (CmpO or Data_*_CmpO) for all coherent read and cancel requests (as well as other requests, such as NonSnpRd requests, that are not involved in cache coherence conflicts).
[0075] Home agents can send forward responses (FwdCnfltO) to coherence agents, which in turn send forward requests (RspCnflt) to indicate conflicts. A coherence agent can generate a forward request whenever it has an unfinished read or cancel request and detects an incoming snoop request for the same cache line as the request. When the coherence agent receives the forward response, it checks the current state of the unfinished request to determine how to process the original snoop. The home agent can send the forward response in a way that orders it with a completion status (such as CmpO or Data_*_CmpO). The coherence agent can use information contained in the snoop to help it process a forward response. For example, a forward response can include any type of information and not an RTID.The nature of the forwarding response can be inferred from information obtained from the preceding sniffer(s). Furthermore, a coherence agent can block pending snooping requests if all of its forwarding resources are waiting for forwarding responses. In some implementations, each coherence agent may be designed to have at least one forwarding resource.
[0076] In some implementations, communication fabric specifications may reside at the routing layer. In one embodiment, the HPI coherence protocol may have a communication fabric specification specific to the routing layer. The coherence protocol may depend on the routing layer to transform a fanout sniffer (SnpF* opcodes – snooping (SNP) channel messages) into the appropriate snoops for all request partners in the fanout set of coherence agents. The fanout set is a routing layer configuration parameter shared with the protocol layer. In this coherence protocol specification, it is described as a home agent configuration parameter.
[0077] In some of the aforementioned implementations, the HPI coherence protocol can utilize four of the virtual channels: REQ, WB, SNP, and RSP. These virtual channels can be used to reverse dependency cycles and prevent blocking. In one embodiment, each message can be delivered without duplication on any virtual channel, while reordering is handled on the RSP virtual channel.
[0078] In some embodiments, the communication fabric can be configured to maintain an order among certain completion messages and the FwdCnfltO message. The completion messages are the CmpO message and any data message with CmpO appended (Data_*_CmpO). Together, all of these messages constitute the "ordered completion responses." The conceptual precedence between ordered completion responses and the FwdCnfltO message is that an FwdCnfltO does not "overtake" an ordered completion. More specifically, among other potential examples, the communication fabric delivers the ordered completion response before the FwdCnfltO if a home agent sends an ordered completion response followed by an FwdCnfltO message, and both messages are destined for the same coherence agent.
[0079] It goes without saying that, although some examples of protocol flow are revealed herein, the examples described are merely intended to provide an intuitive sense of the protocol and do not necessarily encompass all possible scenarios and behaviors that the protocol may exhibit.
[0080] A conflict can occur when requests for the same cache row address are made by more than one coherence agent at approximately the same time. As a specific example, a conflict can arise when a sniffer for a standard request from one coherence agent arrives at a partner coherence agent with an unfulfilled request for the same address. Since any sniffer can cause a conflict, a single request can have multiple conflicts. Conflict resolution can be a coordinated effort between the home agent, the coherence agents, and the communications fabric. However, the primary responsibility lies with the coherence agents, which detect conflicting snifferers.
[0081] In one embodiment, home agents, coherence agents, and the communication fabric can be configured to contribute to successful conflict resolution. For example, home agents can have pending snoops for only one request per address, such that a home agent for a given address might have pending snoops for only one request. This can help prevent race situations involving two conflicting requests. It can also ensure that a coherence agent does not see another snoop for the same address after detecting but not yet resolving a conflict.
[0082] In another example, when a coherence agent processes a sniffer with an address matching an active standard request, it can allocate a forwarding resource and send a forwarding request to the home agent. A coherence agent with an unfinished standard request that receives a sniffer for the same address can respond with an RspCnflt sniffer response. This response can be a forwarding request to the home agent. Because the message is a request, the coherence agent can allocate a resource to process the response sent by the home agent before sending it. (The coherence protocol allows, in some cases, blocking conflicting snifferers if the coherence agent runs out of forwarding resources.) The coherence agent can store information about the conflicting sniffer to use when processing the forwarding response.It can be guaranteed that after detecting a conflict and until processing the forwarding response, a coherence agent will not see any other sniffer for the same address.
[0083] In some examples, a home agent does not record the snooping response when it receives a forwarding response. Instead, the home agent can send a forwarding response to the conflicting coherence agent. In one example, a forwarding request (RspCnflt) looks like a snooping response, but the home agent does not treat it as such. It does not record the message as a snooping response but instead sends a forwarding response. Specifically, for every forwarding request (RspCnflt) it receives, the home agent sends a forwarding response (FwdCnfltO) to the requesting coherence agent.
[0084] The HPI communication fabric orders forwarding responses and orderly completions between the home agent and the target coherence agent. The fabric can thus be used to distinguish between an early and a late conflict at the conflicting coherence agent. From a system-level perspective, an early conflict occurs when a sniffer finds a request that the home agent has not yet processed, and a late conflict occurs when a sniffer finds a request that the home agent has already processed. From a home agent's perspective, an early conflict occurs when a sniffer finds a request for the currently active request that the home agent has not yet received or begun processing, and a late conflict occurs when a sniffer finds a request that it has already processed.In other words, a late conflict involves a request for which the home agent has already sent a completion response. Therefore, when a home agent receives a forward request for a late conflict, it has already sent the completion response for the conflicting agent's unfinished request. By ordering the forward responses and ordered completion responses from the home agent to the coherence agent, the coherence agent can determine whether the conflict was early or late based on the processing state of its conflicting request.
[0085] When a coherence agent receives a forwarding response, it uses the state of its conflicting request to determine whether the conflict was early or late and when to process the original sniffer. Due to the communication fabric's sequencing, the conflicting request's state indicates whether the conflict was early or late. If the request state indicates that completion has been received, it was a late conflict; otherwise, it was an early conflict. Conversely, if the request state indicates that the request is still awaiting its response(s), it was an early conflict; otherwise, it was a late conflict.The conflict type determines when the sniffer should be processed: From a coherence agent's perspective, an early conflict means the sniffer is for a request that is processed before the agent's conflicting request, and a late conflict means the sniffer is for a request that is processed after the agent's conflicting request. Given this order, for an early conflict, the coherence agent processes the original sniffer immediately; and for a late conflict, the coherence agent waits until the conflicting request has received its data (for read operations) and its processor has had an opportunity to act on the completed request before processing the sniffer. When the conflicting sniffer is processed, the coherence agent generates a sniffer response for the home agent to record.
[0086] All conflicts involving write-back requests can be late conflicts. A late conflict, from the perspective of the coherence agent, occurs when the agent's request is processed before the sniffer's request. By this definition, all conflicts involving write-back requests can be treated as late conflicts, since the write-back operation is processed first. Otherwise, data consistency and coherence could be violated if the home agent processed the request before the write-back to memory occurred. Because all conflicts involving write-back operations are considered late conflicts, coherence agents can be configured to block conflicting snifferers until an unfinished write-back request is completed. Furthermore, write-back operations can also block the processing of redirects. Blocking redirects with an active write-back operation can, among other things,It can also be implemented as a protocol specification to support non-caching storage.
[0087] When a coherence agent receives a request to snoop its cache, it can first check if the coherence protocol allows it, and then process the snoop and generate a response. One or more state tables can be defined within a set of state tables that defines the protocol specification. One or more state tables can specify when a coherence agent can process a snoop and whether it snoops the cache or instead generates a conflicting forward request. In one example, there are two conditions under which a coherence agent processes a snoop. The first condition is when the coherence agent has a REQ request (Rd* or Inv*) for the snoop address and it has an available forward resource. In this case, the coherence agent must generate a forward request (RspCnflt).The second condition is that the coherence agent has no REQ, Wb*, or EvctCln request for the snoop address. A state table can define how a coherence agent should process the snoop according to these conditions. In one example, under other conditions, the coherence agent might block the snoop until either a forwarding resource becomes available (first condition) or the blocking Wb* or EvctCln receives its CmpU response (second condition). It's worth noting that NonSnp* requests might not affect snoop processing, and a coherence agent can ignore NonSnp* entries when determining how to process or block a snoop.
[0088] When generating a forwarding request, a coherence agent can reserve a resource for the forwarding response. The HPI coherence protocol, in one example, might not require a minimum number of forwarding response resources (beyond having at least one) and might allow a coherence agent to block snoopers if it has no forwarding response resources available.
[0089] How a coherence agent handles a sniffer in its cache can depend on the sniffer type and the current cache state. However, for a given sniffer type and cache state, there can be many allowed responses. For example, a coherence agent with a fully modified line receiving a non-conflicting SnpMigr (or processing a forward response after a SnpMigr) can perform any of the following, among other potential examples: downgrade to S, send implicit write-back operation to Home and Data_F to Requester; downgrade to S, send implicit write-back operation to Home; downgrade to I, send Data_M to Requester; downgrade to I, send implicit write-back operation to Home and Data_E to Requester; downgrade to I, send implicit write-back operation to Home.
[0090] The HPI coherence protocol allows a coherence agent to store modified rows with partial masks in its cache. All rows for M copies can require a full or empty mask. The HPI coherence protocol can, for example, restrict implicit write-back of partial rows. A coherence agent that wants to remove a partial M row as a result of a snoop request (or a forward response) can first initiate an explicit write-back operation and block the snoop (or forward) until the explicit write-back operation is complete.
[0091] Storing information for forwarding responses: The HPI coherence protocol allows a coherence agent, in one example, to store forwarding response information separately from the outgoing request buffer (ORB). Separating this information allows the ORB to release ORB resources and RTIDs once all responses have been collected, regardless of which entry was involved in a conflict. State tables can be used to specify which forwarding response information should be stored and under what conditions.
[0092] Forward responses in the HPI coherence protocol can contain the address, the NID of the requesting agent, and the home TID. They do not contain the original sniffer type or the RTID. A coherence agent can store the forward type and RTID if it chooses to use them with the forward response, and it can use the address to compare the incoming forward response with the matching forward entry (and generate the home NID). Storing the forward type can be optional. If no type is stored, the coherence agent can treat a forward response as having a FwdInv type. Similarly, storing the RTID can be optional and should only occur if the coherence agent is to support cache-to-cache transfers when processing forward responses.
[0093] As previously mentioned, coherence agents can generate a forwarding request if a snooping request conflicts with an unfulfilled request. Forwarding requests are addressed to the home agent that generated the snoop, which can be determined from the address in the snoop response. Thus, the destNID can identify a home agent. The forwarding request can also include the TID for the home agent's resource assigned to process the original request and the NID of the coherence agent that generated the forwarding request.
[0094] In one embodiment, a coherence agent can block redirects for write-back requests to maintain data consistency. Coherence agents can also use a write-back request to perform a commit on non-cached (UC) data before processing a redirect, and allow the coherence agent to write back partial cached rows instead of the protocol that supports a partial implicit write-back operation for redirects. In fact, in one embodiment, a coherence agent can be allowed to store modified rows with partial masks in its cache (although M copies are intended to include a full or empty mask).
[0095] In one example, early conflicts can be resolved by a forwarding response that arrives at an unfinished standard request before it has received any response. A corresponding log state table can specify, in an example, that a forwarding response can be processed as long as the standard request entry is still in the ReqSent state. Late conflicts can be resolved by a forwarding response that arrives after the unfinished request has received its completion response. When this occurs, either the request is complete (has already received its data or was an Inv* request) or the entry is in its RcvdCmp state. If the request is still waiting for its data, then the coherence agent must block the forwarding until the data is received (and used).Once the conflicting Rd* or Inv* request has completed, the forward response can be processed as long as the coherence agent has not initiated an explicit write-back of the cached line. It may be permissible for a coherence agent to initiate an explicit write-back operation while it has a forward response (or snoop request) for the same address, thus allowing partial lines (e.g., snoop requests for partially modified lines) or non-cachesable stores to be written correctly to memory.
[0096] Turning towards Fig. Figure 7 illustrates a first example of a model conflict handling scheme. A first cache (or coherence) agent 705 Can a read request for a specific row of data be sent to a home agent? 710 send, which leads to a memory readout 715This occurs shortly after the cache agent requests a read. 705 Another cache agent is present. 720 A request for ownership (RFO) on the same line. The home agent 710 However, the Data_S_CmpO is available before receiving the RFO from the cache agent. 720 to the first cache agent 705 sent. The RFO can cause a snoop (SnpFO) to be sent to the cache agent. 705 (as well as to other cache agents), with the sniffer being initiated by the first cache agent. 705 The cache agent is received before the completion message Data_S_CmpO is received. 705 Upon receiving the Snoop SnpO, it can identify a potential conflict involving the memory line requested in its original read request and the Home agent. 710The home agent notifies the system of the conflict by responding to the SnpO with a forwarding response conflict message (RspCnflt). 710 The cache agent can respond to the forward response RspCnflt by sending a forward response (FwdCnfltO). 705 The shared data ready message `Data_S_CmpO` can then be received and transition from an I state to an S state. The forwarding response `FwdCnfltO` can then be processed by the cache agent. 705 be received, and the cache agent 705 Based on the sniffer SnpFO, which triggered the sending of the forward response RspCnflt, the system can determine how to respond to the forward response message FwdClfltO. In this example, the cache agent can 705 For example, consult a protocol state table to determine a response to the forward response message FwdClftO. In the specific example of Fig. 7 can the cache agent705 transition into an F-state and the S-copy of the data it received from the home agent 710 received in the Data_S_CmpO message, in a Data_F message to the second cache agent 720 send. The first cache agent 705 It can also send a reply message RspFwdS to the home agent. 710 send the Home Agent 710 notified that the first cache agent has shared its copy of the data with the second cache agent.
[0097] In another illustrative example, shown in the simplified flowchart of Fig. As shown in 8, the first cache agent can be used. 705 a Request for Ownership (RFO) of a specific row of memory to the home agent 710 Shortly thereafter, a second agent can send an RdInvOwn message as a request for the same row of memory in an M state to the home agent. 710send. In conjunction with the RFO message from the first cache agent. 705 can the home agent 710 a sniffer (SnpFO) on the second cache agent 720 send, which the second cache agent 720 can identify a potential conflict involving the memory row that is subject to both RFO and RdInvOwn requests. Accordingly, the second cache agent can 720 a forwarding request RspCnflt to the home agent 720 send. The Home Agent 720 responds to the forwarding request from the second cache agent 720 with a forwarding response. The second cache agent 720 Determines a response to the forwarded response based on information contained in the original sniffer SnpFO. In this example, the second cache agent responds. 720 with a sniffer response RspI, which indicates that the second cache agent 720is in an I-state. The home agent 710 It receives the snoop response RspI and determines that it is appropriate to send the Data Ready Exclusive message (Data_E_CmpO) to the first cache agent. 705 to send, which causes the first cache agent to enter an E state. After sending the completion message, the home agent can 710 then begin responding to the RdInvOwn request from the second cache agent, while also responding to a SnpInvO snooping request from the first cache agent. 705 begins. The first cache agent 705 can identify that the sniffer is responding to a request from the second cache agent. 720 This leads to obtaining an exclusive M-state copy of the line. Consequently, the first cache agent goes 705 transition to the M state to send its copy of the line as an M-state copy (with a Data_M message) to the second cache agent. 720to send. Additionally, the first cache agent sends 705 also a response message RspFwdI to indicate that the copy of the line was sent to the second cache agent. 720 was sent and that the first cache agent has entered an I state (and ownership of the copy has passed to the second cache agent). 720 has submitted).
[0098] Next, turning to Fig. Figure 9 shows another simplified flowchart. In this example, a cache agent attempts to... 720 , to request exclusive ownership of a non-caching (UC) row without receiving any data (e.g., via an InvItoE message). A first cache agent 705sends a concurrent message (RdInv) for the cache row in the E state. The HPI coherence protocol can specify that if the requested row was previously cached in the M state, the row is written to memory before E data in response to the RdInv of the first cache agent. 705 will be delivered. The Home Agent 710 can send a completion message (CmpO) to the InvItoE request and a sniffer (SnpInv) based on the RdInv request to the cache agent. 720 send. If the cache agent 720 If the sniffer receives the completion message, the cache agent can 720 Identify that the snoop is affecting the same cache line as its exclusive ownership request, and indicate a conflict by making a forward request (RspCnflt). As in previous examples, the Home agent can 710It must be configured to respond to the forwarding request with a forwarding response (FwdCnfltO). Multiple valid responses for the forwarding response are possible. For example, the cache agent can 720 Initiate an explicit write-back operation (e.g., WbMtoI) and block the sniffer (or forwarding) until the explicit write-back operation is complete (e.g., CmpU), as in the example of Fig. Figure 9 illustrates this. The cache agent can then complete the snoop response (RspI). Among other examples, the home agent can 710 the RdInv request of the first cache agent 705 Process and return a completion message Data_E_CmpO.
[0099] In examples, such as the example of Fig.In implementations where a cache agent receives a sniffer when the agent has an unfinished read or cancel request for the same address and has cached a partially modified line (often referred to as a "buried M"), the HPI coherence protocol allows the agent to either 1) perform an explicit (partial) write-back of the line while the sniffer is blocking, or 2) send a forward request (RspCnflt) to the home agent. If (1) is chosen, the agent processes the sniffer after receiving the completion message for the write-back operation. If (2) is chosen, it is possible for the agent to receive a forward response (FwdCnfltO) while its unfinished read or cancel request is still awaiting responses, and the agent still has a partially modified line.If this is the case, the protocol allows the agent to block forwarding while executing an explicit (partial) write-back of the line. The protocol guarantees that the agent will not receive responses to pending read or cancel requests during the write-back. The previously described mechanism (which allows coherence agents to issue explicit write-back operations and block snooperators and forwards, even if the agent has a pending read or cancel request) is also used to ensure that partial or UC write operations are made to memory before the writer acquires global observability.
[0100] Coherence agents use a two-step process for partial and UC write operations. First, they check if they own the cached row and issue an ownership (cancel) request in the protocol if they do not. Second, they execute the write operation. If they made an ownership request in the first step, it's possible that the request will conflict with other agents' requests for the row, meaning the agent could receive a sniffer while the ownership request is pending. According to coherence protocol guidelines, the agent issues a forward request for the conflicting sniffer. While the agent waits for the forward response, it can receive confirmation of the completion of the ownership request, which grants the agent ownership of the row and allows it to initiate the write-back for the partial or UC write operation.While this is happening, the agent might receive the forward response, which it also needs to process. The coherence agent cannot combine these two activities. Instead, the coherence agent should write back the partial or UC write data separately from processing the forward, performing the write-back operation first. For example, a cache agent, among other examples and features, might use a write-back request to commit UC data before processing the forward and writing back partial cache rows.
[0101] HPI can encompass a wide variety of computing facilities and systems, including mainframes, server systems, personal computers, mobile computers (such as tablets, smartphones, personal digital systems, etc.), smart devices, gaming and entertainment consoles, and set-top boxes, among others. For example, with reference to Fig.Figure 10 shows an embodiment of a block diagram for a computer system that includes a multi-core processor. The processor 1000 This includes any processor or processing unit, such as a microprocessor, embedded processor, digital signal processor (DSP), network processor, handheld processor, application processor, coprocessor, system-on-a-chip (SoC), or any other code-executing device. The processor 1000 In one embodiment, it comprises at least two cores, core 1001 and 1002 , which may include asymmetric cores or symmetric cores (the illustrated embodiment). The processor 1000 However, it can include any number of processing elements, which can be symmetrical or asymmetrical.
[0102] In one embodiment, a processing element refers to hardware or logic to support a software thread. Examples of hardware processing elements include: a thread unit, a thread slot, a process unit, a context, a context unit, a logical processor, a hardware thread, a core, and / or any other element capable of holding a state for a processor, such as an execution state or an architectural state. In other words, in one embodiment, a processing element refers to any hardware that can be independently associated with code, such as a software thread, an operating system, an application, or other code.A physical processor (or processor socket) typically refers to an integrated circuit that potentially includes any number of other processing elements, such as cores or hardware threads.
[0103] A core often refers to logic residing on an integrated circuit capable of maintaining independent architectural states, each associated with at least some dedicated execution resources. In contrast to cores, a hardware thread typically refers to any logic on an integrated circuit capable of maintaining independent architectural states, where these independently maintained architectural states share access to execution resources. As can be seen, the line between the nomenclature of a hardware thread and a core overlaps when certain resources are shared and others are tightly allocated to a specific architectural state.Nevertheless, a core and a hardware thread are often regarded by an operating system as individual logical processors, with the operating system being able to schedule operations on each logical processor individually.
[0104] The physical processor 1000 includes, as in Fig. 10 illustrates two kernels, kernel 1001 and 1002 . Here, core 1001 and 1002 These are considered symmetrical cores, i.e., cores with the same configurations, functional units, and / or logic. In another embodiment, the core comprises 1001 an out-of-order processor core, while the core 1002 It includes an in-order processor core. The cores 1001 and 1002However, they can be individually selected from any kernel type, such as a native kernel, a software-defined kernel, a kernel designed to run a native instruction set architecture (ISA), a kernel designed to run a translated instruction set architecture (ISA), a collaboratively developed kernel, or any other known kernel. In a heterogeneous kernel environment (i.e., with asymmetric kernels), some form of translation, such as binary translation, may be used to dispose of or execute code on one or both kernels. To further elaborate on this, the kernels are... 1001 The illustrated functional units are described in more detail below, since the units in the core 1002 in a similar manner in the embodiment shown.
[0105] As shown, the core comprises1001 two hardware threads 1001a and 1001b , which are also known as hardware thread slots 1001a and 1001b can be described as follows. Therefore, software instances, such as an operating system, see the processor in one implementation. 1000 potentially appearing as four separate processors, i.e., four logical processors or processing elements, capable of simultaneously executing four software threads. As mentioned earlier, a first thread is defined by architectural state registers. 1001a associated with a second thread is architecture state registers. 1001b associated, a third thread can be used with architectural state registers 1002a be associated, and a fourth thread can be associated with architectural state registers. 1002b be associated. Each of the architectural state registers can be used in this process ( 1001a , 1001b , 1002a and 1002b) are referred to as processing elements, thread slots, or thread units, as described previously. As illustrated, the architecture state registers are 1001a in the architectural condition registers 1001b replicated so that individual architectural states / contexts are available for the logical processor 1001a and the logical processor 1001b can be stored. In essence 1001 Other, smaller resources, such as instruction pointers and renaming logic, can also be included in an assignment and renaming block. 1030 , for thread 1001a and 1001b be replicated. Some resources, such as reorder buffers in a reorder / reject unit, may be replicated. 1035 , ILTB 1020Load / store buffers and queues can be shared through partitioning. Other resources, such as internal general-purpose registers, page table-based registers, a subordinate data cache, and data TLBs, can be used independently. 1051 , Execution unit(s) 1040 and parts of an out-of-order unit 1035 , are potentially used entirely jointly.
[0106] The processor 1000 It often includes resources that can be fully shared, shared through partitioning, or dedicated to / by processing element(s). In Fig.Figure 10 is an illustrative representation of a purely exemplary processor, with illustrative logical units / resources of a processor shown. It should be noted that a processor may include or omit any of these functional units, as well as any other known functional unit, logic, or firmware not shown. As illustrated, the core includes 1001 a simplified representative out-of-order (OOO) processor kernel. However, in various embodiments, an in-order processor can be used. The OOO kernel includes a branch target buffer (BTB). 1020 , to predict which branches should be executed / taken, and an instruction-translation buffer (I-TLB). 1020 for storing address translation entries for instructions.
[0107] The core 1001 It also includes a decoding module. 1025, which includes a retrieval unit 1020 It is coupled to decode retrieved elements. In one embodiment, the retrieval logic comprises individual sequencers that are linked to the thread slots. 1001a or 1001b are associated. Usually, the core is 1001 associated with an initial ISA that is on the processor 1000 Executable instructions are defined / specified. Machine code instructions, which are part of the first ISA, often include a portion of the instruction (called an opcode) that references / specifies an instruction or operation to be executed. The decoding logic 1025 It comprises a circuit arrangement that recognizes these instructions by their opcodes and forwards the decoded instructions on the pipeline for processing, as defined by the first ISA. As explained in more detail below, the decoders 1025For example, in one embodiment, logic is designed or configured to recognize specific instructions, such as a transaction instruction. As a result of the recognition by the decoders 1025 takes the architecture or the core 1001 Specific predefined actions are provided to execute tasks associated with the corresponding instruction. It is important to note that all tasks, blocks, operations, and procedures described herein can be executed in response to one or more instructions, some of which may be new or old. It should be noted that the decoders 1026 In one embodiment, the same ISA (or a subset thereof) can be recognized. Alternatively, the decoders can 1026 detect a second ISA (either a subset of the first ISA or another ISA) in a heterogeneous core environment.
[0108] In one example, the assignment and renaming block includes 1030 An allocation tool for reserving resources, such as registry files for storing the results of processing instructions. The threads 1001a and 1001b However, they are potentially capable of out-of-order execution, with the assignment and renaming block 1030 It also reserves other resources, such as a reorder buffer for tracking instruction results. The unit 1030 It may also have a register renamer to rename program / instruction reference registers to other registers accessible to the processor. 1000 to rename internal registers. A reorganization / disposal unit. 1035 It includes components such as the previously mentioned reorder buffers, load buffers and store buffers to support out-of-order execution and subsequent in-order rejection of out-of-order executed instructions.
[0109] A scheduler and execution unit(s) block 1040 In one embodiment, it includes a scheduler unit for distributing instructions / operations to execution units. For example, a floating-point instruction is distributed to a port of an execution unit that has an available floating-point execution unit. Register files associated with the execution units are also included to store the results of processing information instructions. Exemplary execution units include a floating-point execution unit, an integer execution unit, a jump execution unit, a load execution unit, a memory execution unit, and other known execution units.
[0110] A data cache and data translation buffer (D-TLB) 1050 At a lower level, the execution unit(s) are involved. 1040coupled. The data cache is used to store the most recently used / processed elements, such as data operands, which are potentially held in memory coherence states. The D-TLB is used to store the most recent translations from virtual / linear to physical addresses. As a specific example, a processor might include a page table structure to partition physical memory into a plurality of virtual pages.
[0111] Here, the kernels divide. 1001 and 1002 access to higher-level or more distant caches, such as a second-level cache connected via an on-chip interface 1010is associated. It should be noted that "higher level" or "further away" refers to cache levels that are higher up or further away from the execution unit(s). In one embodiment, the higher-level cache is a data cache at a final level—a final cache in the memory hierarchy on the processor. 1000 —such as a second- or third-level data cache. However, higher-level caches are not limited to this, as they can be associated with or include an instruction cache. A trace cache—a type of instruction cache—can instead be located after the decoder. 1025 They are coupled to store the last decoded traces. Here, an instruction potentially refers to a macro instruction (i.e., a general instruction recognized by the decoder), which can decode into a number of micro instructions (micro-operations).
[0112] In the configuration shown, the processor includes 1000 in addition, an on-chip interface module 1010 Historically, a memory controller, which will be described in more detail below, was located outside the processor in a computer system. 1000 Included. In this scenario, the chip's internal interface serves this purpose. 1010 to communicate with devices outside the processor 1000 , such as system memory 1075 , a chipset (which often includes a memory controller hub for connecting to the memory 1075 and includes an I / O controller hub for connecting peripherals), a memory controller hub, a northbridge, or another integrated circuit. And in this scenario, the bus can 1005any known intermediate connection, such as a multipoint interconnect bus, a point-to-point interconnect, a serial interconnect, a parallel bus, a coherent (e.g., cache-coherent) bus, a layered protocol architecture, a differential bus, and a GTL bus.
[0113] The storage 1075 can the processor 1000 be permanently assigned or shared with other facilities in a system. Common examples of storage types: 1075 These include DRAM, SRAM, non-volatile memory (NV memory), and other well-known memory devices. It should be noted that the device 1080may include a graphics accelerator, a graphics processor or a graphics card coupled with a memory controller hub, a data storage device coupled with an I / O controller hub, a wireless transceiver, a flash device, an audio controller, a network controller or other known devices.
[0114] However, for some time now, each of these facilities can be integrated into a processor. 1000 They are integrated because more logic and facilities are integrated into a single chip, such as a SoC. For example, in one embodiment, a memory controller hub is on the same package and / or chip as the processor. 1000 . This includes a part of the core (a core-internal part) 1010 one or more controllers for establishing a connection with other devices, such as the storage 1075 or a graphics setup 1080, via an interface. The configuration that includes an intermediate connection and controllers for establishing a connection to such facilities via an interface is often referred to as an in-core (or out-of-core) configuration. As an example, the on-chip interface includes 1010 a ring interconnect for in-chip communication and a serial point-to-point high-speed link 1005 for communication outside the chip. However, the SOC environment can also include other components, such as the network interface, coprocessors, and memory. 1075 , the graphics processor 1080 and other known computer features / interfaces, integrated into a single chip or integrated circuit to provide a small form factor with high functionality and low power consumption.
[0115] In one embodiment, the processor 1000to execute a compiler, optimization and / or translator code 1077 for compiling, translating and / or optimizing application code 1076capable of supporting the devices and procedures described herein or of establishing a connection to them via an interface. A compiler often comprises a program or a set of programs for translating source text / code into target text / code. Typically, the compilation of program / application code with a compiler occurs in multiple stages and passes to convert a higher-level programming language into machine-oriented machine or assembly language code. However, single-pass compilers can still be used for simple compilations. A compiler can employ any known compilation technique and perform all known compiler operations, such as lexical analysis, preprocessing, syntactic analysis, semantic analysis, code generation, code conversion, and code optimization.
[0116] Larger compilers often comprise several phases, but these phases are most commonly contained in two main stages: (1) the frontend, where syntactic processing, semantic processing, and some transformation / optimization generally take place, and (2) the backend, where analysis, transformations, optimizations, and code generation generally occur. Some compilers refer to a middle ground, which illustrates the blurring of the boundaries between a compiler's frontend and backend. As a result, references to insertion, association, generation, or other compiler operations can occur in any of the aforementioned phases or passes, as well as in any other known phases or passes of a compiler. As an illustrative example, a compiler potentially inserts operations, calls, functions, etc.Dynamic compilation involves one or more compilation phases, such as inserting calls / operations in a frontend compilation phase and subsequently converting these calls / operations into lower-level code during a conversion phase. It's worth noting that during dynamic compilation, compiler code or dynamic optimization code can insert such operations / calls and optimize the code for execution at runtime. As a specific example, binary code (already compiled code) can be dynamically optimized at runtime. The program code can include the dynamic optimization code, the binary code, or a combination thereof.
[0117] Similar to a compiler, a translator, such as a binary translator, translates code either statically or dynamically to optimize and / or translate code. Therefore, references to code execution, application code, program code, or other software environment may refer to: (1) either dynamic or static execution of compiler program(s), optimization code optimizers, or translators to compile program code, maintain software structures, perform other operations, optimize code, or translate code; (2) execution of main program code that includes operations / calls, such as application code that has been optimized / compiled; (3) execution of other program code, such as libraries, associated with the main program code to maintain software structures, execute other software-related operations, or optimize code; or (4) a combination thereof.
[0118] Although the present invention has been described in relation to a limited number of embodiments, numerous modifications and variations thereof are apparent to those skilled in the art. It is intended that the appended claims encompass all such modifications and variations which fall within the true essence and scope of protection of this present invention.
[0119] A design can go through various stages, from creation and simulation to manufacturing. Data representing a design can depict it in a number of ways. First, as is common in simulations, the hardware can be represented using a hardware description language or another functional description language. Additionally, a circuit-level model with logic and / or transistor gates can be created at some stages of the design process. Furthermore, most designs reach a level of data at some stage that represents the physical arrangement of various components within the hardware model.If conventional semiconductor fabrication techniques are used, the data representing the hardware model can be the data specifying the presence or absence of various features on different mask layers for masks used to create the integrated circuit. In any representation of the design, the data can be stored on any form of machine-readable medium. Working memory or magnetic or optical storage, such as a disk, can be the machine-readable medium for storing information transmitted via optical or electrical waves that are modulated or otherwise generated to send such information.When an electrical carrier wave displaying or transmitting the code or design is sent, a new copy is created insofar as copying, buffering, or retransmission of the electrical signal is performed. Accordingly, a communications provider or network operator can, by implementing techniques of embodiments of the present invention, at least temporarily store an item, such as information encoded in a carrier wave, on a physical, machine-readable medium.
[0120] A module, as used herein, refers to any combination of hardware, software, and / or firmware. For example, a module comprises hardware, such as a microcontroller, associated with a non-transitory medium for storing code designed to be executed by the microcontroller. Therefore, in one embodiment, the use of a module refers to the hardware specifically configured to recognize and / or execute the code intended to be stored on a non-transitory medium. Furthermore, in another embodiment, the use of a module refers to the non-transitory medium that comprises the code specifically designed to be executed by the microcontroller to perform predetermined operations.Consequently, the term "module" (in this example) can, in yet another embodiment, refer to the combination of the microcontroller and the non-transitory medium. Module boundaries, illustrated as separate, typically vary frequently and potentially overlap. For example, a first and a second module may share hardware, software, firmware, or a combination thereof, while potentially retaining some independent hardware, software, or firmware. In one embodiment, the term "logic" includes hardware such as transistors, registers, or other hardware such as programmable logic devices.
[0121] The use of the phrase “configured to” refers, in one embodiment, to the design, assembly, manufacture, offering for sale, import, and / or engineering of a device, hardware, logic, or element to perform an intended or specific task. In this example, a device or element that is not in operation is nevertheless “configured to” perform an intended task if it is designed, coupled, and / or connected to perform that intended task. As a purely illustrative example, a logic gate can provide a 0 or a 1 during operation. However, a logic gate that is “configured to” provide an enable signal for a clock does not include every potential logic gate that can provide a 1 or a 0.Instead, a logic gate is one that is coupled such that the output of 1 or 0 during operation is to release the clock signal. It should be noted again that the term "configured to / in order to / that" does not require operation, but instead emphasizes the latent state of a device, piece of hardware, and / or element, where the device, hardware, and / or element in the latent state is designed to perform a specific task when the device, hardware, and / or element is operational.
[0122] Furthermore, in one embodiment, the terms "to," "capable of," and / or "designed to" refer to a device, logic, hardware, and / or element that is designed to enable the use of the device, logic, hardware, and / or element in a specified manner. It should be noted that, as previously mentioned, the use of "to," "capable of," or "designed to" refers to the latent state of a device, logic, hardware, and / or element, wherein the device, logic, hardware, and / or element is not in operation but is designed to enable the use of a device in a specified manner.
[0123] A value, as used herein, encompasses any known representation of a number, a state, a logical state, or a logical binary state. The use of logical levels, logical values, or logic values often also refers to 1s and 0s, which simply represent logical binary states. For example, a 1 refers to the logical state H, and 0 refers to the logical state L. In one embodiment, a memory cell, such as a transistor or flash cell, may be capable of holding a single logical value or multiple logical values. However, other representations of values have been used in computer systems. For example, the decimal number ten can also be represented as a binary value of 1010 and a hexadecimal letter A. Therefore, a value encompasses any representation of information that can be held in a computer system.
[0124] Furthermore, states can be represented by values or parts of values. For example, a first value, such as a logical one, can represent a default or initial state, while a second value, such as a logical zero, can represent a non-default value. Additionally, in one embodiment, the terms "reset" and "set" refer to a default and an updated value, respectively. For instance, a default value potentially includes a logical H value, i.e., reset, while an updated value potentially includes a logical L value, i.e., set. It's worth noting that any combination of values can be used to represent any number of states.
[0125] The embodiments of methods, hardware, software, firmware, or code described above can be implemented by means of instructions or code stored on a machine-accessible, machine-readable, computer-accessible, or computer-readable medium that can be executed by a processing element. A non-transitory machine-accessible / readable medium includes any mechanism that provides (i.e., stores and / or transmits) information in a form that can be read by a machine, such as a computer or electronic system. A non-transitory machine-accessible medium includes, for example, random-accessible memory (RAM).random-access memory), such as static RAM (SRAM) or dynamic RAM (DRAM); ROM; a magnetic or optical storage medium; flash memory devices; electrical storage devices; optical storage devices; acoustic storage devices; any other form of storage device, distinct from non-transient media, capable of receiving information from it, for holding information received from transitory (propagated) signals (e.g., carrier waves, infrared signals, digital signals), etc.
[0126] Instructions used to program logic for carrying out executions of the invention can be stored in a memory within the system, such as DRAM, cache, flash memory, or other storage. Furthermore, the instructions can be distributed over a network or through other machine-readable media. Accordingly, a machine-readable medium can be any mechanism for storing or transmitting information in a form that can be interpreted by a machine (e.g., a computer).Computer-readable media include, but are not limited to, floppy disks, optical discs, CD-ROMs and magneto-optical disks, read-only memory (ROMs), random-access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards, flash memory, or any physical machine-readable memory used in the transmission of information over the internet by electrical, optical, acoustic, or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.). Accordingly, computer-readable media includes all types of physical machine-readable media suitable for storing or transmitting electronic instructions or information in a form that can be read by a machine (e.g., a computer).
[0127] The following examples relate to embodiments according to this specification. One or more embodiments may provide a device, a system, a machine-readable memory, a machine-readable medium, and a method for sending a coherence protocol message corresponding to a specific cache line, identifying a potential conflict involving the specific cache line, and sending a forwarding request to a home agent to identify the potential conflict.
[0128] Furthermore, one or more examples can provide a receipt of a sniffer that corresponds to the specific cache line.
[0129] One or more examples can further provide an identification that the sniffer is being received while a request is pending, and the potential conflict is identified based on the identification that the sniffer is being received while the request is pending.
[0130] One or more examples can also provide a receipt of a forwarding response from the home agent based on the forwarding request.
[0131] Furthermore, one or more examples can provide a way to determine an answer to the forwarding response based at least partially on attributes of the sniffer.
[0132] In at least one example, the sniffer corresponds to another coherence protocol message by another agent, which corresponds to the specified cache line, and the attributes of the sniffer include an identification of the other agent, an identification of a command contained in the other coherence protocol message, and a transaction identifier of the other coherence protocol message.
[0133] In at least one example, the response to the forwarding response includes a snooping response, and the protocol layer logic further serves to send the snooping response to the home agent after receiving a completion of the coherence protocol message.
[0134] In at least one example, the response involves performing a write-back to memory before sending a snoop response to the home agent.
[0135] In at least one example, the specific cache line on the agent is partially modified.
[0136] One or more examples may also provide a receipt of completion after receiving the forwarding response.
[0137] Furthermore, one or more examples may provide for receiving a completion notice before receiving the forwarding response.
[0138] One or more examples may also provide an allocation of a resource for responding to the request.
[0139] One or more examples may also provide an assignment of a forwarding resource for a forwarding response to the forwarding request.
[0140] One or more embodiments may provide a device, a system, a machine-readable memory, a machine-readable medium, and a method for receiving a first coherence protocol request from a first cache agent, sending a snooping request to a second cache agent, wherein the snooping request corresponds to the first coherence protocol request, receiving a forwarding request from the second cache agent that corresponds to the snooping request, wherein the forwarding request identifies a potential conflict with the first coherence protocol request, and sending a forwarding response to the second cache agent in response to the forwarding request.
[0141] One or more examples may further provide receiving another coherence protocol request from the second cache agent, where the first coherence protocol request and the other coherence protocol request each concern a common cache row.
[0142] In at least one example, the other coherence protocol request is received by the agent before the first coherence protocol request, and the protocol layer logic further serves to process the other coherence protocol request and send back a completion message for the other coherence protocol request.
[0143] One or more examples may further provide receiving a response from the second agent to the forwarding response and generating a completion for the first coherence protocol request upon receiving the response to the forwarding response.
[0144] In at least one example, the agent includes a home agent.
[0145] One or more examples can further provide a system with an interlink fabric, a home agent for handling requests for a coherent memory area, and a cache agent that is communicatively coupled to the home agent via the interlink fabric. The interlink fabric can ensure a sequence of responses to the sniffer and completion for the other coherence protocol request. The home agent can have a set of resources, and this set of resources is not pre-allocated to cache agents in the system.
[0146] One or more examples may further provide an agent with a layered protocol stack comprising a protocol layer, where the protocol layer is used to initiate an allocation of resources without intervention from the home agent, to process the first request in response to the agent receiving the first request, and to initiate an allocation of resources without intervention from the home agent to process responses to a second request in response to the agent sending the second request.
[0147] In at least one example, the allocation of resources includes one of HTID, RNID, RTID or a combination thereof.
[0148] In at least one example, the allocation of resources includes the allocation of resources to process snooping requests and forwarding requests.
[0149] One or more examples may further provide an agent with a layered protocol stack that includes a protocol layer, the protocol layer being used to mimic the use of an ordered response channel to perform conflict resolution.
[0150] One or more examples may further provide a coherence agent with a layered protocol stack that includes a protocol layer, the protocol layer being used to block a forward for a write-back request in order to maintain data consistency.
[0151] In at least one example, the protocol layer is used to initiate a write-back request in order to perform a commit on non-caching data before processing the forwarding.
[0152] In at least one example, the protocol layer also serves to support explicit writing back of partial cache lines.
[0153] Throughout the specification, references to “a particular embodiment” or “any embodiment” mean that a specific feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Therefore, the occurrence of the expressions “in a particular embodiment” or “in any embodiment” at various points in the specification does not necessarily always refer to the same embodiment. Furthermore, the respective features, structures, or characteristics may be combined appropriately in one or more embodiments.
[0154] The foregoing specification provides a detailed description with reference to specific exemplary embodiments. However, it is obvious that various modifications and changes can be made to it without deviating from the broader nature and scope of the invention as set forth in the appended claims. Accordingly, the specification and the drawings should be considered in an illustrative rather than a restrictive sense. Furthermore, the foregoing use of "embodiment" and other exemplary expressions does not necessarily refer to the same embodiment or example, but may refer to other and different embodiments, as well as potentially to the same embodiment.
Claims
[1] Device comprising: an agent that includes a protocol layer logic for: Sending a coherence protocol message that corresponds to a specific cache line; Identifying a potential conflict involving the specific cache line; and Sending a forwarding request to a home agent to identify the potential conflict. [2] Device according to claim 1, wherein the protocol layer logic further serves to receive a sniffer which corresponds to the specified cache line. [3] Device according to claim 2, wherein the message includes a request which incorporates the specific cache line, the protocol layer logic serves to identify that the sniffer is being received while the request is pending, and the potential conflict is identified based on the identification that the sniffer is being received while the request is pending. [4] Device according to claim 2, wherein the protocol layer logic further serves to receive a forwarding response from the home agent based on the forwarding request. [5] Device according to claim 4, wherein the protocol layer logic further serves to determine a response to the forwarding response at least partly based on attributes of the sniffer. [6] Device according to claim 5, wherein the sniffer corresponds to another coherence protocol message by another agent which corresponds to the specified cache line, and the attributes of the sniffer comprise an identification of the other agent, an identification of a command contained in the other coherence protocol message, and a transaction identifier of the other coherence protocol message. [7] Device according to claim 5, wherein the response to the forwarding response comprises a snooping response, and the protocol layer logic further serves to send the snooping response to the home agent after receiving a completion of the coherence protocol message. [8] Device according to claim 5, wherein the response comprises performing a write-back to memory prior to sending a snoop response to the home agent. [9] Device according to claim 8, wherein the specific cache line on the agent is partially modified. [10] Device according to claim 4, wherein the protocol layer logic further serves to receive a completion after receiving the forwarding response. [11] Device according to claim 4, wherein the protocol layer logic further serves to receive a completion before receiving the forwarding response. [12] Device according to claim 1, wherein the protocol layer logic further serves to allocate a resource for responses to the request. [13] Device according to claim 12, wherein the protocol layer logic further serves to assign a forwarding resource for forwarding responses to the forwarding request. [14] Device comprising: an agent that includes a protocol layer logic for: Receiving an initial coherence protocol request from an initial cache agent; Sending a sniffing request to a second cache agent, where the sniffing request corresponds to the first coherence protocol request; Receiving a forwarding request from the second cache agent that matches the snooping request, where the forwarding request identifies a potential conflict with the first coherence protocol request; and Sending a forwarding response to the second cache agent in response to the forwarding request. [15] Device according to claim 14, wherein the protocol layer logic further serves to receive another coherence protocol request from the second cache agent, wherein the first coherence protocol request and the other coherence protocol request each relate to a common cache row. [16] Device according to claim 15, wherein the other coherence protocol request is received by the agent prior to the first coherence protocol request, and the protocol layer logic further serves to process the other coherence protocol request and return a completion message for the other coherence protocol request. [17] Device according to claim 15, wherein the protocol layer logic further serves to receive a response from the second agent to the forwarding response and to generate a completion for the first coherence protocol request upon receipt of the response to the forwarding response. [18] Device according to claim 14, wherein the agent comprises a home agent. [19] Device according to claim 14, wherein the protocol layer logic further serves to allocate a resource for responses to the request and to allocate a forwarding resource for the forwarding response based on the forwarding request received. [20] Procedures, including: Receiving an initial coherence protocol request from an initial cache agent; Sending a sniffing request to a second cache agent, where the sniffing request corresponds to the first coherence protocol request; Receiving a forwarding request from the second cache agent that matches the snooping request, where the forwarding request identifies a potential conflict with the first coherence protocol request; and Sending a forwarding response to the second cache agent in response to the forwarding request. [21] Method according to claim 20, further comprising receiving another coherence protocol request from the second cache agent, wherein the first coherence protocol request and the other coherence protocol request each relate to a common cache row. [22] The method of claim 21, wherein the other coherence protocol request is received before the first coherence protocol request, the method further comprising: Processing the other coherence protocol requirement; and Sending back a completion message for the other coherence protocol request. [23] The method of claim 21, further comprising: Receiving a response from the second agent to the forwarding response; and Generating a completion for the first coherence protocol request in response to receiving the reply to the forwarding response. [24] System, encompassing: an intermediate compound fabric; a home agent that handles requests for a coherent memory area; a cache agent that is communicatively coupled to the home agent via the intermediate link fabric, where the cache agent serves to: Sending a coherence protocol message corresponding to a specific cache line; and Receiving a sniffer from a home agent that matches the specified cache line, based at least in part on the sniffer; Identifying a potential conflict involving the specific cache line; and Sending a forwarding request to a home agent to identify the potential conflict. [25] System according to claim 24, wherein the home agent serves to: Receiving another coherence protocol request from another cache agent; Sending the sniffer, where the sniffer meets the coherence protocol requirement; Receiving the forwarding request; Sending a forwarding response to the cache agent in response to the forwarding request; Generating a completion for the coherence protocol request; and Generating a completion for the other coherence protocol requirement. [26] System according to claim 25, wherein the cache agent further serves to determine a response to the forwarding response, wherein the response to the forwarding response comprises sending a response to the sniffer to the home agent. [27] System according to claim 26, wherein the intermediate interconnect fabric serves to ensure the sequence of response to the sniffer and completion for the other coherence protocol request. [28] System according to claim 24, wherein the home agent has a set of resources, and the set of resources is not pre-assigned to cache agents in the system. [29] Device comprising: an agent with a layered protocol stack comprising a protocol layer, where the protocol layer is used to initiate a resource allocation without intervention from the home agent to process a first request in response to the agent receiving the first request, and to initiate a resource allocation without intervention from the home agent to process responses to a second request in response to the agent sending the second request. [30] Device according to claim 29, wherein the allocation of resources comprises one of HTID, RNID, RTID or a combination thereof. [31] Device according to claim 29, wherein the allocation of resources includes resources for processing snooping requests and forwarding requests. [32] Device comprising: an agent with a layered protocol stack that includes a protocol layer, the protocol layer being used to mimic the use of an ordered response channel for conflict resolution. [33] Device comprising: a coherence agent with a layered protocol stack comprising a protocol layer, where the protocol layer is for blocking a forward for a write-back request in order to maintain data consistency. [34] Device according to claim 33, wherein the protocol layer serves to initiate a write-back request in order to perform a commit for non-caching data prior to processing the forwarding. [35] Device according to claim 33, wherein the protocol layer is further for supporting explicit write-back of partial cache lines.
Citation Information
Patent Citations
System and method for a 3-hop cache coherency protocol
US20080162661A1
Re-snoop for conflict resolution in a cache coherency protocol
US7721050B2