Non-posted write transactions to a computer bus

By introducing multi-core processors and accelerators in the computing system, using data streaming accelerators and shared work queue technologies, the performance and power balance problems caused by the complexity of interconnect architecture in the existing technology are solved, and high-performance and low-energy computing effects are achieved.

CN112882963BActive Publication Date: 2025-05-13INTEL CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110189872.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-10
Filing Date
2020-03-18
Publication Date
2025-05-13
Estimated Expiration
2040-03-18

AI Technical Summary

Technical Problem

The complexity of interconnect architectures increases when existing computing systems deal with high performance and low energy consumption requirements, making performance and power balance difficult to achieve.

Method used

By introducing multi-core processors and accelerators in computing systems, using data streaming accelerators (DSAs) and shared work queues (SWQs) technologies, high-performance and low-energy computing are achieved.

Benefits of technology

Improves the performance and energy efficiency of the computing system, simplifies access to accelerator and host memory, reduces cache consistency overhead, and achieves higher computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112882963B_ABST
    Figure CN112882963B_ABST
Patent Text Reader

Abstract

The system and device may include a controller and a command queue to buffer incoming write requests entering the device. The controller may receive a non-posted write request (e.g., a deferred memory write (DMWr) request) in a transaction layer packet (TLP) from a client to the command queue over a link; determine that the command queue can accept the DMWr request; identify a successful completion (SC) message from the TLP, the successful completion message indicating that the DMWr request was accepted into the command queue; and send the SC message to the client over the link indicating that the DMWr request has been accepted into the command queue. The controller may receive a second DMWr request in a second TLP; determine that the command queue is full; and send a memory request retry status (MRS) message to be sent to the client in response to the command queue being full.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application with application number 202010193031.9 filed on March 18, 2020.

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This disclosure claims the benefit of U.S. Provisional Patent Application Serial No. 62 / 836,288, filed on April 19, 2019, pursuant to 35 U.S.C. §119(e), the entire contents of which are incorporated herein by reference. Background Art

[0004] Central processing units (CPUs) perform general computing tasks, such as running application software and operating systems. Graphics processors, image processors, digital signal processors, and fixed-function accelerators handle specialized computing tasks such as graphics and image processing. In today's heterogeneous machines, each type of processor is programmed differently. The era of big data processing requires higher performance at lower energy consumption than today's general-purpose processors. Accelerators (e.g., customized fixed-function units or customized programmable units) are helping to meet these demands. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Figure 1 is a schematic diagram illustrating an embodiment of a block diagram of a computing system including a multi-core processor according to an embodiment of the present disclosure.

[0006] Figure 2 is a schematic diagram of an example accelerator device according to an embodiment of the present disclosure.

[0007] Figure 3 is a schematic diagram of an example computer system including an accelerator and one or more computer processor chips coupled to the processor via a multi-protocol link.

[0008] Figure 4 is a schematic diagram of an example work queue implementation according to an embodiment of the present disclosure.

[0009] Figure 5 is a schematic diagram of an example data streaming accelerator (DSA) device including a plurality of work queues that receive descriptors submitted through an I / O fabric interface.

[0010] FIG. 6A to FIG. 6B is a schematic diagram illustrating an example shared work queue implementation according to an embodiment of the present disclosure.

[0011] Figure 7A-7D is a diagram illustrating an example Delayed Memory Write (DMWr) request and response message flow according to an embodiment of the present disclosure.

[0012] Figure 8 is a process flow diagram for performing scalable work submission according to an embodiment of the present disclosure.

[0013] Fig.9A is a schematic diagram of a 64-bit DMWr group definition according to an embodiment of the present disclosure.

[0014] Fig. 9B is a schematic diagram of a 32-bit DMWr group definition according to an embodiment of the present disclosure.

[0015] Fig.10 An embodiment of a computing system including an interconnect architecture is shown.

[0016] Fig.11 An embodiment of an interconnect architecture including a layered stack is shown.

[0017] Fig.12 Embodiments of requests or packets to be generated or received within an interconnect fabric are shown.

[0018] Fig.13 An embodiment of a transmitter and receiver pair for an interconnect architecture is shown.

[0019] Fig.14 Another embodiment of a block diagram of a computing system including a processor is shown.

[0020] Fig.15 An embodiment of a block of a computing system including multiple processor sockets is shown. DETAILED DESCRIPTION

[0021] In the following description, many specific details are set forth, such as specific types of processors and system configurations, specific hardware structures, specific architectural and microarchitectural details, specific register configurations, specific instruction types, specific system components, specific measurements / heights, specific processor pipeline stages and operations, etc., in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that these specific details need not be employed to practice the present disclosure. In other instances, well-known components or methods, such as specific and alternative processor architectures, specific logic circuits / codes for the described algorithms, specific firmware codes, specific interconnect operations, specific logic configurations, specific manufacturing techniques and materials, specific compiler implementations, expressions of specific algorithms in code, specific power-down and gating techniques / logic, and other specific operational details of computer systems are not described in detail in order to avoid unnecessarily obscuring the present disclosure.

[0022] Although the following embodiments may be described with reference to energy conservation and energy efficiency in a particular integrated circuit, such as in a computing platform or microprocessor, other embodiments may be applicable to other types of integrated circuits and logic devices. Similar techniques and teachings of the embodiments described herein may be applied to other types of circuits or semiconductor devices that may also benefit from better energy efficiency and energy conservation. For example, the disclosed embodiments are not limited to desktop computer systems or Ultrabooks. TM . and may also be used in other devices, such as handheld devices, tablet computers, other thin and light notebook computers, system-on-chip (SOC) devices, and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and handheld PCs. Embedded applications typically include microcontrollers, digital signal processors (DSPs), system-on-chips, network computers (NetPCs), set-top boxes, network hubs, wide area network (WAN) switches, or any other system that can perform the functions and operations described below. In addition, the apparatus, methods, and systems described herein are not limited to physical computing devices, but may also involve software optimization for energy conservation and efficiency. As will be apparent in the description below, embodiments of the methods, apparatus, and systems described herein (whether with reference to hardware, firmware, software, or a combination thereof) are critical to a "green technology" future that balances performance.

[0023] As computing systems evolve, the components therein become increasingly complex. As a result, the complexity of the interconnect architecture used to couple and communicate between components is also increasing to ensure that bandwidth requirements are met for optimal component operation. In addition, different market segments require different aspects of the interconnect architecture to meet the needs of the market. For example, servers require higher performance, while the mobile ecosystem can sometimes sacrifice overall performance to save power. However, the sole purpose of most structures is to provide the highest possible performance and maximize power savings. Below, a number of interconnects are discussed that will potentially benefit from aspects of the present disclosure described herein.

[0024] refer to Figure 1 , depicts an embodiment of a block diagram of a computing system including a multi-core processor. Processor 100 includes any processor or processing device, such as a microprocessor, an embedded processor, a digital signal processor (DSP), a network processor, a handheld processor, an application processor, a coprocessor, a system on a chip (SOC), or other device that executes code. In one embodiment, processor 100 includes at least two cores - cores 101 and 102, which may include asymmetric cores or symmetric cores (embodiment shown). However, processor 100 may include any number of processing elements that may be symmetric or asymmetric.

[0025] In one embodiment, a processing element refers to hardware or logic that supports a software thread. Examples of hardware processing elements include: thread units, thread slots, threads, processing units, contexts, context units, logical processors, hardware threads, cores, and / or any other element capable of maintaining a processor state (e.g., an execution state or an architectural state). In other words, in one embodiment, a processing element refers to any hardware that can be independently associated with code such as a software thread, an operating system, an application, or other code. A physical processor (or processor socket) typically refers to an integrated circuit that may include any number of other processing elements, such as cores or hardware threads.

[0026] A core generally refers to logic located on an integrated circuit capable of maintaining an independent architectural state, where each independently maintained architectural state is associated with at least some dedicated execution resources. In contrast to a core, a hardware thread generally refers to any logic located on an integrated circuit capable of maintaining an independent architectural state, where the independently maintained architectural states share access to execution resources. It can be seen that the line between the nomenclature of a hardware thread and a core overlaps when some resources are shared and other resources are dedicated to an architectural state. However, operating systems often view cores and hardware threads as separate logical processors, where the operating system can schedule operations on each logical processor separately.

[0027] like Figure 1 As shown, the physical processor 100 includes two cores, namely cores 101 and 102. Here, cores 101 and 102 are considered to be symmetric cores, i.e., cores with the same configuration, functional units and / or logic. In another embodiment, core 101 includes a disordered processor core, and core 102 includes an ordered processor core. However, cores 101 and 102 can be selected from any type of core, such as a native core, a software-managed core, a core suitable for executing a native instruction set architecture (ISA), a core suitable for executing a converted instruction set architecture (ISA), a co-designed core, or other known cores. In a heterogeneous core environment (i.e., an asymmetric core), some form of conversion (e.g., binary conversion) can be used to schedule or execute code on one or both cores. For further discussion, the functional units shown in core 101 are described in further detail below, because the units in core 102 operate in a similar manner in the illustrated embodiment.

[0028] As shown, core 101 includes two hardware threads 101a and 101b, which may also be referred to as hardware thread slots 101a and 101b. Therefore, in one embodiment, a software entity such as an operating system potentially treats processor 100 as four separate processors, i.e., four logical processors or processing elements capable of executing four software threads simultaneously. As described above, a first thread is associated with architecture state register 101a, a second thread is associated with architecture state register 101b, a third thread may be associated with architecture state register 102a, and a fourth thread may be associated with architecture state register 102b. Here, as described above, each architecture state register (101a, 101b, 102a, and 102b) may be referred to as a processing element, a thread slot, or a thread unit. As shown, architecture state register 101a is copied into architecture state register 101b, so that each architecture state / context can be stored for logical processor 101a and logical processor 101b. Other smaller resources may also be replicated for threads 101a and 101b in core 101, such as instruction pointers and renaming logic in allocator and renamer block 130. Some resources may be shared through partitioning, such as reorder buffers, ILTB 120, load / store buffers, and queues in reorder / retirement unit 135. Other resources, such as general internal registers, page table base registers, low-level data cache and data TLB 115, execution units 140, and portions of out-of-order unit 135, may be fully shared.

[0029] Processor 100 typically includes other resources that may be fully shared, shared through partitioning, or dedicated by / to processing elements. Figure 1 , an embodiment of a purely exemplary processor with illustrative logic units / resources of the processor is shown. Note that the processor may include or omit any of these functional units, as well as include any other known functional units, logic, or firmware not shown. As shown, core 101 includes a simplified, representative out-of-order (OOO) processor core. However, an in-order processor may be used in different embodiments. The OOO core includes: a branch target buffer 120 for predicting branches to be executed / taken; and an instruction translation buffer (I-TLB) 120 for storing address translation entries for instructions.

[0030] The core 101 also includes a decode module 125 coupled to the fetch unit 120 to decode the fetched elements. In one embodiment, the fetch logic includes individual sequencers associated with the thread slots 101a, 101b, respectively. Typically, the core 101 is associated with a first ISA that defines / specifies instructions that can be executed on the processor 100. Machine code instructions that are part of the first ISA typically include a portion of the instruction (called an opcode) that references / specifies the instruction or operation to be performed. The decode logic 125 includes circuitry that identifies these instructions from their opcodes and passes the decoded instructions into the pipeline for processing defined by the first ISA. For example, as discussed in more detail below, in one embodiment, the decoder 125 includes logic designed or adapted to identify specific instructions such as transactional instructions. As a result of the identification by the decoder 125, the architecture or core 101 takes specific, predefined actions to perform tasks associated with the appropriate instructions. It is important to note that any of the tasks, blocks, operations, and methods described herein can be performed in response to a single or multiple instructions; some of which may be new instructions or old instructions. Note that in one embodiment, decoders 126 recognize the same ISA (or a subset thereof). Alternatively, in a heterogeneous core environment, decoders 126 recognize a second ISA (a subset of the first ISA or a different ISA).

[0031] In one example, the allocator and renamer block 130 includes an allocator to reserve resources such as register files to store instruction processing results. However, threads 101a and 101b are potentially capable of out-of-order execution, where the allocator and renamer block 130 also reserves other resources, such as reorder buffers to track instruction results. Unit 130 may also include a register renamer to rename program / instruction reference registers to other registers internal to processor 100. Reorder / retirement unit 135 includes components such as the above-mentioned reorder buffers, load buffers, and store buffers to support out-of-order execution and then retire instructions executed out-of-order in order.

[0032] In one embodiment, the scheduler and execution unit block 140 includes a scheduler unit for scheduling instructions / operations on the execution units. For example, a floating point instruction is scheduled on a port of an execution unit with an available floating point execution unit. A register file associated with the execution unit is also included to store information instruction processing results. Exemplary execution units include floating point execution units, integer execution units, jump execution units, load execution units, store execution units, and other known execution units.

[0033] A low-level data cache and data translation buffer (D-TLB) 150 is coupled to the execution unit 140. The data cache will store recently used / operated elements (e.g., data operands), which may be kept in a memory consistency state. The D-TLB will store recent virtual / linear to physical address translations. As a specific example, the processor may include a page table structure to divide physical memory into multiple virtual pages.

[0034] Here, cores 101 and 102 share access to a higher level or further cache, such as a secondary cache associated with on-chip interface 110. Note that higher level or further refers to increasing or decreasing distance from the execution unit. In one embodiment, the higher level cache is a last level data cache - the last level cache in the memory hierarchy on processor 100 - such as a second or third level data cache. However, the higher level cache is not limited because it can be associated with or include an instruction cache. Instead, a trace cache (an instruction cache) can be coupled after decoder 125 to store recently decoded traces. Here, an instruction potentially refers to a macroinstruction (i.e., a general instruction recognized by a decoder), which can be decoded into multiple microinstructions (micro-operations).

[0035] In the depicted configuration, the processor 100 also includes an on-chip interface module 110. Historically, a memory controller, described in more detail below, has been included in a computing system external to the processor 100. In this case, the on-chip interface 110 is used to communicate with devices external to the processor 100, such as system memory 175, a chipset (typically including a memory controller hub connected to the memory 175 and an I / O controller hub connected to peripheral devices), a memory controller hub, a north bridge, or other integrated circuits. And in this case, the bus 105 may include any known interconnect, such as a multi-drop bus, a point-to-point interconnect, a serial interconnect, a parallel bus, a coherent (e.g., cache coherent) bus, a layered protocol architecture, a differential bus, and a GTL bus.

[0036] Memory 175 may be dedicated to processor 100 or shared with other devices in the system. Common examples of types of memory 175 include DRAM, SRAM, non-volatile memory (NV memory), and other known storage devices. Note that device 180 may include a graphics accelerator, processor or card coupled to a memory controller hub, data storage coupled to an I / O controller hub, a wireless transceiver, a flash memory device, an audio controller, a network controller, or other known devices.

[0037] However, recently, as more logic and devices are integrated on a single die such as a SOC, each of these devices can be incorporated into the processor 100. For example, in one embodiment, a memory controller hub is located on the same package and / or die as the processor 100. Here, a portion of the core (the on-core portion) 110 includes one or more controllers for interfacing with other devices (e.g., memory 175 or graphics device 180). This configuration includes interconnects, and controllers that interface with such devices are generally referred to as on-core (or non-core configurations). As an example, the on-chip interface 110 includes a ring interconnect for on-chip communication and a high-speed serial point-to-point link 105 for off-chip communication. However, in a SOC environment, even more devices (e.g., network interfaces, coprocessors, memory 175, graphics processors 180, and any other known computer devices / interfaces) can be integrated on a single die or integrated circuit to provide high functionality and low power consumption for a small scale factor.

[0038] In one embodiment, processor 100 can execute compiler, optimization and / or converter code 177 to compile, convert and / or optimize application code 176 to support the apparatus and method described herein or engage with it. Compiler generally includes a program or a set of programs to convert source text / code into target text / code. Generally, using a compiler to compile program / application code is divided into multiple stages and passes to convert high-level programming language code into low-level machine or assembly language code. However, a single-pass compiler can still be used for simple compilation. The compiler can utilize any known compilation technology and perform any known compiler operation, such as lexical analysis, preprocessing, parsing, semantic analysis, code generation, code conversion and code optimization.

[0039] Larger compilers typically contain multiple phases, but most often these phases are contained within two general phases: (1) the front end, which is where syntactic processing, semantic processing, and some transformations / optimizations may typically occur, and (2) the back end, which is where analysis, transformations, optimizations, and code generation typically occur. Some compilers reference an intermediate language that illustrates the blurring of the lines between the front end and back end of the compiler. As a result, references to the insertion, association, generation, or other operations of the compiler may occur in any of the aforementioned phases or processes of the compiler as well as any other known phases or processes. As an illustrative example, a compiler may insert operations, calls, functions, and the like in one or more compilation phases, such as inserting the calls / operations in the front end phase of compilation and then converting the calls / operations to low-level code during the transformation phase. Note that during dynamic compilation, compiler code or dynamically optimized code may insert such operations / calls and optimize the code for execution during runtime. As a specific illustrative example, binary code (compiled code) may be dynamically optimized during runtime. Here, program code may include dynamically optimized code, binary code, or a combination thereof.

[0040] Similar to compilers, translators such as binary translators statically or dynamically translate code to optimize and / or convert code. Thus, references to the execution of code, application code, program code, or other software environment may refer to: (1) executing a compiler program, an optimizing code optimizer, or a dynamic or static translator to compile program code, maintain software structure, perform other operations, optimize code, or convert code; (2) executing a main program code including operations / calls, such as optimized / compiled application code; (3) executing other program code (e.g., a library) associated with the main program code to maintain software structure, perform other software-related operations, or optimize code; or (4) a combination thereof.

[0041] Figure 2 is a schematic diagram of an example accelerator device 200 according to an embodiment of the present disclosure. Figure 2 As shown, in one implementation, the accelerator includes PCI configuration registers 204 and MMIO registers 210, which can be programmed to provide access to device backend resources 212. In one implementation, the base address of the MMIO registers 210 is specified by a set of base address registers (BARs) 202 in the PCI configuration space. Unlike previous implementations, an implementation of the data streaming accelerator (DSA) described herein does not implement multiple channels or PCI functions, so there is only one instance of each register in the device. However, there may be more than one DSA device in a single platform.

[0042] An implementation may provide additional performance or debug registers not described here. Any such registers should be considered implementation-specific.

[0043] PCI configuration space accesses are performed as aligned 1-byte, 2-byte, or 4-byte accesses. See the PCI Express Base Specification for the rules for accessing unimplemented registers and reserved bits in PCI configuration space.

[0044] MMIO space accesses to the BAR0 region (capability, configuration, and status registers) are performed as aligned 1, 2, 4, or 8-byte accesses. 8-byte accesses can only be used for 8-byte registers. Software should not read or write unimplemented registers. MMIO space accesses to the BAR 2 and BAR 4 regions should be performed as 64-byte accesses using the ENQCMD, ENQCMDS, or MOVDIR64B instructions (described in detail below). Work queues configured as shared (SWQ) should be accessed using ENQCMD or ENQCMDS, and work queues configured as dedicated (DWQ) must be accessed using MOVDIR64B.

[0045] One implementation of the DSA PCI configuration space implements three 64-bit BARs 202. The device control register (BAR0) is a 64-bit BAR that contains the physical base address of the device control registers. These registers provide information about the device capabilities, controls to configure and enable the device, and the status of the device. The size of the BAR0 area depends on the size of the interrupt message storage 208, which is 32KB plus the number of interrupt message storage entries 208 times 16, rounded to the next power of 2. For example, if the device supports 1024 interrupt message storage entries 208, the interrupt message storage is 16KB, and the size of BAR0 is 64KB.

[0046] BAR2 is a 64-bit BAR that contains the physical base addresses of privileged and non-privileged portals. Each portal is 64 bytes in size and resides on a separate 4KB page. This allows the portals to be independently mapped into different address spaces using the CPU page tables. The portals are used to submit descriptors to the device. Privileged portals are used by kernel-mode software, while non-privileged portals are used by user-mode software. The number of non-privileged portals is the same as the number of work queues supported. The number of privileged portals is work queues (WQs) × (MSI-X-table-size–1). The address of the portal used to submit a descriptor allows the device to determine which WQ to place the descriptor in, whether the portal is privileged or non-privileged, and which MSI-X table entry can be used to complete the interrupt. For example, if the device supports 8 WQs, the WQ for a given descriptor is (portal address >> 12) & 0x7. If portal address >> 15 is 0, it means that the portal is not privileged; otherwise it is privileged and the MSI-X 206 table index used to complete the interrupt is portal address >> 15. Bits 5:0 must be 0. Bits 11:6 are ignored; therefore, any 64-byte aligned address on the page may be used with the same effect.

[0047] Descriptor submissions using the non-privileged portal are subject to the WQ's occupancy threshold, which is configured using the Work Queue Configuration (WQCFG) register. Descriptor submissions using the privileged portal are not subject to this threshold. Descriptor submissions to the SWQ must be submitted using ENQCMD or ENQCMDS. Any other writes to the SWQ portal are ignored. Descriptor submissions to the DWQ must be submitted using 64-byte writes. Software uses MOVDIR64B to ensure uninterrupted 64-byte writes. ENQCMD or ENQCMDS for disabled or dedicated WQ portals returns a retry. Any other writes to the DWQ portal are ignored. Any reads to the BAR2 address space return all 1s. Kernel-mode descriptors should be submitted using the privileged portal in order to receive a completion interrupt. If a kernel-mode descriptor is submitted using the non-privileged portal, no completion interrupt can be requested. User-mode descriptors can be submitted using either the privileged or non-privileged portal.

[0048] The number of portals in the BAR2 region is the number of WQs supported by the device multiplied by the MSI-X 206 table size. The size of the MSI-X table is typically the number of WQs plus 1. So, for example, if the device supports 8 WQs, the useful size of BAR2 would be 8×9×4KB=288KB. The total size of BAR2 would be rounded to the next power of 2, which is 512KB.

[0049] BAR4 is a 64-bit BAR that contains the physical base addresses of the guest portals. Each guest portal is 64 bytes in size and resides in a separate 4KB page. This allows portals to be independently mapped into different address spaces using the CPU Extended Page Tables (EPT). If the Interrupt Message Storage Support field in GENCAP is 0, this BAR is not implemented.

[0050] Guest kernel mode software can submit descriptors to the device using a guest portal. The number of guest portals is the number of entries in the interrupt message store multiplied by the number of WQs supported. The address of the guest portal used to submit the descriptor allows the device to determine the WQ of the descriptor, and can also use the interrupt message store entry to generate a completion interrupt for the descriptor's completion (if it is a kernel mode descriptor and if the request completion interrupt flag is set in the descriptor). For example, if the device supports 8 WQs, the WQ for a given descriptor is (guest portal address >> 12) & 0x7, and the interrupt table entry index for the completion interrupt is guest portal address >> 15.

[0051] In one implementation, MSI-X is the only PCIe interrupt capability provided by the DSA, and the DSA does not implement traditional PCI interrupts or MSI. The details of this register structure are in the PCI Express specification.

[0052] In one implementation, three PCI Express capabilities control address translation. As shown in Table 1, only certain value combinations for these capabilities are supported. These values ​​are checked when the enable bit in the general control register (GENCTRL) is set to 1.

[0053] Table 1. Supported combinations of capabilities and associated values.

[0054]

[0055]

[0056] If any of these features are changed by software while the device is enabled, the device may stop functioning and report an error in the software error register.

[0057] In one implementation, software configures the PASID capability to control whether the device uses PASID to perform address translation. If PASID is disabled, only physical addresses can be used. If PASID is enabled, virtual or physical addresses can be used, depending on the IOMMU configuration. If PASID is enabled, both the Address Translation Service (ATS) and the Page Request Service (PRS) should be enabled.

[0058] In one implementation, software configures the ATS capability to control whether the device should translate addresses before performing memory accesses. If address translation is enabled in the IOMMU 2810, ATS must be enabled in the device to obtain acceptable system performance. If address translation is not enabled in the IOMMU 2810, ATS must be disabled. If ATS is disabled, only physical addresses can be used and all memory accesses are performed using untranslated accesses. If PASID is enabled, ATS must be enabled.

[0059] In one implementation, software configures the PRS capability to control whether the device can request a page when address translation fails. If PASID is enabled, PRS must be enabled; and if PASID is disabled, PRS must be disabled.

[0060] Some implementations utilize virtual memory space, which is seamlessly shared between one or more processor cores, accelerator devices, and / or other types of processing devices (e.g., I / O devices). In particular, an implementation utilizes a shared virtual memory (SVM) architecture, in which the same virtual memory space is shared between kernels, accelerator devices, and / or other processing devices. In addition, some implementations include heterogeneous forms of physical system memory, which use a common virtual memory space to address. Heterogeneous forms of physical system memory can use different physical interfaces for connection with the DSA architecture. For example, an accelerator device can be directly coupled to a local accelerator memory such as a high bandwidth memory (HBM), and each core can be directly coupled to a host physical memory such as a dynamic random access memory (DRAM). In this example, a shared virtual memory (SVM) is mapped to a combined physical memory of HBM and DRAM so that an accelerator, a processor core, and / or other processing devices can access HBM and DRAM using a consistent set of virtual memory addresses.

[0061] These and other feature accelerators are described in detail below. By way of brief overview, different implementations may include one or more of the following infrastructure features:

[0062] Shared Virtual Memory (SVM) : Some implementations support SVM, which allows user-level applications to submit commands directly to the DSA using virtual addresses in descriptors. The DSA may support translation of virtual addresses to physical addresses using an input / output memory management unit (IOMMU), including handling of page faults. The virtual address range referenced by a descriptor may span multiple pages distributed across multiple heterogeneous storage types. Alternatively, an implementation supports the use of physical addresses, as long as the data buffer is contiguous in physical memory.

[0063] Partial descriptor completion:With SVM support, an operation may encounter a failure during address translation. In some cases, the device may terminate processing of the corresponding descriptor at the point where the failure was encountered and provide a completion record to the software indicating partial completion and failure information to allow the software to take remedial action and retry the operation after the failure is resolved.

[0064] Batch Processing : Some implementations support submitting descriptors in "batches". A batch descriptor points to a set of actually contiguous work descriptors (i.e., descriptors that contain actual data operations). When processing a batch descriptor, the DSA fetches the work descriptors from the specified memory and processes them.

[0065] Stateless Devices : Design descriptors in an implementation so that all the information needed to process the descriptor goes into the descriptor payload itself. This allows the device to store very little client-specific state, improving its scalability. One exception is the completion interrupt message, which is configured by trusted software when used.

[0066] Cache allocation control : This allows the application to specify whether to write to the cache or bypass the cache and write directly to memory. In one implementation, completion records are always written to the cache.

[0067] Shared Work Queue (SWQ) support : Some implementations support scalable work submission through a shared work queue (SWQ) by using the Enqueue Command (ENQCMD) and Enqueue Command as Supervisor (ENQCMDS) instructions, as described below. In this implementation, the SWQ is shared by multiple applications. ENQCMD can be executed from either the user (non-ring 0) or supervisor (ring 0) privilege level. ENQCMDS can be executed from the superuser (ring 0) privilege level.

[0068] Dedicated Work Queue (DWQ) support : In some implementations, the MOVDIR64B instruction is used to support high throughput work submission through a dedicated work queue (DWQ). In this implementation, the DWQ is dedicated to a specific application.

[0069] QoS Support : Some implementations allow a Quality of Service (QoS) level to be specified for each work queue (e.g., via a kernel driver). It can then assign different work queues to different applications, allowing work from different applications to be dispatched from work queues with different priorities. Work queues can be programmed to use specific channels for fabric QoS.

[0070] Biased Cache Coherence Mechanism

[0071] An implementation improves the performance of accelerators with directly attached memory (e.g., stacked DRAM or HBM) and simplifies application development for applications using accelerators with directly attached memory. This implementation allows the accelerator-attached memory to be mapped as part of the system memory and accessed using shared virtual memory (SVM) techniques (such as those used in current IOMMU implementations), but without suffering the typical performance drawbacks associated with full system cache coherency.

[0072] The ability to access accelerator-attached memory as part of system memory without the addition of cumbersome cache coherence overhead provides a favorable operating environment for accelerator offloading. The ability to access memory as part of the system address map allows host software to set operands and access computation results without the overhead of traditional I / ODMA data copies. Such traditional copies involve driver calls, interrupts, and memory-mapped I / O (MMIO) accesses that are inefficient relative to simple memory accesses. At the same time, the ability to access accelerator-attached memory without cache coherence overhead is critical to the execution time of offloaded computations. For example, in situations with a large amount of streaming write memory traffic, cache coherence overhead can reduce the effective write bandwidth seen by the accelerator by half. The efficiency of operand setup, the efficiency of result access, and the efficiency of accelerator computation all play a role in determining how accelerator offloading works. If the cost of offloading work (e.g., setting operands; getting results) is too high, offloading may not pay off at all or may limit the accelerator to very large jobs. The efficiency with which the accelerator performs computations can have the same effect.

[0073] An implementation applies different memory access and consistency techniques depending on the entity initiating the memory access (e.g., accelerator, core, etc.) and the memory being accessed (e.g., host memory or accelerator memory). These techniques are generally referred to as "consistency biasing" mechanisms, which provide two sets of cache consistency streams for accelerator-attached memory, one of which is optimized to achieve efficient accelerator access to its attached memory, and the second set is optimized for host access to accelerator or attached memory and access to shared memory / host to accelerator-attached memory. In addition, it includes two techniques for switching between these streams, one driven by application software and the other driven by autonomous hardware hints. In both sets of consistency streams, the hardware maintains full cache consistency.

[0074] Figure 3 is a schematic diagram of an example computer system 300 that includes an accelerator and one or more computer processor chips coupled to a processor via a multi-protocol link. Figure 3, one implementation is suitable for a computer system including an accelerator 302 and one or more computer processor chips having a processor core and I / O circuits 304, wherein the accelerator 302 is coupled to the processor via a multi-protocol link 314. In one implementation, the multi-protocol link 314 is a dynamically multiplexed link that supports a variety of different protocols, including but not limited to the protocols detailed above. However, it should be noted that the underlying principles of the present invention are not limited to any particular set of protocols. In addition, it is noted that, depending on the implementation, the accelerator 302 and the core I / O 304 can be integrated on the same semiconductor chip or on different semiconductor chips.

[0075] In the illustrated implementation, an accelerator memory bus 316 couples the accelerator 302 to the accelerator memory 306, and a separate host memory bus 318 couples the core I / O 304 to the host memory 308. As described above, the accelerator memory 306 may include high bandwidth memory (HBM) or stacked DRAM (some examples of which are described herein), and the host memory 308 may include DRAM such as double data rate synchronous dynamic random access memory (e.g., DDR3 SDRAM, DDR4 SDRAM). However, the underlying principles of the invention are not limited to any particular type of memory or memory protocol.

[0076] In one implementation, both the accelerator 302 and the "host" software running on the processing core within the processor chip 304 use two different sets of protocol flows, referred to as "host biased" flows and "device biased" flows, to access the accelerator memory 306. As described below, one implementation supports multiple options for modulating and / or selecting the protocol flow for a particular memory access.

[0077] The consistency biasing process is implemented in part on a multi-protocol link 314 between the accelerator 302 and one of the processor chips 304 at two protocol layers: the CAC protocol layer and the MA protocol layer. In one implementation, the consistency biasing process is enabled by: (a) using existing opcodes in the CAC protocol in a new way, (b) adding new opcodes to the existing MA standard, and (c) adding support for the MA protocol to the multi-protocol link 302 (the previous link only included CAC and PCDI). Note that the multi-protocol link is not limited to supporting only CAC and MA. In one implementation, only at least those protocols need to be supported.

[0078] As used herein, a "host bias" flow, such as Figure 3Shown is a set of flows that funnel all requests to the accelerator memory 306 through the standard coherence controller 312 in the processor chip 304 to which the accelerator 302 is attached, including requests from the accelerator itself. This allows the accelerator 302 to take circuit routing to access its own memory, but allows access from the accelerator 302 and the processor core I / O 304 to be kept coherent using the processor's standard coherence controller 312. In one implementation, the flows use CAC opcodes to issue requests to the processor's coherence controller 312 over a multi-protocol link in the same or similar manner as the processor core 312 issues requests to the coherence controller 312. For example, the processor chip's coherence controller 312 can issue UPI and CAC coherence messages (e.g., snoops) that are generated on behalf of the accelerator to all peer processor core chips (e.g., 304) and internal processor agents for requests from the accelerator 302, just as they would for requests from the processor core 304. In this way, coherence is maintained between data accessed by the accelerator 302 and the processor core I / O 304.

[0079] In one implementation, the coherence controller 312 also conditionally issues memory access messages to the accelerator's memory controller 310 over the multi-protocol link 314. These messages are similar to the messages that the coherence controller 312 sends to the memory controller local to its processor die, and include new opcodes that allow data to be returned directly to the agent inside the accelerator 302, rather than forcing the data to be returned to the processor coherence controller 312 of the multi-protocol link 314 and then returned to the accelerator 302 over the multi-protocol link 314 as a CAC response.

[0080] exist Figure 3 In one implementation of the “host biased” mode shown in FIG, all requests from the processor core 304 for the accelerator-attached memory 306 are sent directly to the processor coherence controller 312, just as they are for normal host memory 308. The coherence controller 312 can apply its standard cache coherence algorithms and send its standard cache coherency messages, just as they would for accesses to the accelerator 302, and just as they would for accesses to normal host memory 308. The coherence controller 312 also conditionally sends MA commands over the multi-protocol link 314 for such requests, although in this case the MA stream returns data across the multi-protocol link 314.

[0081] Figure 44 is a schematic diagram of an example work queue implementation 400 according to an embodiment of the present disclosure. In some implementations, a work queue (WQ) stores "descriptors" submitted by software, an arbitrator for implementing quality of service (QoS) and fairness policies, a processing engine for processing descriptors, an address translation and cache interface, and a memory read / write interface. The descriptor defines the scope of work to be completed. Figure 4 As shown, in one implementation, there are two different types of work queues: a dedicated work queue 402 and a shared work queue 404. The dedicated work queue 402 stores descriptors for a single application 416, while the shared work queue 404 stores descriptors submitted by multiple applications 410-414. The hardware interface / arbitrator 406 schedules the descriptors from the work queues 402-404 to the accelerator processing engine 408 according to a specified arbitration policy (e.g., based on the processing requirements and QoS / fairness policies of each application 410-416).

[0082] Shared work queue support on endpoint devices

[0083] Figure 4 The concept of a shared work queue (SWQ) is shown, which allows multiple non-cooperative software agents (applications 410-414) to submit work through the shared work queue 404 using the ENQCMD / S instructions described herein.

[0084] The following considerations apply to endpoint devices that implement a shared work queue (SWQ).

[0085] SWQs and their enumeration: A device physical function (PF) may support one or more SWQs. Each SWQ is accessible for non-posted write queuing via a 64-byte aligned and sized register (referred to herein as SWQ_REG) within the device's MMIO address range. It is recommended that each such SWQ_REG on a device be located in a unique system page size (4KB) region. The device driver for a device is responsible for reporting / enumerating the SWQ capabilities, the number of SWQs supported, and the corresponding SWQ_REG addresses to software via the appropriate software interface. The driver may also optionally report the depth of supported SWQs for software tuning or informational purposes (although not required for functional correctness). For devices supporting multiple physical functions, it is recommended to support separate SWQs for each physical function.

[0086] SWQ Support on Single Root I / O Virtualization (SR-IOV) Devices: Devices supporting SR-IOV can support independent SWQs for each virtual function (VF), which are exposed through the SWQ_REG in the corresponding VF base address register (BAR). This design point allows for maximum performance isolation for work submission across VFs and may be suitable for a small to moderate number of VFs. For devices supporting a large number of VFs (where independent SWQs per VF are not feasible), a single SWQ can be shared across multiple VFs. Even in this case, each VF has its own dedicated SWQ_REG in its VF BAR unless they are backed by a common SWQ across VFs that share the SWQ. For such device designs, which VFs share the SWQ can be statically determined by the hardware design, or the mapping between the SWQ_REG of a given VF to the SWQ instance can be dynamically set / reduced by the physical function and its driver. Device designs that share SWQs across VFs require special attention to QoS and protection against denial of service attacks, as described later in this section. When a SWQ is shared across VFs, care must be taken in the device design to identify which VF received an enqueue request accepted by the SWQ. When dispatching work requests from a SWQ, the device should ensure that the upstream request is properly tagged with the requestor ID (bus / device / function number) of the corresponding VF (in addition to the PASID communicated in the enqueue request payload).

[0087] Enqueue non-posted write addresses: SWQ-capable endpoint devices are required to accept queued non-posted writes to any address routed through their PF or VF memory BARs. For any enqueued non-posted write request received by the endpoint device for a non-SWQ_REG address, the device may be required not to treat it as an error (e.g., malformed TLP, etc.) and instead return a completion with a completion status of retry (MRS). This may be done to ensure that non-privileged (ring-3 or ring-0 VMX guest) software that erroneously or maliciously allocates enqueued stores to non-SWQ_REG addresses on SWQ-capable devices using the ENQCMD / S instruction does not result in non-fatal or fatal error reporting with platform-specific error handling consequences.

[0088] Non-queued request handling to SWQ_REG: SWQ-capable endpoint devices may silently discard non-queued requests (normal memory writes and reads) to SWQ_REG addresses without treating them as fatal or non-fatal errors. Read requests to SWQ_REG addresses may return a successful completion response (as opposed to UR or CA) with the data bytes of the request all having a value of 1. Normal memory (posted) write requests to SWQ_REG addresses may be simply discarded by the endpoint device without any action required. This may be done to ensure that non-privileged software cannot generate normal read and write requests to SWQ_REG addresses that could, by mistake or maliciously, cause non-fatal or fatal error reporting with platform-specific error handling consequences.

[0089] SWQ Queue Depth and Storage: SWQ queue depth and storage are device implementation specific. Device designs should ensure that sufficient queue depth is supported for SWQ to achieve maximum utilization of the device. Storage for SWQ can be implemented on the device. Integrated devices on a SoC may utilize stolen main memory (non-OS-visible private memory reserved for device use) as an overflow buffer for SWQ, allowing larger SWQ queue depths than are available with storage on the device. For such designs, the use of the overflow buffer is transparent to software, and the device hardware decides when to overflow (rather than discarding the enqueue request and sending a retry completion status), draws from the overflow buffer to execute a command, and maintains any command-specific ordering requirements. For all purposes, the use of such an overflow buffer is equivalent to a discrete device using local device-attached DRAM for SWQ storage. In device designs with overflow buffers in stolen memory, special care must be taken to ensure that such stolen memory is not protected from any access other than overflow buffer reads and writes by the device to which the buffer was allocated.

[0090] Non-blocking SWQ behavior: For performance reasons, device implementations SHOULD respond quickly to "enqueue a non-posted write request" with a "success" or "retry" completion status, and not block SWQ capacity to be freed up to accept the request for enqueue completion. The decision to accept or reject an enqueue request to a SWQ MAY be based on capacity, QoS / occupancy, or any other policy. Some example QoS considerations are described next.

[0091] SWQ QoS Note: For enqueued non-posted writes targeting a SWQ_REG address, the endpoint device may apply admission control to decide to accept the request to the corresponding SWQ (and send a successful completion status) or to discard the request (and send a retry completion status). Admission control may be device and usage specific, and the specific policy supported / enforced by the hardware may be exposed to software through the Physical Function (PF) driver interface. Because the SWQ is a resource shared with multiple producer clients, device implementations must ensure adequate protection to prevent cross-producer denial of service attacks. QoS for a SWQ refers only to the acceptance of work requests (via enqueued requests) to the SWQ, and is orthogonal to any QoS applied by the device hardware to how QoS is applied to share the device's execution resources when processing work requests submitted by different producers. Some example methods are described below for configuring an endpoint device to enforce an admission policy for accepting enqueued requests to a SWQ. These documents are for illustrative purposes only, and the exact implementation choices will be device specific.

[0092] These CPU instructions generate atomic non-posted write transactions (write transactions for which a completion response is returned to the CPU). The address of a non-posted write transaction is sent just like any normal MMIO write to the target device. A non-posted write transaction carries the following information:

[0093] a. Identification of the address space of the client executing this instruction. For this disclosure, we use the "Process Address Space Identifier" (PASID) as the client address space identifier. Depending on the use of the software, the PASID can be used for any type of client (process, container, VM, etc.). Other implementations of the present invention may use different identification schemes. The ENQCMD instruction uses the PASID associated with the current software thread (hopefully using the XSAVES / XRSTORS instructions on the thread context switcher to save / restore certain contents of the OS). The ENQCMDS instruction allows privileged software executing the instruction to specify the PASID as part of the source operand of the ENQCMDS instruction.

[0094] b. The privileges of the client executing this instruction (hypervisor or user). Execution of ENQCMD always indicates user permissions. ENQCMDS allows supervisor-mode software executing it to specify either user privileges (if it is executing on behalf of some ring-3 client) or supervisor privileges (indicating that the command came from a kernel-mode client) as part of the source operand of the ENQCMDS instruction.

[0095] c. Command payload specific to the target device. The command payload is read from the source operand and delivered as is via the instruction in the non-posted write payload. Depending on the device, the device may use ENQCMD / S in different ways. For example, some devices may treat it as a doorbell command where the payload specifies the actual work descriptor in memory to be pulled from. Other devices may use the actual ENQ command to carry the device specific work descriptor, thus avoiding the latency / overhead of reading the work descriptor from main memory.

[0096] The SWQ on the device handles the received non-posted write request as follows:

[0097] At the entry of the device SWQ, check whether there is room in the SWQ to accept the request. If there is no room, discard the request and return a completion indicating "reject / retry" in the completion status.

[0098] If there is room to accept the command to the SWQ, any required device-specific admission control is performed based on attributes in the request (e.g., the PASID, privileges, or SWQ_PREG addresses to which the request was routed) and various device-specific QoS settings for the particular client. If the admission control method determines that the request cannot be accepted by the SWQ, the request is discarded and completion is returned with a completion status of "reject / retry".

[0099] If the above checks result in a non-posted write command being accepted by the SWQ, a completion status is returned with a completion status of "Completed Successfully". Commands enqueued to the SWQ are processed / scheduled based on a device-specific scheduling mode internal to the device.

[0100] When the work specified by the command is completed, the device generates appropriate synchronization or notification transactions to notify the client of the completion of the work. These can be done by memory writes, interrupt writes, or other methods.

[0101] The ENQCMD and ENQCMDS instructions will block until the CPU returns a completion response. The instruction returns the status (success v / s rejection / retry) in the EFLAGS.ZF flags before exiting.

[0102] Software can enqueue work via SWQ as follows:

[0103] a. Map the device's SWQ_PREG into the client's CPU virtual address space (either user or kernel virtual address, depending on whether the client is running in user mode or kernel mode). This is similar to memory mapping any MMIO resource on the device.

[0104] b. Format descriptor in memory

[0105] c. Execute ENQCMD / S using the memory virtual address of the descriptor as the source and the virtual address to which SWQ_PREG is mapped as the target.

[0106] d. Perform a conditional jump (JZ) to check if the ENQCMD / S instruction returns success or retry. If it is a retry condition, retry from step C (with an appropriate return

[0107] Any agent between the CPU and the device (uncores, coherency fabrics, I / O fabrics, bridges, etc.) may discard an ENQ non-posted write. If any agent discards an ENQ request, that agent must return a completion response with a retry status code. Software sees and treats it just like a retry response received from the target SWQ on the device. This is possible because ENQ* requests are self-contained and do not have any ordering constraints associated with them. System design and architecture can take advantage of this property to handle timing congestion or backpressure situations in the hardware / fabric.

[0108] This property can also be used to build optimized designs where the SWQ can be moved closer to the CPU and a dedicated posted / credit-based approach is used to forward accepted commands to the SWQ to the target device. Such an approach can be useful for improving the round-trip latency of ENQ instructions, which otherwise require sending a non-posted write request all the way to the SWQ on the device and waiting for completion.

[0109] SWQ can be implemented on the device by using dedicated on-chip storage (SRAM) or using extended memory on the device (DRAM) or even reserved / stolen memory in the platform. In these cases, the SWQ is just a front end for clients to submit work using non-posted semantics, and the device writes all accepted commands to the SWQ to a work queue in memory, from which the various operations of the engine can fetch work just like any regular memory-based work queue. This memory-backed SWQ implementation can allow for much larger SWQ capacity than with dedicated on-chip SRAM storage.

[0110] Figure 5 is a schematic diagram of an example data streaming accelerator (DSA) device including a plurality of work queues that receive descriptors submitted through an I / O fabric interface. Figure 5An implementation of a data streaming accelerator (DSA) device including multiple work queues 510-512 is shown, which receive descriptors submitted through an I / O fabric interface 504 (e.g., a multi-protocol link 314 as described above). The DSA uses the I / O fabric interface 504 to receive downstream work requests from clients (e.g., processor cores, peer input / output (IO) agents (e.g., network interface controllers (NICs)), and / or software chain offload requests) and read, write, and address translation operations for upstream. The illustrated implementation includes an arbitrator 514 that arbitrates between work queues and dispatches work descriptors to one of multiple engines 518. The operation of the arbitrator 514 and the work queues 510-512 can be configured through the work queue configuration registers 502. For example, the arbitrator 514 can be configured to implement various QoS and / or fairness policies for allocating descriptors from each work queue 510-512 to each engine 518.

[0111] In one implementation, some of the descriptors enqueued in the work queues 510-512 are batch descriptors 520, which contain / identify a batch of work descriptors. The arbiter 514 forwards the batch descriptors to the batch unit 524, which processes the batch descriptors by reading the descriptor array 526 from the memory using the addresses translated by the translation cache 506 (or other address translation services possible on the processor). Once the physical address is identified, the data read / write circuit 508 reads the batch descriptors from the memory.

[0112] The second arbitrator 528 arbitrates between the batch work descriptors 526 provided by the batch processing unit 524 and the individual work descriptors 522 obtained from the work queues 510-512, and outputs the work descriptors to the work descriptor processing unit 530. The work descriptor processing unit 530 has stages of reading memory (via the data R / W unit 508), performing requested operations on the data, generating output data and writing the output data (via the data R / W unit 508), completing recording and interrupt messages.

[0113] In one implementation, the work queue configuration allows software (via WQ configuration registers 502) to configure each WQ as either a shared work queue (SWQ) that uses non-posted ENQCMD / S instruction receive descriptors or a dedicated work queue (DWQ) that uses posted MOVDIR64B instruction receive descriptors. Figure 4As mentioned, a DWQ can handle work descriptors and batch descriptors submitted from a single application, while a SWQ can be shared between multiple applications. The WQ configuration registers 502 also allow software to control which WQ 510-512 feeds which accelerator engine 518 and the relative priority of the WQs 510-512 fed to each engine. For example, an ordered set of priorities (e.g., high, medium, low; 1, 2, 3, etc.) can be specified, and descriptors can generally be scheduled earlier or more frequently than a higher priority work queue than from a low priority work queue. For example, for two work queues to be identified as high priority and low priority, for every 10 descriptors to be allocated, 8 of the 10 descriptors can be allocated from the high priority work queue, while 2 of the 10 descriptors can be allocated from the low priority work queue. Various other techniques can be used to implement different priorities between work queues 510-512.

[0114] In one implementation, the Data Streaming Accelerator (DSA) is software compatible with the PCI Express configuration mechanism and implements the PCI header and expansion space in its configuration mapped register set. The configuration registers can be programmed via CFC / CF8 or MMCFG from the root complex. All internal registers can also be accessed via the JTAG or SMBus interface.

[0115] In one implementation, the DSA device uses memory mapped registers to control its operation. Capability, configuration, and work submission registers (portals) are accessible through MMIO regions defined by the BAR0, BAR2, and BAR4 registers (described below). Each portal can be on a separate 4K page so that they can be independently mapped into different address spaces (clients) using the processor page tables.

[0116] As mentioned above, software specifies work to the DSA via descriptors. The descriptor specifies the type of operation to be performed for the DSA, the addresses of the data and status buffers, immediate operands, completion attributes, etc. (Additional details and details of the descriptor format are listed below). The completion attributes specify the address where the completion record will be written, as well as the information required to generate an optional completion interrupt.

[0117] In one implementation, DSA avoids maintaining client-specific state on the device. All information for processing a descriptor comes from the descriptor itself. This improves its shareability between user-mode applications and between different virtual machines (or machine containers) in a virtualized system.

[0118] A descriptor can contain an operation and associated parameters (called a work descriptor) or it can contain the address of an array of work descriptors (called a batch descriptor). Software prepares the descriptor in memory and then submits the descriptor to the device's work queue (WQ) 510-512. Descriptors are submitted to the device using the MOVDIR64B, ENQCMD, or ENQCMDS instructions, depending on the mode of the WQ and the privilege level of the client.

[0119] Each WQ 510-512 has a fixed number of slots and therefore may become full under heavy load. In one implementation, the device provides the required feedback to help the software implement flow control. The device allocates descriptors from the work queues 510-512 and submits them to the engine for further processing. When the engine 518 completes a descriptor or encounters some fault or error that causes an abort, it notifies the host software by writing a completion record in host memory, issuing an interrupt, or both.

[0120] In one implementation, each work queue is accessible via multiple registers, each located in a separate 4KB page in the device MMIO space. One work submission register per WQ is called the "unprivileged portal" and is mapped into user space for use by user-mode clients. Another work submission register is called the "privileged portal" and is used by kernel-mode drivers. The rest are guest portals for use by kernel-mode clients in virtual machines.

[0121] As described above, each work queue 510-512 can be configured to operate in one of two modes, dedicated or shared. The DSA exposes the capability bits in the work queue capability register to indicate support for dedicated and shared modes. It also exposes controls in the work queue configuration register 502 to configure each WQ to operate in one mode. The mode of the WQ can only be changed when the WQ is disabled (i.e. (WQCFG.Enabled = 0)). Other details of the WQ capability register and the WQ configuration register are as follows.

[0122] In one implementation, in shared mode, a DSA client submits descriptors to a work queue using the ENQCMD or ENQCMDS instructions. ENQCMD and ENQCMDS use a 64-byte non-posted write and wait for a response from the device before completing. The DSA returns "success" (e.g., to the requesting client / application) if there is room in the work queue and "retry" if the work queue is full. The ENQCMD and ENQCMDS instructions may return the status of the command submission with a zero flag (0 for success and 1 for retry). Using the ENQCMD and ENQCMDS instructions, multiple clients can submit descriptors to the same work queue directly and simultaneously. Because the device provides this feedback, clients can tell whether their descriptors were accepted.

[0123] In shared mode, the DSA may reserve some SWQ capabilities for submissions via the privileged portal for kernel mode clients. Work submissions via non-privileged portals are accepted until the number of descriptors in the SWQ reaches a threshold configured for the SWQ. Work submissions via privileged portals will be accepted until the SWQ is full. Work submissions via guest portals are limited by thresholds in the same manner as non-privileged portals.

[0124] If the ENQCMD or ENQCMDS instruction returns "success", the descriptor has been accepted by the device and queued for processing. If the instruction returns "retry", the software can attempt to resubmit the descriptor to the SWQ, or if it is a user-mode client using the unprivileged portal, it can request the kernel-mode driver to submit the descriptor on its behalf using the privileged portal. This helps avoid denial of service and provides forward progress guarantees. Alternatively, if the SWQ is full, the software can use other methods (e.g., use the CPU to perform the work).

[0125] The device uses a 20-bit ID called the process address space ID (PASID) to identify the client / application. The device uses the PASID to look up the address in the device TLB 1722 and sends the address translation or page request to the IOMMU 1710 (e.g., over the multi-protocol link 2800). In shared mode, the PASID to be used with each descriptor is contained in the PASID field of the descriptor. In one implementation, ENQCMD copies the PASID of the current thread from a specific register (e.g., the PASID MSR) into the descriptor, while ENQCMDS allows the hypervisor mode software to copy the PASID into the descriptor.

[0126] Although dedicated mode cannot share a single DWQ by multiple clients / applications, a DSA device can be configured with multiple DWQs, and each DWQ can be independently assigned to a client. In addition, DWQs can be configured with the same or different QoS levels to provide different performance levels for different clients / applications.

[0127] In one implementation, the data streaming accelerator (DSA) includes two or more engines 518 that process descriptors submitted to the work queues 510-512. One implementation of the DSA architecture includes four engines, numbered 0 to 3. Engines 0 and 1 are each able to utilize the full bandwidth of the device (e.g., 30 GB / s read and 30 GB / s write). The combined bandwidth of all engines is also limited to the maximum bandwidth available to the device.

[0128] In one implementation, software configures WQs 510-512 and engines 518 into groups using group configuration registers. Each group contains one or more WQs and one or more engines. The DSA can use any engine in a group to process descriptors delivered to any WQ in the group, and each WQ and each engine can only be in one group. The number of groups can be the same as the number of engines, so each engine can be in a separate group, but if any group contains more than one engine, not all groups need to be used.

[0129] Although the DSA architecture provides great flexibility in configuring work queues, groups, and engines, the hardware can be narrowly designed for a specific configuration. Engines 0 and 1 can be configured in one of two different ways, depending on the software requirements. A recommended configuration is to place engines 0 and 1 in the same group. The hardware uses either engine to process descriptors from any work queue in the group. In this configuration, if one engine is stalled due to high latency memory address translation or page faults, the other engine can continue to run and maximize the throughput of the entire device.

[0130] Figure 6A-6B is a schematic diagram illustrating an example DMWr request scenario according to an embodiment of the present disclosure. Figure 6A-6B is a schematic diagram illustrating an example DMWr request scenario according to an embodiment of the present disclosure. Fig. 6AA system 600 is shown that includes two work queues 610-612 and 614-616 in each group 606 and 608, respectively, but there can be any number up to the maximum number of supported WQs. The WQs in a group can be shared WQs with different priorities, or one shared WQ and other dedicated WQs, or multiple dedicated WQs with the same or different priorities. In the example shown, group 606 is served by engines 0 and 1 602, and group 608 is served by engines 2 and 3 604. Engines 0, 1, 2, and 3 can be similar to engine 518.

[0131] like Figure 6B As shown, another system 620 using engines 0 622 and 1 624 places them in separate groups 630 and 632, respectively. Similarly, group 2 634 is assigned to engine 2 626, while group 3 is assigned to engine 3 628. In addition, group 0 630 is composed of two work queues 638 and 640; group 1 632 is composed of work queue 642; work queue 2 634 is composed of work queue 644; and group 3 636 is composed of work queue 646.

[0132] This configuration can be selected when software wants to reduce the possibility that latency-sensitive operations are blocked behind other operations. In this configuration, software submits latency-sensitive operations to work queue 642 connected to engine 1 626, and submits other operations to work queues 638-640 connected to engine 0 622.

[0133] Engine 2 626 and Engine 3 628 may be used, for example, to write to high bandwidth non-volatile memory (e.g., phase change memory). The bandwidth capabilities of these engines may be sized to match the expected write bandwidth of this type of memory. For this usage, bits 2 and 3 of the Engine Configuration Register should be set to 1, indicating that Virtual Channel 1 (VC1) should be used for traffic from these engines.

[0134] In platforms without high-bandwidth, nonvolatile memory (e.g., phase-change memory), or when DSA devices are not used to write to this type of memory, engines 2 and 3 may not be used. However, software can use them as additional low-latency paths as long as the submitted operations can tolerate the limited bandwidth.

[0135] As each descriptor reaches the head of the work queue, it may be removed by the scheduler / arbiter 514 and forwarded to one of the engines in the group. For batch descriptors 520 that reference work descriptors 526 in memory, the engine fetches the array of work descriptors from memory (i.e., using batch unit 524).

[0136] In one implementation, for each work descriptor 522, the engine 518 pre-fetches the translation of the completion record address and passes the operation to the work descriptor processing unit 530. The work descriptor processing unit 530 uses the device TLB and IOMMU for source and target address translation, reads the source data, performs the specified operation, and then writes the target data back to the memory. After the operation is completed, if the work descriptor requests, the engine will write the completion record to the pre-translated completion address and generate an interrupt.

[0137] In one implementation, multiple work queues of a DSA may be used to provide multiple levels of Quality of Service (QoS). The priority of each WQ may be specified in the WQ configuration register 502. The priority of a WQ is relative to other WQs in the same group (e.g., the priority of a WQ in the group itself is meaningless). Work queues in a group may have the same or different priorities. However, it does not make sense to configure multiple shared WQs with the same priority in the same group, since a single SWQ may achieve the same purpose. The scheduler / arbitrator 514 assigns work descriptors from the work queues 510-512 to the engines 518 according to their priorities.

[0138] The accelerator device and the high-performance I / O device support service requests directly from multiple clients. In this case, the term "client" (also referred to as an entity in this article) can include any of the following:

[0139] Multiple user-mode (ring 3) applications that are submitting direct user-mode I / O requests to the device;

[0140] Multiple kernel-mode (ring 0) drivers running in multiple virtual machines (VMs) sharing the same device;

[0141] Multiple software agents running in multiple containers (with an OS that supports container technology);

[0142] Any combination of the above (e.g., Ring 3 apps inside a VM, containers hardened by running in a VM, etc.);

[0143] A peer I / O agent submits work directly for efficient inline acceleration (e.g., a NIC device using a cryptographic device for cryptographic acceleration, or a touch controller or image processing unit using a GPU for advanced sensor processing); or

[0144] Host software offloads requests across accelerator links, where an accelerator device can forward work to another designated host software accelerator to chain work without bouncing through the host (e.g., a compression accelerator first compresses and then chains the work to a bulk encryption accelerator for encrypting the compressed data).

[0145] The term “directly” above means sharing the device with multiple clients without intermediate software layers in the control and data paths, such as sharing with a common kernel driver in the case of user applications, or sharing with a common VMM / hypervisor layer in the case of VMs) to reduce software overhead.

[0146] Examples of multi-client accelerators / high-performance devices may include a programmable GPU, a configurable offload device, a reconfigurable offload device, a fixed-function offload device, a high-performance host fabric controller device, or a high-performance I / O device.

[0147] As described above, scalability of work submission can be addressed by using a shared work queue. This disclosure describes a mechanism to apply a shared work queue and achieve scalability of work submission using an interconnect protocol based on the PCIe specification.

[0148] This disclosure describes a PCIe packet type, referred to herein as a "delayed memory write request" (DMWr). In an embodiment, the term "acknowledged memory write" (AMWr) request is used, and it should be understood that these concepts are similar. A DMWr packet may include the following features and functions:

[0149] DMWr packets are non-posted transactions and are therefore handled differently by PCIe flow control than posted MWr packets;

[0150] As a non-posted transaction, the DMWr transaction needs to send completion back to the requester; and

[0151] Unlike MWr transactions (where completers must accept all well-formed requests), DMWr completers may choose to accept or reject AMWr requests for any implementation-specific reasons.

[0152] FIG. 7A to FIG. 7D is a schematic diagram illustrating an example acknowledge memory write (DMWr) request and response message flow according to an embodiment of the present disclosure. Typically, an I / O device (e.g., an accelerator device) supports a common command interface for work submission from its clients. The common command interface may be referred to as a shared work queue (SWQ). In Figure 7A-7D , the SWQ is shown as command queue 704. The SWQ may be a command queue 704 implemented on an accelerator device. The depth of the command queue 704 is implementation specific and may be sized based on the number of outstanding commands required to feed the device to achieve its full throughput potential.

[0153] The accelerator 702 with a fixed length FIFO command queue 704 can directly receive uncoordinated commands from multiple software (or hardware) entities (e.g., entity A and entity B in the following example). Each entity can issue commands (e.g., work descriptors) by issuing AMWr requests to a single fixed memory address. Note that each entity can issue these commands directly to the device without any kind of coordination mechanism between the other entities issuing the commands.

[0154] exist Fig. 7A In the example shown, the accelerator command queue 704 is almost full. 1) Entity A issues a command to the queue 704 via a DMWr packet. 2) The accelerator may accept the command into the command queue. 3) The accelerator 702 may respond with a successful completion (SC). Note that for implementation-specific reasons (e.g., implementing a TCP-like flow control mechanism), such a device may also choose to reject packets before the queue is completely full.

[0155] exist Figure 7B , command queue 704 is full. 4) Entity B issues a command to command queue 704 while command queue 704 is full. 5) Accelerator 702 rejects the request, and 6) Accelerator 702 sends a completion with a request retry status (RRS). 7) If entity B attempts to reissue the command while the queue is still filled, accelerator 702 may continue to send completions associated with RRS in the status field.

[0156] exist Figure 7C In , accelerator 702 completes processing of the command and space becomes available on its queue. Fig.7D 9) Entity B (or any other entity) reissues the command to the accelerator 702. 10) The accelerator 702 may now accept the command into the command queue 704. 11) The accelerator 702 may respond with a Successful Completion (SC) message.

[0157] Figure 8 800 is a process flow diagram for performing scalable work submission according to an embodiment of the present disclosure. First, an I / O device such as an accelerator can receive a command from an entity into its command queue (802). If the command queue is full (804), the accelerator can send a completion (RRS) message with a request retry status to the entity (806). If the command queue is not full (804), the accelerator can send a successful completion (SC) message to the entity (808) and can accept the command into the command for processing (810).

[0158] Delayed Memory Write (DMWr):

[0159] In some embodiments, DMWr can be used to send data over PCIe. Deferrable memory writes require the completer to return an acknowledgment to the requester and provide a mechanism for the recipient to defer received data. DMWr TLPs facilitate use cases that were previously impossible (or more difficult to implement) on PCIe-based devices and links. DMWr provides a mechanism for endpoints and hosts to choose to execute or defer incoming DMWr requests. Endpoints and hosts can use this mechanism to simplify the design of flow control and fixed-length queue mechanisms. Using DMWr, a device can have a shared work queue and accept work items from multiple non-cooperative software agents in a non-blocking manner. DMWr can be defined as a memory write in which the requester attempts to write to a given location in the memory space. The completer can accept or reject this write (similar to the above-mentioned RRS) by returning a completion with a status of SC or a memory request retry status (MRS), respectively.

[0160] Deferrable Memory Write (DMWr) is an optional non-posted request that enables scalable, high-performance mechanisms for shared work queues and similar functionality. Using DMWr, a device can have a shared work queue and accept work items from multiple non-cooperating software agents in a non-blocking manner.

[0161] The following requirements apply to DMWr completers (a completer can be any entity that completes a non-posted write request transaction):

[0162] A completer that supports DMWr requests treats a properly formatted DMWr request as a Successful Completion (SC), Request Retry Status (MRS), Unsupported Request (UR), or Completion Abort (CA) for any location in its target memory space.

[0163] A completer that supports DMWr shall treat any well-formed DMWr request with a type or operand size that it does not support as an Unsupported Request (UR). The value at the destination location MUST remain unchanged.

[0164] Implementers that support DMWr can implement the restricted programming model. Optimizations based on the restricted programming model are defined in the PCIe specification.

[0165] If any function in a multifunction device supports DMWr Completer or DMWr Routing capabilities, then all functions in that device with memory space BARs must decode well-formed DMWr requests and handle any that are not supported as Unsupported Requests (UR). Note that in such devices, functions that lack DMWr Completer capability must not handle well-formed DMWr requests as malformed TLPs.

[0166] Unless there is a higher priority error, a DMWr-aware completer MUST handle the poisoned DMWr request as an error that a poisoned TLP was received, and MUST also return a completion with a completion status of Unsupported Request (UR). The value of the target location MUST remain unchanged.

[0167] If the completer of a DMWr request encounters an uncorrectable error while accessing the target location, the completer MUST handle this as a Completer Abort (CA). The subsequent state of the target location is implementation specific.

[0168] Completers are allowed to support DMWr requests on subsets of their target memory space as required by their programming model. PCI Express defined or inherited memory space structures (e.g., MSI-X table structures) are not required to be supported as DMWr targets unless explicitly stated in the structure description.

[0169] If the RC has any root ports that support DMWr routing capabilities, then all RCiEPs in the RC that are reachable by the forwarded DMWr request must decode the well-formed DMWr request and process any RCiEPs that they do not support as Unsupported Requests (UR).

[0170] The following requirements apply to root complexes and switches that support DMWr routing:

[0171] If a switch supports DMWr routing for any of its ports, the switch supports DMWr routing for all ports.

[0172] For a switch or RC, when DMWr egress blocking is enabled in an egress port and the target of a DMWr request exits that egress port, then the egress port handles the request as a DMWr egress blocking error and must also return a completion with a completion status of UR. If the severity of the DMWr egress blocking error is not fatal, then this situation is handled as a recommended non-fatal error as described in the PCIe specification, which governs completers that send completions with a status of UR / CA.

[0173] For the RC, support for DMWr requested peer routing and completion between root ports is optional and implementation dependent. If the RC supports DMWr routing functionality between two or more root ports, the RC indicates the capability in each associated root port via the DMWr Routing Supported bit in the Device Capability 2 Register.

[0174] The RC is not required to support DMWr routing between all root port pairs with the DMWr routing support bit set. Software should not assume that DMWr routing is supported between all root port pairs with the DMWr routing support bit set.

[0175] If the RC supports DMWr routing capability between two or more root ports, the RC indicates the capability in each associated root port via the DMWr Routing Support bit in the Device Capability 2 Register.

[0176] A DMWr request routed through a root port that does not support DMWr is processed as an Unsupported Request (UR).

[0177] Implementation Notes: Design Considerations for Delayable Memory Writes

[0178] In some embodiments, devices and device drivers may use DMWr to implement control mechanisms.

[0179] As a non-posted request, a DMWr TLP completes the transaction with a Completion TLP. In addition, PCIe ordering rules dictate that a non-posted TLP cannot pass a posted TLP, making posted transactions preferable for improved performance. Because a DMWr TLP cannot pass a Memory Read Request TLP, and because a DMWr TLP can be deferred by a completer, device and device driver manufacturers should be careful when attempting to read a memory location that is also the target of an uncompleted DMWr transaction.

[0180] Implementation Note: Ensuring Forward Progress with Deferrable Memory Writes

[0181] When a single shared work queue is created using DMWr transactions, care must be taken to ensure that no requestor is denied access to the queue indefinitely due to congestion. Software entities writing to such a queue may choose to implement a flow control mechanism or may rely on a specific programming model to ensure that all entities are able to make forward progress. This programming model may include feedback mechanisms or indications from functions to software on the state of the queue, or timers that delay DMWr requests after completion with an MRS condition.

[0182] Memory transactions include the following types:

[0183] Read request / completion;

[0184] Write request;

[0185] Delaying write requests; and

[0186] AtomicOp request / completion.

[0187] Memory transactions use two different address formats:

[0188] Short address format: 32-bit address; and

[0189] Long address format: 64-bit address.

[0190] Certain memory transactions may optionally have a PASID TLP prefix that contains a Process Address Space ID (PASID).

[0191] Fig.9A and Fig. 9B is a visual representation of the AMWr grouping. Fig.9A is a schematic diagram of a 64-bit DMWr (or AMWr) packet definition 900 according to an embodiment of the present disclosure. Fig. 9B is a schematic diagram of a 32-bit DMWr (or AMWr) packet definition 920 according to an embodiment of the present disclosure.

[0192] Table 2 below shows the TLP definitions supporting DMWr.

[0193] Table 2. TLP definitions supporting DMWr.

[0194]

[0195] Table 3 below provides the definition of completion status.

[0196] Table 3. Definition of completion status.

[0197]

[0198]

[0199] Address-based routing rules can also facilitate DMWr requests. Address routing is used with memory and I / O requests. Two address formats are specified: a 64-bit format for 4DW headers and a 32-bit format for 3DW headers. Fig.9A – Fig. 9B Indicates 64-bit and 32-bit formats respectively.

[0200] For Memory Read, Memory Write, Deferrable Memory Write, and AtomicOp requests, the Address Type (AT) field is encoded as shown in Table 10-1 Address Type (AT) Field Encoding. For all other requests, the AT field is a reserved field unless explicitly stated otherwise. LN Read and LN Write have special requirements.

[0201] Memory Read, Memory Write, Deferrable Memory Write, and AtomicOp requests can use either format.

[0202] For addresses below 4GB, the requester MUST use the 32-bit format. If a 64-bit format request is received for an address below 4GB (i.e., the upper 32 bits of the address are all 0), the receiver's behavior is unspecified.

[0203] I / O read requests and I / O write requests use a 32-bit format.

[0204] The proxy decodes all address bits in the header - address aliasing is not allowed.

[0205] Request handling rules MAY support DMWr. If the device supports being the target of an I / O write request (i.e., a non-posted request), then each associated completion shall be returned within the same time limit as if a posted request were accepted. If the device supports being the target of a deferrable memory write request, then each associated completion shall be returned within the same time limit as if a "posted request" were accepted.

[0206] Flow control rules can also support DMWr. Each virtual channel has independent flow control. Flow control distinguishes three types of TLPs:

[0207] Posted Request (P) - messages and memory writes;

[0208] Non-posted requests (NP) - all reads, I / O writes, configuration writes, AtomicOps, and deferrable memory writes; and

[0209] Completion count (Cpl) - associated with the corresponding NP request

[0210] Furthermore, flow control distinguishes the following types of TLP information in each of three types: header (H); and data (D).

[0211] Therefore, there are six types of information tracked by flow control for each virtual channel, as shown in the flow control credit types of Table 4.

[0212] Table 4. Flow control credit types.

[0213]

[0214] TLP consumes flow control credits, as shown in Table 5 TLP flow control credit consumption.

[0215] Table 5. TLP flow control credit consumption.

[0216]

[0217] Table 6 shows the minimum announcement for NPD credit types.

[0218] Table 6. Minimum initial flow control advertisements.

[0219]

[0220] For errors detected in the transaction layer and uncorrectable internal errors, it is permitted and recommended that no more than one error be reported for a single received TLP, using the following priority levels (highest to lowest):

[0221] Uncorrectable internal error

[0222] Receiver overflow;

[0223] Malformed TLP;

[0224] ECRC check failed;

[0225] AtomicOp or DMWr exit is blocked;

[0226] The TLP prefix is ​​blocked;

[0227] ACS;

[0228] MC blocked TLP;

[0229] Unsupported Request (UR), Completion Aborted (CA), or Unexpected Completion; and

[0230] A poisoned TLP has been received or the poisoned TLP export is blocked

[0231] Table 7 shows the transaction layer error list. The transaction layer error list can be used by a detection agent (eg, an egress port) to identify errors when processing DMWr packets.

[0232] Table 7. List of transaction layer errors.

[0233]

[0234] Table 8 provides the bit mapping to the Device Capability 2 register:

[0235] Table 8. Device Capability 2 Registers.

[0236]

[0237] Table 9 provides the bit mapping to the Device Capability 4 register:

[0238] Table 9. Device Capability 3 Register.

[0239]

[0240] Table 10 shows the definition of the uncorrectable error status register bit values.

[0241] Table 10. Uncorrectable Error Status Register.

[0242]

[0243] Table 11 shows the definition of the uncorrectable error mask register bit values.

[0244] Table 11. Uncorrectable Error Mask Register.

[0245]

[0246] Table 12 shows the definition of the uncorrectable error severity register bit values.

[0247] Table 12. Uncorrectable Error Severity Register.

[0248]

[0249] In an embodiment, AMWr can be used for non-posted write requests to make work requests scalable to completers. The following features can facilitate AMWr messaging over PCIe-based interconnects:

[0250] Definition of the transaction layer packet (TLP) type "AMWr" (e.g., using the value 0b11011, which was previously used for the deprecated TCfgWr TLP type);

[0251] Extension of the CRS completion condition to apply to other non-configuration packets (referred to herein as the "Request Retry Condition (RRS)");

[0252] Add AMWr routing and egress blocking support for PCIe ports;

[0253] Adds AMWr requester and completer support for endpoints and root complexes; and

[0254] Added AMWr exit blocking reporting to Advanced Error Reporting (AER).

[0255] Acknowledged Memory Writes (AMWr) are non-posted requests and enable simple scalability of shared work queues or acknowledged doorbell writes for high-performance, low-latency devices such as accelerators. The number of unique application clients that can submit work to such devices typically depends on the number of queues and doorbells supported by the device. With AMWr, accelerator-type (and other I / O) devices can have a single shared work queue and accept work items from multiple non-cooperating software agents in a non-blocking manner. Devices can expose such a shared work queue in an implementation-specific manner.

[0256] As with other PCIe transactions, support for peer routing of AMWr requests and completions between root ports is optional and implementation dependent. If a root complex (RC) supports AMWr routing capability between two or more root ports, the RC indicates the capability in each associated root port via the AMWr Routing Support bit in the Device Capability 2 Register.

[0257] If a switch supports AMWr routing for any of its ports, it can support AMWr routing for all of those ports.

[0258] For a switch or RC, when AMWr egress blocking is enabled in an egress port and an AMWr request target is issued from that egress port, the egress port shall error the request as an unsupported request and must also return a completion with a completion status of UR.

[0259] In an embodiment, for a switch or RC, when AMWr egress blocking is enabled in an egress port and the AMWr request target goes out of the egress port, the egress port handles the request as an AMWr egress blocking error and returns a completion with a completion status of UR. If the severity of the AMWr egress blocking error is not fatal, this situation must be handled as an advisory non-fatal error.

[0260] The RC is not required to support AMWr routing between all Root Port pairs that have the AMWr Routing Support bit set. AMWr requests requesting routing between unsupported Root Port pairs MUST be handled as Unsupported Requests (UR) and reported by the "sending" port. If the RC supports AMWr routing capability between two or more Root Ports, the capability MUST be indicated in each associated Root Port via the "AW Routing Support" bit in the Device Capability 2 Register. Software MUST NOT assume that AMWr routing is supported between all Root Port pairs that have the AMWr Routing Support bit set.

[0261] A completer that supports AMWr requests (e.g., an entity that completes an AMWr request) is required to process a well-formed AMWr request as a successful completion (SC), a request retry condition (RRS), an unsupported request (UR), or a completer abort (CA) for any location in its target memory space. The following characteristics apply to AMWr completers:

[0262] Unless a higher priority error exists, an AMWr-aware completer shall treat a poisoned or corrupted AMWr request as an error that a poisoned TLP was received and return a completion with a completion status of Unsupported Request (UR). The value of the destination location MUST remain unchanged.

[0263] If the completer of an AMWr request encounters an uncorrectable error while accessing the target location or performing an acknowledged write, the completer MUST handle this as a Completer Abort (CA). The subsequent state of the target location is implementation specific.

[0264] An AMWr-aware completer is required to treat any well-formed AMWr request with a type or operand size it does not support as an Unsupported Request (UR).

[0265] If any function in a multifunction device supports AMWr completer or AMWr routing capabilities, all functions in the device with memory space BARs can decode well-formed AMWr requests and treat any supporting functions as unsupported requests (UR). Note that in such devices, functions lacking AMWr completer capabilities are prohibited from treating well-formed AMWr requests as malformed TLPs.

[0266] If the RC has any root ports that support AMWr routing capabilities, then all RCiEPs in the RC that are reachable by the forwarded AMWr request must decode the well-formed AMWr request and process anything they do not support as an Unsupported Request (UR).

[0267] For an AMWr request with a supported type and operand size, an AMWr-aware completer is required to either execute the request or handle it as a Completer Abort (CA) for any location in its target memory space. Completers are permitted to support AMWr requests on a subset of their target memory space as required by their programming model (see Section 2.3.1, "Request Handling Rules"). PCI Express-defined or inherited memory space structures (such as the MSI-X table structures) are not required to be supported as AMWr targets unless explicitly stated in the structure description.

[0268] Implementing AMWr completer support is optional.

[0269] In some embodiments, devices and device drivers may use AMWr to implement control mechanisms.

[0270] As a non-posted request, AMWr TLP uses a completion TLP to complete the transaction. In addition, PCIe ordering rules stipulate that non-posted TLPs cannot pass through posted TLPs, making posted transactions more preferred for improved performance.

[0271] The grouping definitions, which may represent changes to the PCIe specification, may be interpreted at least in part based on the following definitions:

[0272] Memory transactions include the following types:

[0273] Read request / completion;

[0274] Write request;

[0275] Confirmed write request (add AMWr packet here);

[0276] AtomicOp request / completion; etc.

[0277] Table 13 below is a representation of Tables 2-3 of the PCIe specification. Here, Table 13 includes TLP definitions that support AMWr.

[0278] Table 13. TLP definitions supporting AMWr.

[0279]

[0280] Table 14 below provides the definition of completion status.

[0281] Table 14. Definition of completion status.

[0282]

[0283]

[0284] Various other changes may be made to the PCIe specification to facilitate AMWr messaging for scalable workflow submission.

[0285] The following registers and bit values ​​can be defined to support AMWr:

[0286] In a Device Control 3 register (eg, as defined in the PCIe specification), Table 15 may provide definitions or register bits.

[0287] Table 15. Device Control 3 Registers.

[0288]

[0289] Table 16 shows the definition of the uncorrectable error status register bit values.

[0290] Table 16. Uncorrectable Error Status Register.

[0291]

[0292] Table 17 shows the definition of the uncorrectable error mask register bit values.

[0293] Table 17. Uncorrectable Error Mask Register.

[0294]

[0295] Table 18 shows the definition of the uncorrectable error severity register bit values.

[0296] Table 18. Uncorrectable Error Severity Register.

[0297]

[0298] Fig.10 An embodiment of a computing system including an interconnect architecture is shown. Fig.10 , shows an embodiment of a structure consisting of point-to-point links interconnecting a set of components. System 1000 includes a processor 1005 and a system memory 1010 coupled to a controller hub 1015. Processor 1005 includes any processing element, such as a microprocessor, a main processor, an embedded processor, a coprocessor, or other processor. Processor 1005 is coupled to controller hub 1015 via a front-end bus (FSB) 1006. In one embodiment, FSB 1006 is a serial point-to-point interconnect as described below. In another embodiment, link 1006 includes a serial, differential interconnect architecture that conforms to different interconnect standards.

[0299] System memory 1010 includes any memory device, such as random access memory (RAM), non-volatile (NV) memory, or other memory accessible to devices in system 1000. System memory 1010 is coupled to controller hub 1015 through memory interface 1016. Examples of memory interfaces include double data rate (DDR) memory interfaces, dual channel DDR memory interfaces, and dynamic RAM (DRAM) memory interfaces.

[0300] In one embodiment, controller hub 1015 is a root hub, root complex or root controller in a Peripheral Component Interconnect Express (PCIe or PCIE) interconnect hierarchy. Examples of controller hub 1015 include chipsets, memory controller hubs (MCH), north bridges, interconnect controller hubs (ICH), south bridges and root controllers / hubs. The term chipset typically refers to two physically separate controller hubs, i.e., a memory controller hub (MCH) coupled to an interconnect controller hub (ICH). Note that current systems typically include an MCH integrated with processor 1005, and controller 1015 will communicate with I / O devices in a similar manner as described below. In some embodiments, peer-to-peer routing is optionally supported through root complex 1015.

[0301] Here, controller hub 1015 is coupled to switch / bridge 1020 via serial link 1019. Input / output modules 1017 and 1021 (also referred to as interfaces / ports 1017 and 1021) include / implement a layered protocol stack to provide communications between controller hub 1015 and switch 1020. In one embodiment, multiple devices can be coupled to switch 1020.

[0302] Switch / bridge 1020 routes packets / messages from device 1025 from upstream (i.e., upward toward the hierarchy of the root complex) to controller hub 1015, and from processor 1005 or system memory 1010 downstream (i.e., away from the hierarchy of the root controller) to device 1025. In one embodiment, switch 1020 is referred to as a logical component of multiple virtual PCI-PCI bridge devices. Device 1025 includes any internal or external device or component to be coupled to an electronic system, such as an I / O device, a network interface controller (NIC), an add-on card, an audio processor, a network processor, a hardware driver, a storage device, a CD / DVD ROM, a monitor, a printer, a mouse, a keyboard, a router, a portable storage device, a FireWire device, a Universal Serial Bus (USB) device, a scanner, and other input / output devices. In PCIe local devices (e.g., devices), it is usually referred to as an endpoint. Although not specifically shown, device 1025 may include a PCIe to PCI / PCI-X bridge to support traditional or other versions of PCI devices. Endpoint devices in PCIe are generally classified as legacy, PCIe, or root complex endpoints.

[0303] Graphics accelerator 1030 is also coupled to controller hub 1015 via serial link 1032. In one embodiment, graphics accelerator 1030 is coupled to MCH, which is coupled to ICH. Switch 1020 and corresponding I / O devices 1025 are then coupled to the ICH. I / O modules 1031 and 1018 are also used to implement a layered protocol stack to communicate between graphics accelerator 1030 and controller hub 1015. Similar to the MCH discussed above, the graphics controller or graphics accelerator 1030 itself can be integrated into processor 1005.

[0304] Fig.11 An embodiment of an interconnect architecture including a layered stack is shown. Fig.11 , shows an embodiment of a layered protocol stack. The layered protocol stack 1100 includes any form of layered communication stack, such as a Quick Path Interconnect (QPI) stack, a PCie stack, a next generation high performance computing interconnect stack, or other layered stack. Although the following reference Figures 10 to 13 The discussion is related to the PCIe stack, but the same concepts can be applied to other interconnect stacks. In one embodiment, the protocol stack 1100 is a PCIe protocol stack, which includes a transaction layer 1105, a link layer 1110, and a physical layer 1120. Interfaces such as Figure 1 The interfaces 1017, 1018, 1021, 1022, 1026, and 1031 in may be represented as a communication protocol stack 1100. What is represented as a communication protocol stack may also be referred to as a module or interface that implements / includes the protocol stack.

[0305] PCI Express uses packets to pass information between components. Packets are formed in the Transaction Layer 1105 and Data Link Layer 1110 to carry information from a sending component to a receiving component. As the transmitted packets flow through the other layers, they are expanded with additional information required to process the packets at those layers. On the receiving side, the reverse process occurs, and the packet is converted from its Physical Layer 1120 representation to a Data Link Layer 1110 representation and finally (for a Transaction Layer packet) to a form that can be processed by the Transaction Layer 1105 of the receiving device.

[0306] Transaction Layer

[0307] In one embodiment, the transaction layer 1105 will provide an interface between the processing core of the device and the interconnect architecture (e.g., the data link layer 1110 and the physical layer 1120). In this regard, the primary responsibility of the transaction layer 1105 is the assembly and disassembly of packets (i.e., transaction layer packets or TLPs). The translation layer 1105 typically manages credit flow control for TLPs. PCIe implements split transactions, i.e., transactions where requests and responses are separated in time, allowing the link to carry other traffic while the target device is collecting data for the response.

[0308] Additionally, PCIe utilizes credit-based flow control. In this scheme, a device advertises an initial credit limit for each receive buffer in the transaction layer 1105. External devices at the opposite end of the link, such as Figure 1 The controller hub 115 in the embodiment of the present invention calculates the number of credits consumed by each TLP. If the transaction does not exceed the credit limit, the transaction can be transmitted. After receiving the response, the credit limit will be restored. The advantage of the credit plan is that as long as the credit limit is not encountered, the delay of credit return will not affect the performance.

[0309] In one embodiment, the four transaction processing address spaces include a configuration address space, a memory address space, an input / output address space, and a message address space. Memory space transactions include one or more of read requests and write requests to transfer data to / from a memory mapped location. In one embodiment, memory space transactions can use two different address formats, for example, a short address format (e.g., a 32-bit address) or a long address format (e.g., a 64-bit address). Configuration space transactions are used to access the configuration space of a PCIe device. Transactions to the configuration space include read requests and write requests. Message space transactions (or simply messages) are defined to support in-band communication between PCIe agents.

[0310] Thus, in one embodiment, transaction layer 1105 assembles a packet header / payload 1106. The current format of the packet header / payload can be found in the PCIe specification at the PCIe specification website.

[0311] Quick Reference Fig.12 : Fig.12 An embodiment of a request or packet to be generated or received within an interconnect architecture is shown. An embodiment of a PCIe transaction descriptor is shown. In one embodiment, the transaction descriptor 1200 is a mechanism for carrying transaction information. In this regard, the transaction descriptor 1200 supports the identification of transactions in the system. Other potential uses include tracking modifications to the default transaction order and the association of transactions with channels.

[0312] Transaction descriptor 1200 includes a global identifier field 1202, an attribute field 1204, and a channel identifier field 1206. In the example shown, global identifier field 1202 is depicted including a local transaction identifier field 1208 and a source identifier field 1210. In one embodiment, global transaction identifier 1202 is unique for all outstanding requests.

[0313] According to one implementation, the local transaction identifier field 1208 is a field generated by the requesting agent and is unique to all outstanding requests that the requesting agent is required to complete. In addition, in this example, the source identifier 1210 uniquely identifies the requester agent within the PCIe hierarchy. Therefore, the local transaction identifier 1208 field, together with the source ID 1210, provides a global identification of transactions within the hierarchy domain.

[0314] Attribute fields 1204 specify the characteristics and relationships of a transaction. In this regard, attribute fields 1204 are potentially used to provide additional information that allows modification of the default handling of a transaction. In one embodiment, attribute fields 1204 include a priority field 1212, a reserved field 1214, an order field 1216, and a non-snooping field 1218. Here, priority subfield 1212 can be modified by the initiator to assign a priority to a transaction. Reserved attribute fields 1214 are reserved for future use or vendor-defined usage. Possible usage models that use priority or security attributes can be implemented using reserved attribute fields.

[0315] In this example, the sort attribute field 1216 is used to provide optional information that conveys the type of sorting that can modify the default sorting rules. According to an example implementation, a sort attribute of "0" indicates that the default sorting rules are to be applied, where a sort attribute of "1" indicates a loose sorting, where writes can be passed in the same direction as writes, and read completions can be passed in the same direction as writes. The snoop attribute field 1218 is used to determine whether to snoop the transaction. As shown, the channel ID field 1206 identifies the channel associated with the transaction.

[0316] Link Layer

[0317] The link layer 1110, also referred to as the data link layer 1110, acts as an intermediate stage between the transaction layer 1105 and the physical layer 1120. In one embodiment, the responsibility of the data link layer 1110 is to provide a reliable mechanism for exchanging transaction layer packets (TLPs) between two components in a link. One side of the data link layer 1110 accepts the TLPs assembled by the transaction layer 1105, applies a packet sequence identifier 1111 (i.e., an identification number or packet number), calculates and applies an error detection code (i.e., CRC 1112), and submits the modified TLP to the physical layer 1120 for transmission across the physical device to an external device.

[0318] Physical Layer

[0319] In one embodiment, the physical layer 1120 includes a logical sub-block 1121 and an electronic block 1122 to physically send packets to an external device. Here, the logical sub-block 1121 is responsible for the "digital" functions of the physical layer 1121. In this regard, the logical sub-block includes a transmit portion that prepares outgoing information for transmission by the physical sub-block 1122 and a receiver portion that identifies and prepares received information before passing it to the link layer 1110.

[0320] Physical block 1122 includes a transmitter and a receiver. Logical sub-block 1121 provides symbols to the transmitter, which serializes and sends them to an external device. Serialized symbols are provided to the receiver from the external device and the received signal is converted into a bit stream. The bit stream is deserialized and provided to logical sub-block 1121. In one embodiment, an 8b / 10b transmission code is used, in which ten-bit symbols are sent / received. Here, special symbols are used to form packets in frames 1123. In addition, in one example, the receiver also provides a symbol clock recovered from the incoming serial stream.

[0321] As described above, although the transaction layer 1105, link layer 1110, and physical layer 1120 are discussed with reference to a specific embodiment of the PCIe protocol stack, the layered protocol stack is not limited thereto. In fact, any layered protocol may be included / implemented. As an example, a port / interface represented as a layered protocol includes: (1) a first layer for assembling packets, i.e., a transaction layer; a second layer for sequencing packets, i.e., a link layer; and a third layer for transmitting packets, i.e., a physical layer. As a specific example, a common standard interface (CSI) layered protocol is used.

[0322] Fig.13An embodiment of a transmitter and receiver pair for an interconnect architecture is shown. An embodiment of a PCIe serial point-to-point structure is shown. Although an embodiment of a PCIe serial point-to-point link is shown, the serial point-to-point link is not limited thereto because it includes any transmission path for transmitting serial data. In the illustrated embodiment, the basic PCIe link includes two, low-voltage, differential drive signal pairs: a transmit pair 1306 / 1311 and a receive pair 1312 / 1307. Therefore, device 1305 includes a transmission logic 1306 for sending data to device 1310 and a receive logic 1307 for receiving data from device 1310. In other words, two transmit paths, namely paths 1316 and 1317, and two receive paths, namely paths 1318 and 1319, are included in the PCIe link.

[0323] A transmission path refers to any path for transmitting data, such as a transmission line, copper wire, optical fiber, wireless communication channel, infrared communication link, or other communication path. A connection between two devices (e.g., device 1305 and device 1310) is called a link (e.g., link 415). A link can support one channel - each channel represents a set of differential signal pairs (one pair for transmission and one pair for reception). To expand bandwidth, a link can aggregate multiple channels represented by xN, where N is any supported link width, such as 1, 2, 4, 8, 12, 16, 32, 64, or wider.

[0324] A differential pair refers to two transmission paths, such as lines 1316 and 1317, for transmitting differential signals. For example, when line 1316 switches from a low voltage level to a high voltage level, i.e., a rising edge, line 1317 drives from a high logic level to a low logic level, i.e., a falling edge. Differential signals potentially exhibit better electrical characteristics, such as better signal integrity, i.e., cross-coupling, voltage overshoot / undershoot, ringing, etc. This allows for better timing windows, thereby achieving faster transmission frequencies.

[0325] Fig.14 Another embodiment of a block diagram of a computing system including a processor is shown. Fig.14 , shows a block diagram of an exemplary computer system formed by a processor including an execution unit for executing instructions, wherein one or more interconnections implement one or more features according to an embodiment of the present invention. System 1400 includes components according to the present disclosure, such as processor 1402, which employs an execution unit including logic according to the present invention to execute an algorithm for process data. System 1400 represents a system based on PENTIUM III available from Intel Corporation in Santa Clara, California. TM 、PENTIUM 4 TM , Xeon TM , Itanium, XScaleTM and / or StrongARM TM 1400 is a processing system based on a microprocessor, but other systems may be used (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.). In one embodiment, the exemplary system 1400 executes WINDOWS 1000 available from Microsoft Corporation of Redmond, Washington. TM The operating system may be a version of the UNIX® operating system, but other operating systems (eg, UNIX and Linux), embedded software and / or graphical user interfaces may also be used. Therefore, embodiments of the present invention are not limited to any specific combination of hardware circuitry and software.

[0326] Embodiments are not limited to computer systems. Alternative embodiments of the present invention may be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet protocol devices, digital cameras, personal digital assistants (PDAs), and handheld PCs. Embedded applications may include microcontrollers, digital signal processors (DSPs), systems on chips, network computers (NetPCs), set-top boxes, network hubs, wide area network (WAN) switches, or any other system that can execute one or more instructions according to at least one embodiment.

[0327] In the illustrated embodiment, the processor 1402 includes one or more execution units 1408 to implement an algorithm that will execute at least one instruction. An embodiment may be described in the context of a single processor desktop or server system, but alternative embodiments may be included in a multi-processor system. System 1400 is an example of a "hub" system architecture. Computer system 1400 includes a processor 1402 to process data signals. As an illustrative example, processor 1402 includes a complex instruction set computer (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor that implements an instruction set combination, or, for example, any other processor device, such as a digital signal processor. Processor 1402 is coupled to a processor bus 1410 that transmits data signals between processor 1402 and other components in system 1400. The elements of system 1400 (e.g., graphics accelerator 1412, memory controller hub 1416, memory 1420, I / O controller hub 1424, wireless transceiver 1426, flash BIOS 1428, network controller 1434, audio controller 1436, serial expansion port 1438, I / O controller 1440, etc.) perform conventional functions well known to those skilled in the art.

[0328] In one embodiment, processor 1402 includes a level 1 (L1) internal cache 1404. Depending on the architecture, processor 1402 may have a single internal cache or multiple levels of internal cache. Other embodiments include a combination of both internal and external caches depending on the specific implementation and requirements. Register file 1406 is used to store various types of data in various registers, including integer registers, floating point registers, vector registers, bank registers, shadow registers, checkpoint registers, status registers, and instruction pointer registers.

[0329] Execution unit 1408, including logic to perform integer and floating point operations, is also located in processor 1402. In one embodiment, processor 1402 includes a microcode (ucode) ROM for storing microcode, which, when executed, will execute certain macro instructions or algorithms to handle complex scenarios. Here, the microcode may be updateable to handle logic errors / fixes of processor 1402. For one embodiment, execution unit 1408 includes logic for processing a packaged instruction set 1409. By including a packaged instruction set 1409 in the instruction set 1402 of a general-purpose processor, along with associated circuitry for executing the instructions, operations used by many multimedia applications can be performed using packaged data in a general-purpose processor 1402. Therefore, many multimedia applications are accelerated and executed more efficiently by using the full width of the processor's data bus for performing operations on packaged data. This potentially eliminates the need to transfer smaller units of data on the processor's data bus to perform one or more operations (one data element at a time).

[0330] Alternative embodiments of execution unit 1408 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. System 1400 includes memory 1420. Memory 1420 includes a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, or other storage device. Memory 1420 stores instructions and / or data represented by data signals to be executed by processor 1402.

[0331] Note that any of the foregoing features or aspects of the present invention may be Fig.14For example, an on-die interconnect (ODI) (not shown) for coupling internal units of processor 1402 implements one or more aspects of the above invention. Or the present invention is associated with the following: processor bus 1410 (e.g., Intel Quick Path Interconnect (QPI) or other known high-performance computing interconnects), high-bandwidth memory path 1418 to memory 1420, point-to-point link to graphics accelerator 1412 (e.g., Peripheral Component Interconnect Express (PCIe) compatible fabric), controller hub interconnect 1422, I / O or other interconnects (e.g., USB, PCI, PCIe) for coupling other components shown. Some examples of such components include an audio controller 1436, a firmware hub (flash BIOS) 1428, a wireless transceiver 1426, data storage 1424, a conventional I / O controller 1410 including a user input and keyboard interface 1442, a serial expansion port 1438 (e.g., a universal serial bus (USB), and a network controller 1434. The data storage device 1424 may include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0332] Fig.15 Another embodiment of a block diagram of a computing system is shown. Fig.15 , shows an embodiment of a system on chip (SOC) design according to the present invention. As a specific illustrative example, SOC 1500 is included in a user equipment (UE). In one embodiment, UE refers to any device that an end user wants to use to communicate, such as a handheld phone, a smart phone, a tablet computer, an ultra-thin notebook, a notebook with a broadband adapter, or any other similar communication device. UE is typically connected to a base station or node, which may essentially correspond to a mobile station (MS) in a GSM network.

[0333] Here, SOC 1500 includes two cores - 1506 and 1507. Similar to the above discussion, cores 1506 and 1507 may conform to an instruction set architecture, such as based on Architecture Core TM , Advanced MicroDevices, Inc. (AMD) processors, MIPS-based processors, ARM-based processor designs, or their customers, and their licensees or adopters. Cores 1506 and 1507 are coupled to cache control 1508 associated with bus interface unit 1509 and L2 cache 1510 to communicate with other parts of system 1500. Interconnect 1510 includes an on-chip interconnect, such as IOSF, AMBA, or other interconnects discussed above, which may implement one or more aspects of the described invention.

[0334] Interface 1510 provides a communication channel to other components, such as a subscriber identity module (SIM) 1530 that interfaces with a SIM card, a boot rom 1535 for storing boot code executed by the kernels 1506 and 1507 to initialize and boot the SOC 1500. An SDRAM controller 1540 that interfaces with external memory (e.g., DRAM 1560), a flash controller 1545 that interfaces with non-volatile memory (e.g., flash memory 1565), a peripheral control Q1650 (e.g., a serial peripheral interface) that interfaces with peripherals, a video codec 1520 and a video interface 1525 to display and receive input (e.g., touch-enabled input), a GPU 1515 that performs graphics-related calculations, etc. Any of these interfaces can be combined with aspects of the invention described herein.

[0335] In addition, the system shows peripheral devices for communication, such as Bluetooth module 1570, 3G modem 1575, GPS 1585 and WiFi 1585. As mentioned above, it is noted that the UE includes a radio device for communication. As a result, not all of these peripheral communication modules are required. However, in the UE, some form of radio for external communication will be included.

[0336] While the present invention has been described with respect to a limited number of embodiments, those skilled in the art will appreciate numerous modifications and variations therefrom. It is intended that the appended claims cover all such modifications and variations that fall within the true spirit and scope of the present invention.

[0337] A design may go through various stages from creation to simulation to manufacturing. Data representing a design can represent the design in a variety of ways. First, as useful in simulation, a hardware description language or another functional description language can be used to represent the hardware. In addition, a circuit-level model with logic and / or transistor gates can be generated at some stage of the design process. In addition, most designs reach a data level representing the physical placement of various devices in the hardware model at some stage. In the case of using traditional semiconductor manufacturing technology, the data representing the hardware model can be data that specifies the presence or absence of various features on different mask layers of a mask used to manufacture an integrated circuit. In any representation of the design, the data can be stored in any form of machine-readable medium. A memory or a magnetic or optical storage such as a disk can be a machine-readable medium for storing information transmitted via light waves or electric waves, which are modulated or otherwise generated to transmit such information. When an electric carrier wave indicating or carrying a code or design is sent, a new copy can be made as long as the extent of copying, buffering or retransmitting the electric signal is performed. Therefore, a communication provider or network provider can store an article embodying the technology of an embodiment of the present invention, such as information encoded as a carrier wave, at least temporarily on a tangible, machine-readable medium.

[0338] As used herein, a module refers to any combination of hardware, software, and / or firmware. As an example, a module includes hardware associated with a non-transitory medium, such as a microcontroller, to store code suitable for execution by the microcontroller. Therefore, in one embodiment, a reference to a module refers to hardware specifically configured to identify and / or execute code to be stored on a non-transitory medium. In addition, in another embodiment, the use of a module refers to a non-transitory medium including a code that is particularly suitable for execution by a microcontroller to perform a predetermined operation. And it can be inferred that in yet another embodiment, the term module (in this example) may refer to a combination of a microcontroller and a non-transitory medium. Typically, the boundaries of modules shown as separate typically change and may overlap. For example, a first and a second module may share hardware, software, firmware, or a combination thereof, while potentially retaining some independent hardware, software, or firmware. In one embodiment, the use of the term "logic" includes hardware such as transistors, registers, or other hardware such as programmable logic devices.

[0339] In one embodiment, the use of the phrase "to" or "configured to" refers to a device, hardware, logic, or element that is arranged, combined, manufactured, offered for sale, imported, and / or designed to perform a specified or determined task. In this example, a device or element thereof that is still not in operation is still "configured to" perform the specified task if it is designed, coupled, and / or interconnected to perform the specified task. As an illustrative example only, a logic gate can provide a 0 or 1 during operation. But "a logic gate that provides an enable signal to a clock is configured to" does not include every potential logic gate that may provide a 1 or 0. Instead, the logic gate is coupled in a certain way, that is, a 1 or 0 output is an enabled clock during operation. Note again that the use of the term "configured to" does not require operation, but focuses on the potential state of the device, hardware, and / or element, where when the device, hardware, and / or element is operating, in the potential state, the device, hardware, and / or element is designed to perform a specific task.

[0340] Additionally, in one embodiment, the use of the phrases "capable of" and / or "operable" refers to certain devices, logic, hardware, and / or elements that are designed to be able to use the device, logic, hardware, and / or elements in a specific manner. As described above, in one embodiment, "for", "capable of" or "operable" refers to a potential state of a device, logic, hardware, and / or element, wherein the device, logic, hardware, and / or element is incapable of functioning, but is designed in a manner that allows the device to be used in a specified manner.

[0341] As used herein, a value includes any known representation of a number, state, logic state, or binary logic state. Typically, values ​​using logic levels, logic values, or logic are also referred to as 1 and 0, which simply represent binary logic states. For example, 1 represents a logic high level, and 0 represents a logic low level. In one embodiment, a storage unit such as a transistor or a flash memory cell may be able to store a single logic value or multiple logic values. However, other representations of values ​​in computer systems have been used. For example, the decimal number ten may also be represented as a binary value 1010 and the hexadecimal letter A. Therefore, the value includes any representation of information that can be stored in a computer system.

[0342] In addition, a state can be represented by a value or a portion of a value. As an example, a first value, such as a logical 1, can represent a default or initial state, while a second value, such as a logical 0, can represent a non-default state. Additionally, in one embodiment, the terms "reset" and "set" refer to a default value and an updated value or state, respectively. For example, a default value may include a high logic value, i.e., reset, and an updated value may include a low logic value, i.e., set. Note that any combination of values ​​can be used to represent any number of states.

[0343] Embodiments of the above methods, hardware, software, firmware or code may be implemented by instructions or code stored on a machine-accessible, machine-readable, computer-accessible or computer-readable medium that can be executed by a processing element. Non-transitory machine-accessible / readable media include any mechanism that provides (i.e., stores and / or transmits) information in a form that is readable by a machine such as a computer or electronic system. For example, non-transitory machine-accessible media include random access memory (RAM), such as static RAM (SRAM) or dynamic RAM (DRAM); ROM; magnetic or optical storage media; flash memory devices; electrical storage devices; optical storage devices; acoustic storage devices; other forms of storage devices for storing information received from transient (propagating) signals (e.g., carrier waves, infrared signals, digital signals); etc., which will be distinguished from non-transitory media from which information can be received.

[0344] Instructions for programming logic to perform embodiments of the present invention may be stored in a memory in the system, such as a DRAM, cache, flash memory, or other memory. In addition, instructions may be distributed via a network or by other computer-readable media. Therefore, a machine-readable medium may include any mechanism for storing or transmitting information in a machine (e.g., computer) readable form, but is not limited to a floppy disk, an optical disk, a compressed disk, a read-only memory (CD-ROM) and a magneto-optical disk, a read-only memory (ROM), a random access memory (RAM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic card or an optical card, a flash memory, or a tangible machine-readable storage for transmitting information through the Internet by other forms of electrical, optical, acoustic, or propagating signals (e.g., carrier waves, infrared signals, digital signals, etc.). Therefore, a computer-readable medium includes any type of tangible machine-readable medium suitable for storing or transmitting electronic instructions or information in a machine (e.g., computer) readable form.

[0345] Throughout the specification, references to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, the phrases "in one embodiment" or "in an embodiment" appearing in various places throughout the specification are not necessarily all referring to the same embodiment. Furthermore, in one or more embodiments, the particular features, structures, or characteristics may be combined in any suitable manner.

[0346] In the foregoing description, detailed descriptions have been given with reference to specific exemplary embodiments. However, it will be apparent that various modifications and changes may be made thereto without departing from the broader spirit and scope of the invention as set forth in the appended claims. Therefore, the description and the accompanying drawings should be considered illustrative rather than restrictive. In addition, the foregoing use of embodiments and other exemplary language does not necessarily refer to the same embodiment or the same example, but may refer to different and distinct embodiments as well as potentially the same embodiments.

[0347] The system, method, and apparatus may include one or a combination of the following examples:

[0348] Example 1 is a device including a controller and a command queue for buffering incoming write requests into the device. The controller is configured to receive a non-posted write request in a transaction layer packet (TLP) from a client to the command queue via a link; determine that the command queue can accept the non-posted write request; generate a completion message with a completion status bit set to successfully completed (SC); and send a completion message to the client via the link indicating that the non-posted write request has been accepted into the command queue.

[0349] Example 2 may include the subject matter of Example 1, wherein the non-posted write request comprises a deferred memory write (DMWr) request message.

[0350] Example 3 may include the subject matter of any of Examples 2-3, wherein the controller is used to receive a second non-posted write request in a second TLP; determine that the command queue is full; generate a completion message having a completion status bit set to a memory request retry status (MRS); and send a completion message to the client via a link indicating that the non-posted write request has been rejected into the command queue.

[0351] Example 4 may include the subject matter of Example 3, wherein the controller is used to receive a third TLP, the third TLP including a retry of a second non-posted write request; determine that the command queue can accept the second non-posted write request; generate a completion message with the completion status bit set to successfully completed (SC); and send a completion message to the client via the link indicating that the non-posted write request has been accepted into the command queue.

[0352] Example 5 may include the subject matter of any of Examples 1-3, wherein the device comprises an input / output device, such as an accelerator.

[0353] Example 6 may include the subject matter of any of Examples 1-5, wherein the non-posted write request is received in a transaction layer packet (TLP) including a type field indicating an acknowledged memory write request.

[0354] Example 7 may include the subject matter of Example 6, wherein the TLP includes a completion field indicating completion of the data without memory write completion for acknowledgement.

[0355] Example 8 may include the subject matter of any of Examples 1-7, wherein the link is based on a Peripheral Component Interconnect Express (PCIe) protocol.

[0356] Example 9 is a system including a host device, the host device including a port located on the host device, the port being used to send a TLP containing a non-posted write request to an I / O device via a link; and an input / output (I / O) device coupled to the port via the link. The I / O device can receive a TLP including a non-posted write request from the link, determine that a command queue has space available to accept the non-posted write request, and generate a completion message with a completion status bit set to a successful completion (SC). The port is used to send a completion message to a client via the link indicating that the non-posted write request has been accepted into the command queue.

[0357] Example 10 may include the subject matter of Example 9, wherein the non-posted write request is a first non-posted write request. A second TLP including a second non-posted write request is sent to the I / O device. The I / O device is used to receive the second non-posted write request, determine that the command queue at the I / O device is full, and generate a completion message with a completion status bit set to a memory request retry status (MRS) for transmission by the port to the client through the link, the completion message indicating that the non-posted write request was rejected into the command queue. The port is used to forward a message indicating the completion message of the non-posted write request to a link partner of the I / O device.

[0358] Example 11 may include the subject matter of any of Example 9, wherein the non-posted write request is a first non-posted write request.

[0359] Example 12 may include the subject matter of Example 11, wherein the port is a first port, the host device includes a second port, the second port includes enabled egress blocking, the second port is used to receive a second non-posted write request from a link partner; and send a completion message including a status of an unsupported request status identifier.

[0360] Example 13 may include the subject matter of Example 11, wherein the I / O device is to receive a second non-posted write request including a corrupted transaction layer packet; and send a completion message including an unsupported request status indicator to the first port.

[0361] Example 14 may include the subject matter of Example 11, wherein the I / O device is to receive a second non-posted write request; determine that the second non-posted write request includes an uncorrectable error; and send a complete abort message to the port.

[0362] Example 15 may include the subject matter of Example 11, wherein the I / O device is to receive a second non-posted write request; determine that the second non-posted write request includes an unsupported type or operand; and send a completion message with an unsupported request status indicator to the port.

[0363] Example 16 may include the subject matter of any of Examples 9-15, wherein the I / O device comprises an accelerator device.

[0364] Example 17 may include the subject matter of any of Examples 9-16, wherein the non-posted write request comprises a deferred memory write (DMWr) request.

[0365] Example 18 may include the subject matter of any of Examples 9-17, wherein the command queue comprises a shared work queue (SWQ), and the non-posted write request is received at the SWQ of the I / O device.

[0366] Example 19 may include the subject matter of any of Examples 9-18, wherein the link is based on a Peripheral Component Interconnect Express (PCIe) protocol.

[0367] Example 20 is a computer-implemented method comprising: receiving a non-posted write request in a transaction layer packet (TLP) from a client to a command queue via a link at an input / output (I / O) device; determining, by a controller of the I / O device, that the command queue can accept the non-posted write request; generating a completion message having a completion status bit set to successfully completed (SC); and sending a completion message to the client via the link indicating that the non-posted write request has been accepted into the command queue.

[0368] Example 21 may include the subject matter of Example 20, wherein the non-posted write request comprises a deferred memory write (DMWr) request message.

[0369] Example 22 may include the subject matter of any of Examples 20-21, and may further include receiving a second non-posted write request in a second TLP at the I / O device; determining, by a controller of the I / O device, that the command queue is full; generating a completion message having a completion status bit set to a memory request retry status (MRS); and sending a completion message to the client over the link indicating that the non-posted write request was rejected into the command queue.

[0370] Example 23 may include the subject matter of Example 22, and may further include receiving a third TLP comprising a retry of a second non-posted write request; determining that the command queue can accept the second non-posted write request; generating a completion message having a completion status bit set to SC; and sending the completion message to the link partner.

[0371] Example 24 may include the subject matter of any of Examples 20-23, and may further include receiving a second non-posted write request; determining that the second non-posted write request contains an uncorrectable error; and sending a complete abort message to the link partner.

[0372] Example 25 may include the subject matter of any of Examples 20-24, and may further include receiving a second non-posted write request; determining that the non-posted write request includes an unsupported type or operand; and sending a completion message with an unsupported request status indicator to the link partner.

[0373] Example 26 is a device comprising a command queue for buffering incoming work requests, and means for determining a delayed memory write (DMWr) request and one or more response messages to the DMWr from a received transaction layer packet based on a status of the command queue.

Claims

1. An apparatus for responding to a deferrable memory write request, comprising: Logic for managing a shared work queue, The logic is used to: receiving a deferrable memory write (DMWr) request in a transaction layer packet (TLP) from an agent, wherein the TLP includes a peripheral component interconnect express (PCIe) TLP type field encoding to indicate the DMWr request; Determining whether the DMWr request can be successfully completed; and If the DMWr request cannot be completed successfully, a request retry status RRS completion status is returned to the agent.

2. The device according to claim 1, wherein: The shared work queue is used to accept work items from multiple non-cooperative agents in a non-blocking manner.

3. The device according to claim 1, wherein: The RRS completion status includes a completion status field message coded as 010.

4. The apparatus of any one of claims 1 to 3, wherein the logic is configured to: receiving a second DMWr request in another TLP from the agent; determining that the second DMWr request includes an uncorrectable error; and The Completer Aborted CA completion status is returned to the Agent.

5. The device according to claim 4, wherein: The CA completion status includes a completion status field message coded as 100.

6. The apparatus of any one of claims 1-3, wherein the logic is configured to: receiving a second DMWr request in a second TLP from the agent; Determining that the second DMWr can be successfully completed; and A successful completion SC completion status is returned to the agent.

7. The device according to claim 6, wherein: The SC Completion Status includes a Completion Status field encoded as 000.

8. The apparatus of any one of claims 1-3, wherein the logic is configured to: receiving a second DMWr request in a transaction layer packet TLP from an agent; determining that the second DMWr request is not supported; and Returns an unsupported request UR completion status to the proxy.

9. The apparatus of any one of claims 1-3, wherein the logic is configured to: receiving a second DMWr request in a transaction layer packet TLP from an agent; determining that the second DMWr request includes a poisoned DMWr; and Returns an unsupported request UR completion status to the proxy.

10. The device according to claim 9, wherein: The UR completion status includes a completion status field value of 001.

11. The device according to any one of claims 1 to 3, wherein: The DMWr request is an atomic, non-posted write request.

12. The device according to any one of claims 1 to 3, wherein: The TLP includes a format encoded as 010 or 011.

13. A method for responding to a deferrable memory write request, comprising: Receive transaction layer packets TLP from the agent; The TLP includes a peripheral component interconnect express (PCIe) type field encoding to indicate that the TLP includes a deferrable memory write request (DMWr); Determine whether the DMWr request can be completed; as well as If the DMWr cannot be completed, a request retry status RRS completion status is returned to the agent.

14. The method according to claim 13, wherein: Determining that the TLP includes a DMWr request includes reading a TLP format field code and a TLP type field code, the format field code includes one of 010 or 011, and the type field code includes 11011.

15. The method according to claim 13 or 14, wherein: The RRS completion status includes a completion status field message coded as 010.

16. The method according to claim 13 or 14, further comprising: If the TLP includes uncorrectable errors, the Completer Abort CA Completion status is returned.

17. The method according to claim 16, wherein: The CA completion status includes a Completion Status field coded as 100.

18. The method according to claim 13 or 14, further comprising: If the DMWr is able to complete, a successful completion SC completion status is returned to the agent.

19. The method according to claim 18, wherein: The SC completion status includes a completion status field message encoded as 000.

20. The method according to claim 13 or 14, further comprising: If the DMWr is not supported, or if the TLP includes a poisoned DMWr, an unsupported request UR completion status is returned.

21. The method according to claim 20, wherein: The UR completion status includes a completion status field value of 001.

22. A computing system comprising: Requester; Completer; The requester includes logic for sending a peripheral component interconnect express (PCIe) transaction layer packet (TLP) including a deferrable memory write (DMWr) request, The PCIe TLP includes a type field encoding to indicate that the TLP includes a DMWr request; and The completer includes logic to: Determine whether the DMWr can be completed, and If the DMWr cannot be completed, a request retry status RRS completion status is returned to the requester.

23. The computing system of claim 22, wherein: The TLP includes a format field code and a type field code to indicate that the TLP includes a DMWr, the format field code includes one of 010 or 011, and the type field code includes 11011.

24. A computing system according to claim 22 or 23, wherein: The RRS completion status includes a completion status field message coded as 010.

25. The computing system of claim 22 or 23, the completer to return a completer aborted CA completion status if the TLP includes an uncorrectable error.

26. The computing system of claim 25, wherein: The CA completion status includes a completion status field message coded as 100.

27. The computing system of claim 22 or 23, the logic of the completer to return a successful completion (SC) completion status to the requester if the DMWr is able to complete.

28. The computing system of claim 27, wherein: The SC completion status includes a completion status field message encoded as 000.

29. The computing system of claim 22 or 23, the completer to return an unsupported request UR completion status if the DMWr is not supported or if the TLP includes a poisoned DMWr.

30. The computing system of claim 29, wherein: The UR completion status includes a completion status field value of 001.

31. A computer program product comprising instructions for implementing the method according to any one of claims 13-21 or realizing the apparatus according to any one of claims 1-12 when the instructions are executed by a processor.

Citation Information

Patent Citations

  • On-chip bus

    EP1544743A2

  • Distributed transaction processing

    US8838534B2