Method, device and system for controlling the performance of unused hardware of a link interface
The method and system optimize power consumption and performance in interconnect interfaces by using a configuration memory to manage hardware buffers, addressing the inefficiencies in existing technologies.
Patent Information
- Application Number
- DE112014006490
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2014-03-20
- Publication Date
- 2025-12-24
- Estimated Expiration
- 2034-03-20
AI Technical Summary
Existing technologies fail to efficiently manage power consumption and performance in the context of interconnects in computer systems, particularly in the context of mobile devices, by not addressing the need for simplified configuration control of link interfaces.
A method and system for managing the performance of interconnect interfaces with a simplified configuration control, involving the use of a configuration memory to manage the state of hardware buffers in link interfaces, thereby optimizing power consumption and performance.
This approach reduces power consumption and enhances performance by dynamically managing the state of hardware buffers in link interfaces, ensuring efficient energy usage and optimal performance.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical field
[0001] This disclosure concerns computer systems and in particular (but not exclusively) the performance management of link interfaces in such systems. State of the art
[0002] US 7,472,299 B2 concerns methods and devices for reducing power consumption in arbiters of interconnect routers. For example, an arbiter can be switched off for a selected number of clock cycles when no arbitration is to be performed for the corresponding buffer.
[0003] US 8 379 659 B2 concerns a heterogeneous router microarchitecture. For example, one method involves comparing the utilization level of a buffer in a router port with a threshold and controlling the port to operate at least partially based on the comparison at a first voltage and frequency, with at least one other port of the router being controlled to operate at a second voltage and frequency. Summary of the invention
[0004] It can be considered an object of the present invention to propose devices, methods and systems for managing the performance of interconnect interfaces with a simplified configuration control.
[0005] The foregoing problem is solved according to the invention with the device according to main claim 1, the method according to dependent claim 10, the device according to dependent claim 15, and the system according to dependent claim 21. The dependent claims define further developments of the solutions according to the invention. Brief description of the drawings Fig. Figure 1 is an embodiment of a block diagram for a computer system that includes a multi-core processor. Fig. 2 is an embodiment of a Fabric consisting of point-to-point links that connect a set of components. Fig. Figure 3 is an embodiment of a shift log stack. Fig. 4 is an embodiment of a PCle transaction descriptor. Fig. 5 is an embodiment of a serial point-to-point PCB fabric. Fig. Figure 6 is a block diagram of a SoC design according to one embodiment. Fig. Figure 7 is a block diagram of a system according to an embodiment of the present invention. Fig. Figure 8 is a flowchart of a configuration procedure according to an embodiment of the present invention. Fig. Figure 9A is a block diagram of a configuration memory according to one embodiment. Fig. Figure 9B is a block diagram of a section of a voltage control circuit according to one embodiment. Fig. Figure 10 is a block diagram of a section of a system according to one embodiment. Detailed description
[0006] The following description sets forth numerous specific details, such as examples of specific types of processors and system configurations, specific hardware structures, specific architectural and microarchitectural details, specific register configurations, specific instruction types, specific system components, specific measurements / altitudes, specific processor pipeline stages and operation, etc., to provide a comprehensive understanding of the present invention. However, it will be apparent to a person skilled in the art that these specific details need not be employed to implement the present invention in practice. In other cases, known components and methods, such as specific and alternative processor architectures, specific logic circuits, etc., have been used.Specific logic code for described algorithms, specific firmware code, specific interconnect operation, specific logical configurations, specific manufacturing techniques and materials, specific compiler implementations, specific expression of algorithms in code, specific shutdown and gate control techniques or specific shutdown and gate control logic, and other specific operational details of a computer system are not described in order to avoid unnecessary complication of the present invention.
[0007] Although the following embodiments may be described with respect to energy saving and energy efficiency in specific integrated circuits, such as computer platforms or microprocessors, other embodiments are applicable to other types of integrated circuits and logic devices. Similar techniques and teachings of embodiments described herein can be applied to other types of circuits or semiconductor devices that also benefit from improved energy efficiency and energy saving. For example, the disclosed embodiments are not limited to desktop computer systems or Ultrabooks™. They can also be used in other devices, such as portable devices, tablets, other thin notebooks, system-on-a-chip (SoC) devices, and embedded applications.Some examples of portable equipment include cellular phones, Internet Protocol equipment, digital cameras, personal digital assistants (PDAs), and portable PCs. Embedded applications typically include a microcontroller, a digital signal processor (DSP), a system-on-a-chip, network computers (NetPCs), set-top boxes, network hubs, wide area network (WAN) switches, or any other system capable of performing the functions and operations described below. Furthermore, the devices, methods, and systems described herein are not limited to physical computing equipment but may also relate to software optimizations for power saving and energy efficiency.As can be seen from the following description, the embodiments of methods, devices and systems described herein (whether in relation to hardware, firmware, software or a combination thereof) are vital for a future in which environmentally friendly technology and performance considerations are in balance.
[0008] As computer systems advance, their components become more complex. Consequently, the complexity of the interconnect architecture for coupling and communicating between components also increases to ensure that bandwidth requirements for optimal component operation are met. Furthermore, different market segments demand that various aspects of interconnect architectures be adapted to market needs. For example, servers require higher performance, while the mobile ecosystem is sometimes able to sacrifice overall performance for energy savings. Nevertheless, a key objective of most fabrics is to provide the highest possible performance with maximum energy savings. The following section discusses a number of interconnects that would potentially benefit from aspects of the invention described herein.
[0009] With reference to Fig. Figure 1 shows an embodiment of a block diagram for a computer system comprising a multi-core processor. The processor 100 comprises any processor or processing unit, such as a microprocessor, an embedded processor, a digital signal processor (DSP), a network processor, a handheld processor, an application processor, a coprocessor, a system-on-a-chip (SoC), or any other code-executing device. In one embodiment, the processor 100 comprises at least two cores, cores 101 and 102, which may be asymmetric cores or symmetric cores (the illustrated embodiment). However, the processor 100 may comprise any number of processing units, which may be symmetric or asymmetric.
[0010] In one embodiment, a processing element refers to hardware or logic to support a software thread. Examples of hardware processing elements include: a thread unit, a thread slot, a process unit, a context, a context unit, a logical processor, a hardware thread, a core, and / or any other element capable of holding a state for a processor, such as an execution state or an architectural state. In other words, in one embodiment, a processing element refers to any hardware that can be independently associated with code, such as a software thread, an operating system, an application, or other code.A physical processor (or processor socket) typically refers to an integrated circuit that potentially includes any number of other processing elements, such as cores or hardware threads.
[0011] A core often refers to logic residing on an integrated circuit capable of maintaining independent architectural states, each associated with at least some dedicated execution resources. In contrast to cores, a hardware thread typically refers to any logic on an integrated circuit capable of maintaining independent architectural states, where these independently maintained architectural states share access to execution resources. As can be seen, the line between the nomenclature of a hardware thread and a core overlaps when certain resources are shared and others are tightly allocated to a specific architectural state.However, often an operating system treats a core and a hardware thread as individual logical processors, with the operating system being able to schedule operations on each logical processor individually.
[0012] The physical processor 100 comprises, as in Fig. Figure 1 illustrates two cores, core 101 and 102. Here, cores 101 and 102 are considered symmetric cores, i.e., cores with the same configurations, functional units, and / or logic. In another embodiment, core 101 comprises an out-of-order processor core, while core 102 comprises an in-order processor core. However, cores 101 and 102 can be individually selected from any core type, such as a native core, a software-defined core, a core designed to run a native instruction set architecture (ISA), a core designed to run a translated instruction set architecture (ISA), a collaboratively developed core, or any other known core. In a heterogeneous core environment (i.e.,(In the case of asymmetric kernels) some form of translation, such as binary translation, can be used to dispose of or execute code on one or both kernels. To further the discussion, the functional units illustrated in kernel 101 are described in more detail below, since the units in kernel 102 function similarly in the embodiment shown.
[0013] As shown, the core 101 comprises two hardware threads, 101a and 101b, which can also be referred to as hardware thread slots 101a and 101b. Therefore, in one embodiment, software instances, such as an operating system, potentially view the processor 100 as four separate processors, i.e., four logical processors or processing units, capable of executing four software threads simultaneously. As previously mentioned, a first thread is associated with architecture state registers 101a, a second thread is associated with architecture state registers 101b, a third thread can be associated with architecture state registers 102a, and a fourth thread can be associated with architecture state registers 102b. Each of the architecture state registers (101a, 101b, 102a, and 102b) can be referred to as a processing unit, thread slot, or thread unit, as previously described.As illustrated, the architecture state registers 101a are replicated in the architecture state registers 101b, allowing individual architecture states / contexts to be stored for logical processor 101a and logical processor 101b. In core 101, other, smaller resources, such as instruction pointers and rename logic in an allocation and rename block 130, can also be replicated for threads 101a and 101b. Some resources, such as reorder buffers in a reorder / deselect unit 135, I-TLB 120, load / store buffers, and queues, can be shared through partitioning. Other resources, such as internal universal registers, page table-based registers, a subordinate data cache and data TLB 115, execution unit(s) 140 and parts of an out-of-order unit 135, are potentially shared in their entirety.
[0014] The processor 100 can often include other resources that can be fully shared, shared through partitioning, or dedicated to / for processing elements. In Fig. Figure 1 is an illustrative representation of a purely exemplary processor, with illustrative logical units / resources of a processor shown. It should be noted that a processor may include or omit any of these functional units, as well as any other known functional unit, logic, or firmware not shown. As illustrated, core 101 includes a simplified representative out-of-order (OOO) processor core. However, in various embodiments, an in-order processor may be used. The OOO core includes a branch target buffer (BTB) 120 to predict branches to be executed / taken, and an instruction translation buffer (I-TLB) 120 to store address translation entries for instructions.
[0015] The core 101 further comprises a decode module 125, which is coupled to a retrieval unit 120 to decode retrieved elements. In one embodiment, the retrieval logic comprises individual sequencers associated with thread slots 101a and 101b, respectively. Typically, the core 101 is associated with a first ISA that defines / specifies instructions executable on the processor 100. Machine code instructions that are part of the first ISA often include a portion of the instruction (referred to as an opcode) that references / specifies an instruction or operation to be executed. The decode logic 125 comprises circuitry that recognizes these instructions by their opcodes and forwards the decoded instructions on the pipeline for processing as defined by the first ISA.As explained in more detail below, the Decoder 125, for example, in one embodiment, has logic designed or configured to recognize specific instructions, such as a transaction instruction. As a result of recognition by the Decoder 125, the Architecture or Core 101 performs specific predefined actions to execute tasks associated with the corresponding instruction. It is important to note that all of the tasks, blocks, operations, and procedures described herein can be executed in response to a single instruction or multiple instructions, some of which may be new or old. It should be noted that in one embodiment, the Decoder 126 recognizes the same ISA (or a subset thereof). Alternatively, in a heterogeneous core environment, the Decoder 126 can recognize a second ISA (either a subset of the first ISA or a different ISA).
[0016] In one example, the allocation and renaming block 130 includes an assignor for reserving resources, such as register files for storing the results of instruction execution. Threads 101a and 101b are potentially capable of out-of-order execution, however, and allocation and renaming block 130 also reserves other resources, such as a reorder buffer for tracking instruction results. Unit 130 may also include a register renamer to rename program / instruction reference registers to other, processor 100-internal registers. A reorder / elimination unit 135 includes components, such as the previously mentioned reorder buffers, load buffers, and memory buffers, to support out-of-order execution and subsequent in-order elimination of out-of-order instructions.
[0017] In one embodiment, a scheduler and execution unit block 140 comprises a scheduler unit for scheduling instructions / operations on execution units. For example, a floating-point instruction is scheduled on a port of an execution unit that has an available floating-point execution unit. Register files associated with the execution units are also included to store the results of processing information instructions. Exemplary execution units include a floating-point execution unit, an integer execution unit, a jump execution unit, a load execution unit, a memory execution unit, and other known execution units.
[0018] A lower-level data cache and data translation buffer (D-TLB) 150 are coupled to the execution unit(s) 140. The data cache stores the most recently used / processed elements, such as data operands, which are potentially held in memory coherence states. The D-TLB stores the most recent translations from virtual / linear to physical addresses. As a specific example, a processor might include a page table structure to partition physical memory into a multitude of virtual pages.
[0019] In this configuration, cores 101 and 102 share access to higher-level or more distant caches, such as a second-level cache associated with an on-chip interface 110. It is important to note that "higher-level" or "more distant" refers to cache levels that are higher or further away from the execution unit(s). In one embodiment, the higher-level cache is a last-level data cache—the final cache in the memory hierarchy on processor 100—such as a second- or third-level data cache. However, the higher-level cache is not limited to this, as it can be associated with or include an instruction cache. A trace cache—a type of instruction cache—can instead be coupled downstream of decoder 125 to store the most recently decoded traces.In this context, an instruction potentially refers to a macro instruction (i.e., a general instruction that is recognized by the decoders), which can be decoded into a number of micro instructions (micro-operations).
[0020] In the configuration shown, the processor 100 also includes an on-chip interface module 110. Historically, a memory controller, which is described in more detail below, was contained outside the processor 100 in a computer system. In this scenario, the on-chip interface 11 is used to communicate with devices outside the processor 100, such as system memory 175, a chipset (which often includes a memory controller hub for connecting to the memory 175 and an I / O controller hub for connecting peripheral devices), a memory controller hub, a northbridge, or another integrated circuit. And in this scenario, the bus 105 can be any known intermediate connection, such as a multipoint interconnect bus, a point-to-point interconnect, a serial interconnect, a parallel bus, a coherent (e.g.,include a cache-coherent bus, a layered protocol architecture, a differential bus, and a GTL bus.
[0021] Memory 175 can be dedicated to the processor 100 or shared with other components in a system. Common examples of memory types 175 include DRAM, SRAM, non-volatile memory (NV memory), and other known memory devices. It should be noted that the device 180 can include a graphics accelerator, a graphics processor or graphics card coupled with a memory controller hub, a data storage device coupled with an I / O controller hub, a wireless transceiver, a flash device, an audio controller, a network controller, or other known devices.
[0022] For some time now, however, each of these facilities can be integrated into a processor 100, as more logic and facilities are integrated into a single chip such as a SoC. For example, in one embodiment, a memory controller hub is on the same package and / or chip as the processor 100. In this configuration, a portion of the core (an in-core portion) 110 includes one or more controllers for connecting to other facilities, such as memory 175 or graphics 180, via an interface. The configuration that includes an intermediary connection and controllers for connecting to such facilities via an interface is often referred to as an in-core (or out-of-core) configuration.As an example, the in-chip interface 110 includes a ring interconnect for in-chip communication and a serial point-to-point high-speed link 105 for out-of-chip communication. However, in the SoC environment, even more features, such as the network interface, coprocessors, memory 175, graphics processor 180, and other familiar computer features / interfaces, can be integrated into a single chip or integrated circuit to provide a small form factor with high functionality and low power consumption.
[0023] In one embodiment, the processor 100 is capable of executing a compiler, an optimizer, and / or a translator code 177 for compiling, translating, and / or optimizing application code 176 to support the devices and methods described herein or to connect to them via an interface. A compiler often comprises a program or a set of programs for translating source text / code into target text / code. Typically, the compilation of program / application code with a compiler is performed in several stages and passes to convert a higher-level programming language into machine-oriented machine or assembly language code. However, single-pass compilers can still be used for simple compilations.A compiler can use any known compilation technique and perform all known compiler operations, such as lexical analysis, preprocessing, syntactic analysis, semantic analysis, code generation, code transformation, and code optimization.
[0024] Larger compilers often comprise several phases, but these phases are most commonly contained in two main stages: (1) the frontend, where syntactic processing, semantic processing, and some transformation / optimization generally take place, and (2) the backend, where analysis, transformations, optimizations, and code generation generally occur. Some compilers refer to a middle ground, which illustrates the blurring of the boundaries between a compiler's frontend and backend. As a result, references to insertion, association, generation, or other compiler operations can occur in any of the aforementioned phases or passes, as well as in any other known phases or passes of a compiler. As an illustrative example, a compiler potentially inserts operations, calls, functions, etc.Dynamic compilation involves one or more compilation phases, such as inserting calls / operations in a frontend compilation phase and subsequently converting these calls / operations into lower-level code during a conversion phase. It's worth noting that during dynamic compilation, compiler code or dynamic optimization code can insert such operations / calls and optimize the code for execution at runtime. As a specific example, binary code (already compiled code) can be dynamically optimized at runtime. The program code can include the dynamic optimization code, the binary code, or a combination thereof.
[0025] Similar to a compiler, a translator, such as a binary translator, translates code either statically or dynamically to optimize and / or translate code. Therefore, references to code execution, application code, program code, or other software environment may refer to: (1) either dynamic or static execution of compiler program(s), optimization code optimizer(s), or translator(s) to compile program code, maintain software structures, perform other operations, optimize code, or translate code; (2) execution of main program code that includes operations / calls, such as application code that has been optimized / compiled; (3) execution of other program code, such as libraries associated with the main program code, to maintain software structures, perform other software-related operations, or optimize code; or (4) a combination thereof.
[0026] An intermediate interconnect fabric architecture encompasses the PCIe architecture. A primary goal of PCIe is to enable interoperability between components and devices from different vendors within an open architecture spanning multiple market segments: desktop and mobile clients, servers (standard and enterprise), and embedded and communications devices. PCI Express is a high-performance, general-purpose I / O intermediate interconnect designed for a wide variety of future computing and communications platforms. Some PCI attributes, such as its usage model, load / store architecture, and software interfaces, have remained consistent throughout its revisions, while earlier parallel bus implementations have been replaced by a highly scalable, fully serial interface.The latest versions of PCI Express leverage advancements in point-to-point interconnects, switch-based technology, and packaged protocol to deliver new levels of performance and features. Power management, quality of service (QoS), hot-plug / hot-swap support, data integrity, and error handling are among the advanced features supported by PCI Express.
[0027] With reference to Fig. Figure 2 illustrates an embodiment of a fabric consisting of point-to-point links connecting a set of components. A system 200 comprises a processor 205 and a system memory 210 coupled to a controller hub 215. The processor 205 comprises any processing element, such as a microprocessor, a host processor, an embedded processor, a coprocessor, or another processor. The processor 205 is coupled to the controller hub 215 by a front-side bus (FSB) 206. In one embodiment, the FSB 206 is a serial point-to-point link, as described below. In another embodiment, the link 206 comprises a serial differential link architecture compatible with a different link standard.
[0028] System memory 210 comprises any memory device, such as random-access memory (RAM), non-volatile (NV) memory, or other memory accessible by devices in system 200. System memory 210 is coupled to the controller hub 215 by a memory interface 216. Examples of memory interfaces include a dual-data-rate (DDR) memory interface, a dual-channel DDR memory interface, and a DRAM (dynamic RAM) memory interface.
[0029] In one embodiment, the controller hub 215 can comprise a root hub, root complex, or root controller in a PCIe (Peripheral Component Interconnect Express) interconnect hierarchy. Examples of the controller hub 215 include a chipset, a memory controller hub (MCH), a northbridge, an interconnect controller hub (ICH), a southbridge, and a root controller / hub. The term chipset often refers to two physically separate controller hubs, i.e., a memory controller hub (MCH) coupled to an interconnect controller hub (ICH). It is worth noting that current systems often have the MCH integrated into the processor 205, while the controller 215 communicates with I / O devices in a similar manner, as described below.In some embodiments, partner-to-partner routing is optionally supported by the stem complex 215.
[0030] Here, the controller hub 215 is coupled to a switch or bridge 220 via a serial link 219. Input / output modules 217 and 221, which can also be referred to as interfaces / ports 217 and 221, comprise / implement a layered protocol stack to provide communication between the controller hub 215 and the switch 220. In one embodiment, multiple devices are capable of being coupled to the switch 220.
[0031] The switch or bridge 220 forwards packets / messages from a device 225 upstream, i.e., one hierarchy upwards to a root complex, the controller hub 215, and downstream, i.e., one hierarchy downwards from a root controller, the processor 205, or the system memory 210 to the device 225. In one embodiment, the switch 220 is referred to as a logical array of multiple virtual PCI-to-PCI bridge devices. The device 225 comprises any internal or external device or component intended to be coupled to an electronic system, such as an I / O device, a network interface controller (NIC), or a network interface card (NIC).Network Interface Controller), an add-in card, an audio processor, a network processor, a hard disk drive, a storage device, a CD / DVD-ROM drive, a monitor, a printer, a mouse, a keyboard, a router, a portable storage device, a FireWire device, a USB (Universal Serial Bus) device, a scanner, and other input / output devices. In PCIe terminology, such a device is often referred to as an endpoint. Although not specifically shown, the device may include a PCIe-to-PCI / PCI-X bridge to support legacy or other versions of PCI devices. Endpoint devices in PCIe are often classified as integrated legacy, PCIe, or root complex endpoints.
[0032] Furthermore, a graphics accelerator 230 is coupled to the controller hub 215 via a serial link 232. In one embodiment, the graphics accelerator 230 is coupled to an MCH, which is coupled to an ICH. The switch 220, and consequently the I / O device 225, are then coupled to the ICH. Additionally, I / O modules 231 and 218 are used to implement a layered protocol stack for communication between the graphics accelerator 230 and the controller hub 215. Similar to the MCH discussion above, a graphics controller or the graphics accelerator 230 itself can be integrated into the processor 205.
[0033] Turning towards Fig. Figure 3 illustrates an embodiment of a layered protocol stack. The layered protocol stack 300 comprises any form of layered communications stack, such as a QPI (Quick Path Interconnect) stack, a PCle stack, a next-generation high-performance computer interconnect stack, or any other layered stack. Although the discussion immediately following refers to Fig. Since the concepts described in sections 2 to 5 refer to a PCle stack, the same concepts can be applied to other intermediate link stacks. In one embodiment, the protocol stack 300 is a PCle protocol stack comprising a transaction layer 305, a link layer 310, and a physical layer 320. An interface can be represented as a communication protocol stack 300. The representation as a communication protocol stack can also be referred to as a module or interface that implements / comprises a protocol stack.
[0034] PCI Express uses packets to communicate information between components. The packets are formed in the transaction layer 305 and the data link layer 310 to transfer information from the sending component to the receiving component. As the sent packets pass through the other layers, they are augmented with additional information necessary for handling the packets at those layers. On the receiving side, the reverse process takes place, and the packets are transformed from their physical layer 320 representation to the data link layer 310 representation and finally (for transaction layer packets) into a form that can be processed by the transaction layer 305 of the receiving device.
[0035] In one embodiment, transaction layer 305 serves to provide an interface between a facility's processing core and the intermediate link architecture, such as the data link layer 310 and the physical layer 320. In this respect, a primary responsibility of transaction layer 305 is the packaging and depackaging of packets (i.e., transaction layer packets, or TLPs). Transaction layer 305 typically manages credit-based flow control for TLPs. PCIe implements split transactions, meaning transactions with a time-separated request and response, which allow one link to carry other traffic while the destination facility gathers data for the response.
[0036] Furthermore, PCIe uses credit-based flow control. In this scheme, an institution declares an initial amount of credit for each of the receive buffers in transaction layer 305. An external institution at the opposite end of the link, such as a controller hub, counts the number of credits consumed by each TLP. A transaction can be sent if the transaction does not exceed a credit limit. Upon receiving a response, a credit amount is restored. One advantage of such a credit scheme is that the latency of credit restoration does not impact performance, provided the credit limit is not reached.
[0037] In one embodiment, four transaction address spaces comprise a configuration address space, a memory address space, an input / output address space, and a message address space. Memory space transactions comprise one or more read and write requests to transfer data to and from a memory location. In one embodiment, memory space transactions are capable of using two different address formats, such as a short address format like a 32-bit address or a long address format like a 64-bit address. Configuration space transactions are used to access the configuration space of the PCIe devices. Configuration space transactions include read and write requests. Message space transactions (or simply messages) are defined to support in-band communication between PCIe agents.
[0038] Therefore, in one implementation, transaction layer 805 packages packet headers / payloads 806. The format for current packet headers / payloads can be found in the PCle specification on the PCle specification website.
[0039] Briefly referring to Fig. Figure 4 illustrates an embodiment of a PCIe transaction descriptor. In one embodiment, the transaction descriptor 400 is a mechanism for transmitting transaction information. In this respect, the transaction descriptor 400 supports the identification of transactions in a system. Other potential uses include tracking modifications to the default transaction order and associating transactions with channels.
[0040] The transaction descriptor 400 comprises a global identifier field 402, an attribute field 404, and a channel identifier field 406. In the illustrated example, the global identifier field 402 is represented as comprising a local transaction identifier field 408 and a source identifier field 410. In one embodiment, the global transaction identifier 402 is unique for all pending requests.
[0041] According to one implementation, the local transaction identifier field 408 is a field generated by a requesting agent and is unique for all pending requests that require completion by that requesting agent. Additionally, in this example, the source identifier 410 uniquely identifies the requesting agent within a PCle hierarchy. Therefore, the local transaction identifier field 408, together with the source ID 410, provides a global identification of a transaction within a hierarchy scope.
[0042] Attribute field 404 specifies characteristics and relationships of the transaction. In this respect, attribute field 404 is potentially used to provide additional information that allows for modification of the standard handling of transactions. In one embodiment, attribute field 404 comprises a priority field 412, a reserved field 414, an order field 416, and a do-not-snoop field 418. Here, the priority subfield 412 can be modified by an initiator to assign a priority to the transaction. The reserved attribute field 414 is left reserved for future or vendor-defined use. Possible usage models that utilize priority and security attributes can be implemented using the reserved attribute field.
[0043] In this example, the order attribute field 416 is used to provide optional information that transmits the order type, which can modify default order rules. According to one example implementation, an order attribute of "0" means that default order rules should be applied, while an order attribute of "1" means random ordering, where write operations can overtake write operations in the same direction and read operations can overtake write operations in the same direction. The snoop attribute field 418 is used to determine whether transactions are snooped. As shown, the channel ID field 406 identifies a channel with which a transaction is associated.
[0044] A link layer 310, also referred to as data link layer 310, acts as an intermediate layer between the transaction layer 305 and the physical layer 320. In one embodiment, one responsibility of the data link layer 310 is to provide a reliable mechanism for exchanging transaction layer packets (TLPs) between two components on a link. One side of the data link layer 310 accepts packetized TLPs from the transaction layer 305, applies a packet sequence identifier 311 (i.e., an identification number or packet number), calculates and applies an error detection code (CRC 312), and transmits the modified TLPs to the physical layer 320 for transmission over a physical layer to an external device.
[0045] In one embodiment, the physical layer 320 comprises a logical subblock 321 and an electrical subblock 322 for physically transmitting a packet to an external device. The logical subblock 321 is responsible for the "digital" functions of the physical layer 321. Specifically, the logical subblock includes a transmit section for preparing outgoing information for transmission by the physical subblock 322 and a receive section for identifying and preparing received information before forwarding it to the link layer 310.
[0046] The physical block 322 comprises a transmitter and a receiver. The transmitter is supplied with symbols by the logical subblock 321, which the transmitter serializes and forwards to an external device. The receiver is supplied with serialized symbols from an external device and converts the received signals into a bitstream. The bitstream is deserialized and delivered to the logical subblock 321. In one embodiment, an 8b / 10b transmission code is used, whereby ten-bit symbols are sent / received. Special symbols are used to frame a packet with frame 323. In one example, the receiver also provides a symbol clock, which is recovered from the incoming serial stream.
[0047] Although, as mentioned earlier, the transaction layer (305), the link layer (310), and the physical layer (320) are discussed in relation to a specific implementation of a PCIe protocol stack, a layered protocol stack is not limited to these. In fact, any layered protocol can be included / implemented. As an example, a port or interface represented as a layered protocol comprises: (1) a first layer for packaging packets, i.e., a transaction layer; (2) a second layer for sequencing packets, i.e., a link layer; and (3) a third layer for sending the packets, i.e., a physical layer. A QPL layered protocol is used as a specific example.
[0048] Next, with reference to Fig. Figure 5 illustrates an embodiment of a serial point-to-point PCIe fabric. Although one embodiment of a serial point-to-point PCIe link is illustrated, a serial point-to-point link is not limited to this, as it includes any transmission path for sending serial data. In the illustrated embodiment, a basic PCIe link comprises two differentially controlled low-voltage signal pairs: a transmit pair 506 / 511 and a receive pair 512 / 507. Accordingly, a device 505 comprises transmit logic 506 to send data to a device 510 and receive logic 507 to receive data from the device 510. In other words, two transmit paths, i.e., paths 516 and 517, and two receive paths, i.e., paths 518 and 515, are included in a PCIe link.
[0049] "Transmission path" refers to any path for sending data, such as a transmission line, copper wire, optical line, wireless communication channel, infrared communication link, or other communication path. A connection between two facilities, such as Facility 505 and Facility 510, is called a link, such as Link 415. A link can support one lane—each lane being a set of differential signal pairs (one pair for transmitting, one pair for receiving). To scale bandwidth, a link can aggregate multiple lanes, denoted by xN, where N is any supported link width, such as 1, 2, 4, 8, 12, 16, 32, 64, or wider.
[0050] "Differential pair" refers to two transmission paths, such as lines 516 and 517, used to transmit differential signals. For example, line 517 switches from a logic high (H) to a logic low (L) level (i.e., a falling edge) when line 516 switches from a low to a high level (i.e., a rising edge). Differential signals potentially exhibit better electrical characteristics, such as improved signal integrity (i.e., reduced cross-coupling, overshoot / undershoot, ringing, etc.). This allows for a narrower clock window, which in turn enables faster transmission frequencies.
[0051] Next, with reference to Fig. Figure 6 illustrates an embodiment of a SoC design according to one embodiment. As a specific illustrative example, the SoC 2000 is included in the user equipment (UE). In one embodiment, UE refers to any equipment used by an end user for communication, such as a mobile phone, smartphone, tablet, ultra-thin notebook, notebook with broadband adapter, or any other similar communication equipment. A UE often connects to a base station or node that, in its nature, is potentially equivalent to a mobile station (MS) in a GSM network.
[0052] Here, the SoC 2000 comprises two cores—2006 and 2007. Similar to the preceding discussion, cores 2006 and 2007 can correspond to an instruction set architecture such as an Intel® Architecture Core™-based processor, an Advanced Micro Devices, Inc. (AMD) processor, a MIPS-based processor, an ARM-based processor design, or a customer thereof, as well as their licensees or subcontractors. Cores 2006 and 2007 are coupled to a cache controller 2008, which is associated with a bus interface unit 2009 and an L2 cache 2010 to communicate with other parts of the System 2000. An intermediate 2010 comprises an on-chip intermediate, such as an IOSF, AMBA, or another intermediate discussed previously, which potentially implements one or more of the aspects described herein.
[0053] The intermediate interface 2010 provides communication channels to the other components, such as a Subscriber Identification Module (SIM) 2030 for connecting to a SIM card, a Start ROM 2035 for storing a start code for execution by the cores 2006 and 2007 to initialize and start the SoC 2000, an SDRAM controller 2040 for connecting to external memory (e.g., DRAM 2060), a Flash controller 2045 for connecting to non-volatile memory (e.g., Flash 2065), a Peripheral Controller 2050 (e.g., a serial peripheral interface) for connecting to peripheral devices, video codecs 2020 and a video interface 2025 for displaying and receiving input (e.g., touch input), a GPU 2015 for performing graphics-related calculations, etc. Each of these interfaces can handle aspects described herein. include.
[0054] Furthermore, the system illustrates communication peripherals, such as a Bluetooth module 2070, a 3G modem 2075, a GPS 2080, and WiFi 2085. A power controller 2055 is also included in the system. As mentioned earlier, a UE includes a radio for communication. Consequently, not all of these communication peripherals are required. However, a UE must contain some form of radio for external communication.
[0055] In various embodiments, at least sections of a circuit arrangement of one or more devices coupled by a specific intermediate connection can be power-controlled (e.g., power gate-controlled) if the configuration of the device determines that such a circuit arrangement is not used. As an example of the embodiments described herein, a circuit arrangement associated with one or more virtual channels that provide communication via the intermediate connection can be placed in a shutdown state (e.g., by not providing an operating voltage to such a circuit arrangement) if the configuration of a system determines that such virtual channels are not used for communication.Of course, embodiments are not limited to this example; the techniques described herein apply equally to the power control of other circuit arrangements.
[0056] Fig. 7 is a block diagram of a system according to an embodiment of the present invention. As in Fig. As shown in Figure 7, System 700 is an implementation of a PCle™ system with various units connected to a Switch 720. Each unit is connected to the Switch 720 via a corresponding link (links 1 to 4, respectively). It should be noted that in one embodiment, each link can have different characteristics and operating parameters.
[0057] For example, devices 730, 740, and 750 could be different types of peripheral devices. For instance, device 730 could be a graphics acceleration device, device 740 could be a storage device, and device 750 could be another type of portable device, such as a capture device. Switch 720 is further connected to a root complex 710 via another link (Link 4). For example, root complex 710 could be the system's main data processor, such as a multi-core processor. Of course, other examples of such complexes are possible.
[0058] With particular reference to the connection between the switch 720 and the device 730, it should be noted that a variable number of virtual channels (VCs) are provided in the different devices. As can be seen, in this example, the switch 720 comprises 8 virtual channels, each with a corresponding hardware buffer in a link interface 725. In contrast, the device 730 comprises only 4 virtual channels and therefore has a link interface 735 that includes only 4 hardware buffers. Since these devices have unequal numbers of virtual channels and buffers, at least some of the buffers within the link interface 725 of the switch 720 are not used. Accordingly, these buffers can be deactivated in hardware using an embodiment of the present invention to avoid power consumption for these buffers.It is understood that despite the representation with this particular implementation in the embodiment of . Fig. 7 many variations are possible.
[0059] Now with reference to Fig. Figure 8 shows a flowchart of a configuration procedure according to an embodiment of the present invention. In a particular embodiment, the procedure 800 can be performed during device initialization by configuration logic of devices that are linked to each other. Furthermore, the procedure can also be performed dynamically whenever there is a change to a device or hardware that is linked to a device. For example, when a new device is linked to an endpoint, the link is retrained, and configuration logic accordingly reassigns certain specific values. With reference to Fig. Procedure 800 begins by reading an extended VC count field from a configuration memory of both a local facility or local endpoint and a remote facility or remote endpoint located at the far end of a link coupling these two facilities (Block 810 and Block 820). In one embodiment, this extended VC count field may be stored in a memory of the corresponding facility, for example, within a PCle™ configuration space. For discussion purposes, it is assumed that the local facility (Endpoint 1) is connected to the switch 720 of Fig. 7 corresponds, and that the remote facility (endpoint 2) of facility 730 of Fig. 7 corresponds.
[0060] In the previously described configuration (with 8 virtual channels and buffers present in the switch 720 and 4 virtual channels and buffers present in the device 730), the value returned for the extended VC count field from the switch device 720 is 8, and the value returned from the device 730 is 4. More precisely, in an embodiment where this count field is a 3-bit field, a value of zero corresponds to a single virtual channel being supported (e.g., VC0), and the values 1 through 7 of this 3-bit binary value correspond to the additional number of supported VCs. Thus, in this embodiment, the extended VC count field for the switch 720 has a value of 111b, and the extended VC count field for the device 730 has a count of 011b. Other representations are of course possible.
[0061] Still referring to Fig. The control then moves to block 830, where it can be determined whether the VC count for the local facility is greater than the VC count for the remote facility. If so, the control moves to block 835, where a maximum Link VC value can be set for the extended VC count field of the remote facility. It's worth noting that this maximum Link VC value, or VC ID value, corresponds to a minimum value for the extended VC count field of connected endpoints. Therefore, in this example, this maximum Link VC value is set to 011 b.
[0062] If this is not the case, however, control proceeds from block 830 to block 840, where it can be determined whether the VC count for the local facility equals the VC count for the remote facility. If so, control proceeds to block 845, where the maximum link VC value can be set to the extended VC count field from the local facility. Otherwise, control proceeds to block 850, where it is determined whether the VC count for the local facility is less than the VC count for the remote facility. In this case, control proceeds to block 855, where the maximum link VC value can be set to the extended VC count field from the local facility.
[0063] Regardless of the maximum Link VC value set in any of blocks 835, 845, and 855, control next moves to block 860, where a configuration memory, such as a table, containing these maximum Link VC values can be accessed. More precisely, accessing this table can provide a different representation of the selected maximum Link VC value. As with regard to Fig. As shown in more detail in Figure 9A, the table can contain a multitude of entries, each providing a 3-bit representation of a maximum Link VC value and a corresponding 8-bit representation of the same value. Control then proceeds to block 870, where individual bits of the retrieved entry can be obtained, each bit corresponding to one of the available virtual channels and representing an enabled state of a corresponding hardware buffer. That is, in one example, a logical value of one indicates an active buffer and, accordingly, a corresponding enabled state, and a logical value of zero indicates an inactive buffer and, accordingly, a corresponding disabled state. Other representations are, of course, possible.
[0064] Now with reference to Fig. 9A shows a block diagram of a configuration memory according to one embodiment. As shown in Fig. As shown in Figure 9A, a memory 900 can be located at a desired location within the system, for example, in a separate non-volatile memory. Alternatively, a copy of the information can be stored in the configuration memory 900, for example, in a configuration space of each of the system's facilities. Still other embodiments can store this information in yet another location, such as external registers or read-only memory. As can be seen, the memory 900 comprises a plurality of entries 9100 to 910. nEach entry comprises a first field 920 and a second field 930. The first field 920 can correspond to the maximum Link VC value and is thus used as an addressable means of accessing a selected entry of memory 900. The second field 930, in turn, represents a corresponding 8-bit representation of the maximum Link VC value. In one embodiment, the 8-bit value mapped from the VC ID can be such that each bit corresponds to a buffer of a virtual channel of that endpoint or another lane mapping of the endpoint; for example, bit 0 is mapped to VC0, bit 1 is mapped to VC1, and finally, bit 7 is mapped to VC7. In this embodiment, a binary 1 indicates an "on" state, while 0 indicates "off".In one embodiment, these states are then used to control a voltage control circuit arrangement of each buffer in order to control the provision of an operating voltage for the buffer (or alternatively, not to provide the operating voltage).
[0065] Next, with reference to Fig. Section 9B presents a selected entry 910x, obtained from the table, and its use for controlling the provision of an operating voltage to corresponding hardware buffers of a facility. More precisely, it represents Fig. 9B a set of AND gates 9500 to 950 nEach of these logic gates is configured to receive a corresponding bit from the selected entry and a corresponding supply voltage, which is obtained, for example, from an external voltage regulator. Of course, the supply voltage can instead be received from other locations—either internal or external. When a particular bit has a logical high (H), the AND gate is active, and the supply voltage is provided to the corresponding hardware buffer. Otherwise, the AND gate does not allow the supply voltage to pass to the buffer, and the buffer is deactivated during normal operation, thus reducing power consumption. It is understood that despite the representation, at this high level in the embodiment of Fig. 9B variants are possible.
[0066] Now with reference to Fig. Figure 10 shows a block diagram of a section of a system according to one embodiment. As in Fig. As shown in Figure 10, a system comprises 1000 different components. For discussion purposes, two components are shown here: a device 1010, which may be of any type of integrated circuit, coupled via a link 1005 to another circuit (not shown), and a non-volatile memory 1080, coupled via a link 1090. In the embodiment shown, the link 1005 may be a PCle™ link with unidirectional serial links in a transmit and a receive direction.
[0067] The device 1010 can be any type of device, including a root complex, switch, peripheral device, and so on. For the sake of simplicity, only sections of the device 1010 are shown. More precisely, it is a set of receive buffers 10250 to 1025.n These hardware buffers can each be assigned to a specific virtual channel VC0 to VC1. n correspond. In addition, there is also a set of transmit buffers from 10200 to 1020. n The invention provides for a system in which each buffer is associated with a specific virtual channel. It should be noted that different traffic classes (TCs) can be assigned to be routed through specific virtual channels (VCs) and corresponding hardware buffers. Using one embodiment of the present invention, only activated virtual hardware buffers of these virtual channels are supplied with an operating voltage via a gate logic 1060, which in turn receives control information from a configuration logic 1050, the details of which are discussed below.
[0068] Still referring to the device 1010, buffers 1020 and 1025 can be used with a different circuit arrangement of a physical layer (for the sake of simplicity, in Fig. (10 not shown) communicate. From there, communication can proceed with a link logic 1030 to perform various link layer processing operations. Communication can then proceed with a transaction logic 1035, which can perform transaction layer processing operations. After that, communication can occur with a core logic 1040, which can be the main logic circuit arrangement of the device 1010. For example, in the context of a multi-core processor, the core logic 1040 could be one or more processor cores or other processing circuits. As another example, if the device 1010 is a graphics acceleration device, the core logic 1040 could be a graphics processing unit.
[0069] Furthermore, in Fig. 10. A configuration logic 1050 is represented, which may consist of hardware, software, and / or firmware (or combinations thereof) used to perform configuration operations when the system 1000 is powered on, when the facility 1010 is reset, or when other dynamic changes occur during operation. In one embodiment, the configuration logic 1050 may include logic for performing the power management control described herein. Accordingly, in the embodiment described in Fig. In the embodiment shown in Figure 10, the configuration logic 1050 is a power management logic (PML) 1055, which can be configured to execute a procedure such as the previously discussed procedure 800.
[0070] For this purpose, the configuration logic 1050 can communicate with a configuration memory 1070. In various embodiments, the configuration memory 1070 can be a non-volatile memory of the facility 1010, which includes a PCle™ configuration memory space. Among the various configuration information stored therein is an extended VC counting field 1075, as described herein. It is understood, of course, that additional configuration information is also stored in the memory 1070.
[0071] In a particular embodiment, an assignment table can be provided which associates a maximum 3-bit virtual circuit (VC) count value with a corresponding 8-bit value. The number of encoded bits and single-bit representations can, of course, vary depending on the number of possible virtual channels, hardware buffers, or other circuit arrangements to be controlled. In the illustrated embodiment, the separate non-volatile memory 1080 can include an assignment table 1085. The assignment table 1085 can thus associate a maximum VC count value with lane assignments. In other words, an indication of the maximum number of supported virtual channels can be assigned to a corresponding set of activation indicators, which can be used to control whether an operating voltage is supplied to one or more hardware buffers, each associated with a corresponding virtual channel.It is understood, of course, that although this embodiment discusses the control of hardware buffers based on enabled virtual channels, additional hardware within a facility can be controlled similarly. For example, such additional hardware, controlled on a lane-to-lane or virtual channel basis, could include a graphics card, a daughterboard such as a USB-to-PCIe™ or SATA-to-PCIe™ card, or any other such facility.
[0072] According to this implementation, unused hardware, such as VC hardware buffers, can be switched off according to configuration control, e.g., via configuration logic. This reduces system performance, as only hardware that is actually in use is switched on.
[0073] The following examples relate to further embodiments.
[0074] In one example, a device comprises: a plurality of hardware buffers, each for storing information associated with one or more virtual channels; configuration logic for determining an identifier corresponding to a maximum number of virtual channels jointly supported by a first facility and a second facility coupled via a link, and for obtaining a control value based on the identifier; and gate logic for providing an operating voltage to corresponding to the plurality of hardware buffers based on the control value.
[0075] In one embodiment, the gate logic can be configured to prevent the provision of operating voltage to at least one of the plurality of hardware buffers if the maximum number of virtual channels is less than the plurality of hardware buffers.
[0076] In one example, the configuration logic for determining the maximum number of virtual channels is based on a first count of virtual channels associated with the first facility and a second count of virtual channels associated with the second facility.
[0077] In one example, the configuration logic is to obtain the first count of virtual channels from a virtual channel count field of a configuration store of the first facility and to obtain the second count of virtual channels from a virtual channel count field of a configuration store of the second facility.
[0078] In one example, gate logic comprises a multitude of logic circuits, each designed to receive a bit of the control value and the operating voltage, and to provide the operating voltage to one of the multitude of hardware buffers based on a value of the bit.
[0079] In one example, a non-volatile memory comprising an assignment table with a plurality of entries, each for associating an identifier with a control value, can be coupled to the device. In one embodiment, the configuration logic for obtaining the control value from an entry of the assignment table is accessed using the identifier. The control value can comprise a plurality of bits, each associated with one of the plurality of hardware buffers, with each of the bits representing a first state to indicate that the associated hardware buffer should be enabled and a second state to indicate that the associated hardware buffer should be disabled.
[0080] In one example, the first device includes a configuration memory for storing a count of the maximum number of virtual channels supported by the first device, and also for storing a copy of one or more entries in the mapping table. In one embodiment, the non-volatile memory is a separate component from the first device and is coupled to the first device via a second link.
[0081] It should be mentioned that the aforementioned setup can be implemented using various means.
[0082] In one example, a processor comprises a system-on-a-chip (SoC) that is integrated into a touch-enabled device of a user facility.
[0083] In another example, a system includes a display and a memory, and it includes the setup according to one or more of the preceding examples.
[0084] In another example, a procedure includes: determining a common number of virtual channels that can be supported by a first endpoint and a second endpoint coupled via a link; accessing a memory using the common number of virtual channels to obtain a control setting equal to the common number of virtual channels; and providing an operating voltage to selected first hardware buffers of the first endpoint and selected second hardware buffers of the second endpoint based on the control setting.
[0085] In one example, providing the operating voltage includes providing the operating voltage for the selected first and second hardware buffers and not providing the operating voltage for unselected first hardware buffers and unselected second hardware buffers.
[0086] In one example, the procedure further includes communicating data between the first endpoint and the second endpoint using the selected first hardware buffer and the selected second hardware buffer.
[0087] In one example, the procedure further includes accessing the memory during the configuration of the link using the common number of virtual channels, wherein the memory is separate from the first and second endpoints and includes a plurality of entries, each of which is for storing a common number of virtual channels and a control setting.
[0088] In one example, the procedure in response to a reconfiguration of the link further includes: determining a second common number of virtual channels that can be supported by the first and second endpoints; accessing memory using the second common number of virtual channels to obtain a second control setting; and providing operating voltage to first hardware buffers other than the selected first hardware buffers and to second hardware buffers other than the selected second hardware buffers based on the second control setting.
[0089] In another example, a computer-readable medium containing instructions for carrying out the procedure according to one of the preceding examples is used.
[0090] In another example, a device comprises means for carrying out the method according to one of the preceding examples.
[0091] In another example, a device comprises: a first link interface for connecting the device to a link coupled between the device and a second facility, wherein the first link interface comprises a plurality of independent circuits, each for communicating data of a corresponding traffic class; a first configuration memory for storing a maximum supported value corresponding to the number of the plurality of independent circuits; configuration logic for determining a maximum link value corresponding to a minimum of the maximum supported value stored in the first configuration memory and a maximum supported value stored in a second configuration memory of the second facility, and for obtaining another representation of the maximum link value;and a control circuit for activating a first set of the plurality of independent circuits and for deactivating a second set of the plurality of independent circuits in response to the other representation when the maximum link value is less than a number of the plurality of independent circuits.
[0092] In one example, the maximum supported value stored in the first configuration memory also corresponds to a count of virtual channels for the device.
[0093] In one example, a non-volatile memory coupled to the device comprises a mapping table with a multitude of entries, each associating a maximum link value with a different representation of the maximum link value. The mapping table can be accessed using the maximum link value, which is determined by the configuration logic.
[0094] In one example, the other representation comprises a plurality of bits, each associated with one of the plurality of independent circuits, with each of the bits of a first state being to indicate that the associated independent circuit should be enabled, and of a second state being to indicate that the associated independent circuit should be disabled.
[0095] In one example, the control circuit comprises a multitude of logic circuits, each of which is designed to receive one bit of the multitude of bits of the other representation and an operating voltage from a voltage regulator, and to provide the operating voltage to one of the multitude of independent circuits based on a value of the bit.
[0096] In one example, the multitude of independent circuits each includes a hardware buffer associated with a virtual channel.
[0097] In yet another example, a system comprises: a first facility, which includes a first link interface with a first plurality of hardware buffers, each for storing information associated with one or more virtual channels, and a second facility, which is coupled to the first facility via a link.In one embodiment, the second device comprises: a second link interface with a second plurality of hardware buffers, each for storing information associated with one or more virtual channels, wherein there are more of the second plurality of hardware buffers than of the first plurality of hardware buffers; a controller for determining a maximum number of virtual channels jointly supported by the first and second devices, wherein the maximum number corresponds to the number of the plurality of first hardware buffers, and for obtaining a control value based on the maximum number; and gate logic for activating fewer than all of the plurality of second hardware buffers in response to the control value.
[0098] In one example, the first setup includes a first configuration store containing a first maximum count of virtual channels, and the second setup includes a second configuration store containing a second maximum count of virtual channels.
[0099] In one example, the controller is used to determine the maximum number of jointly supported virtual channels using the first maximum count of virtual channels and the second maximum count of virtual channels.
[0100] In one example, a non-volatile memory comprises an allocation table with a multitude of entries, each of which is used to associate a maximum number of jointly supported virtual channels.
[0101] In one example, the controller can obtain the control value from an entry in the mapping table, accessed using the specified maximum number of shared virtual channels. The control value comprises a set of bits, each associated with one of the second set of hardware buffers, where each bit represents a first state to indicate that the associated second hardware buffer should be enabled, and a second state to indicate that the associated second hardware buffer should be disabled.
[0102] In one example, gate logic comprises a multitude of logic circuits, each designed to receive a bit of the control value and an operating voltage, and to provide the operating voltage to one of the multitude of hardware buffers based on a value of the bit.
[0103] It goes without saying that various combinations of the above examples are possible.
[0104] The embodiments can be used in many different types of systems. For example, in one embodiment, a communication device may be designed to perform the various methods and techniques described herein. The scope of protection of the present invention is, of course, not limited to a communication device; other embodiments may instead be directed to other types of devices for processing instructions or to one or more machine-readable media comprising instructions which, in response to their execution on a computer device, cause the device to execute one or more of the methods and techniques described herein.
[0105] Implementations can be written in code and stored on a non-transitory storage medium containing instructions that can be used to program a system to execute those instructions. The storage medium can be any type of disk, including floppy disks, optical disks, solid-state drives (SSDs), compact disk read-only memories (CD-ROMs), rewritable compact disks (CD-RWs), and magneto-optical disks; semiconductor devices such as read-only memories (ROMs), random-access memories (RAMs), such as dynamic random-access memories (DRAMs), static random-access memories (SRAMs), and erasable programmable read-only memories (EPROMs).This includes, but is not limited to, erasable programmable read-only memories, flash memory, electrically erasable programmable read-only memories (EEPROMs), magnetic or optical cards, or any other type of media capable of storing electronic instructions.
Claims
[1] Device comprising: a multitude of hardware buffers (1020, 1025), each for storing information associated with one or more virtual channels; a configuration logic (1050) for determining an identifier corresponding to a maximum number of virtual channels jointly supported by a first facility (1010) and a second facility coupled via a link (1005), and for obtaining a control value based on the identifier; and a gate logic (1060) for providing an operating voltage for corresponding to the plurality of hardware buffers (1020, 1025) based on the control value, wherein the gate logic (1060) is for preventing the provision of the operating voltage for at least one of the plurality of hardware buffers (1020, 1025) if the maximum number of virtual channels is less than the plurality of hardware buffers (1020, 1025). [2] Device according to claim 1, wherein the configuration logic (1050) is for determining the maximum number of virtual channels based on a first count of virtual channels associated with the first device (1010) and a second count of virtual channels associated with the second device. [3] Device according to claim 2, wherein the configuration logic (1050) is for obtaining the first count of virtual channels from a count field of virtual channels of a configuration memory (1070) of the first device (1010) and for obtaining the second count of virtual channels from a count field of virtual channels of a configuration memory of the second device. [4] Device according to claim 1, wherein the gate logic (1060) comprises a plurality of logic circuits (950), each of which is for receiving a bit of the control value and the operating voltage and for providing the operating voltage to one of the plurality of hardware buffers (1020, 1025) based on a value of the bit. [5] Device according to claim 1, further comprising a non-volatile memory (1080) comprising an assignment table (1085) with a plurality of entries, each of which is for associating an identifier with a control value. [6] Device according to claim 5, wherein the configuration logic (1060) is for obtaining the control value from an entry of the mapping table (1085) which is accessed using the identifier. [7] Device according to claim 6, wherein the control value comprises a plurality of bits, each of which is associated with one of the plurality of hardware buffers (1020, 1025), wherein each of the bits of a first state is to indicate that the associated hardware buffer is to be activated, and of a second state is to indicate that the associated hardware buffer is to be deactivated. [8] Device according to claim 5, wherein the first device (1010) comprises a configuration memory (1070) for storing a count of the maximum number of virtual channels supported by the first device (1010), and the configuration memory (1070) further comprises storing a copy of one or more entries of the mapping table (1085). [9] Device according to claim 8, wherein the non-volatile memory (1080) is a separate component from the first device (1010) and is coupled to the first device (1010) via a second link (1090). [10] Procedures, including: Determining a common number of virtual channels that can be supported by a first endpoint and a second endpoint that are coupled via a link; Accessing a memory using the shared number of virtual channels to obtain a control setting corresponding to the shared number of virtual channels; and Providing an operating voltage for selected first hardware buffers of the first endpoint and selected second hardware buffers of the second endpoint based on the control setting. [11] Method according to claim 10, wherein providing the operating voltage comprises providing the operating voltage for the selected first and second hardware buffers and not providing the operating voltage for unselected first hardware buffers and unselected second hardware buffers. [12] Method according to claim 10, further comprising communicating data between the first endpoint and the second endpoint using the selected first hardware buffer and the selected second hardware buffer. [13] Method according to claim 10, further comprising accessing the memory during the configuration of the link using the common number of virtual channels, wherein the memory is separate from the first and second endpoints and comprises a plurality of entries, each of which is for storing a common number of virtual channels and a control setting. [14] Method according to claim 13, further comprising in response to a reconfiguration of the link: Determine a second common set of virtual channels that can be supported by the first and second endpoints; Accessing memory using the second shared number of virtual channels to obtain a second control setting; and Providing operating voltage for first hardware buffers other than the selected first hardware buffers and for second hardware buffers other than the selected second hardware buffers based on the second control setting. [15] Device (1010), comprising: a first link interface for connecting the device to a link (1005) coupled between the device (1010) and a second device, wherein the first link interface comprises a plurality of independent circuits, each of which is for communicating data of a corresponding traffic class; a first configuration memory (1070) for storing a maximum supported value, which corresponds to the number of the plurality of independent circuits; a configuration logic (1050) for determining a maximum link value corresponding to a minimum of the maximum supported value stored in the first configuration memory (1070) and a maximum supported value stored in a second configuration memory of the second facility, and for obtaining another representation of the maximum link value; and a control circuit (1060) for activating a first set of the plurality of independent circuits and for deactivating a second set of the plurality of independent circuits in response to the other representation when the maximum link value is less than a number of the plurality of independent circuits. [16] Device (1010) according to claim 15, wherein the maximum supported value stored in the first configuration memory (1070) further corresponds to a count value of virtual channels for the device (1010). [17] Device (1010) according to claim 15, further comprising a non-volatile memory (1080) coupled to the device (1010), wherein the non-volatile memory (1080) comprises an assignment table (1085) with a plurality of entries, each of which is for associating a maximum link value with another representation of the maximum link value, wherein the assignment table (1085) is accessed using the maximum link value determined by the configuration logic (1050). [18] Device (1010) according to claim 17, wherein the other representation comprises a plurality of bits, each of which is associated with one of the plurality of independent circuits, wherein each of the bits of a first state is to indicate that the associated independent circuit is to be activated, and of a second state is to indicate that the associated independent circuit is to be deactivated. [19] Device (1010) according to claim 18, wherein the control circuit (1060) comprises a plurality of logic circuits, each of which is for receiving a bit of the plurality of bits of the other representation and an operating voltage from a voltage regulator and for providing the operating voltage to one of the plurality of independent circuits based on a value of the bit. [20] Device (1010) according to claim 15, wherein the plurality of independent circuits each comprises a hardware buffer (1020, 1025) associated with a virtual channel. [21] System (1000), comprising: a first setup (1010) comprising a first link interface with a first set of hardware buffers (1020, 1025), each for storing information associated with one or more virtual channels; and a second facility that is linked to the first facility (1010) via a link (1005), the second facility comprising the following: a second link interface with a second set of hardware buffers, each for storing information associated with one or more virtual channels, with more of the second set of hardware buffers than the first set of hardware buffers (1020, 1025); a controller for determining a maximum number of virtual channels jointly supported by the first and second facilities (1010), the maximum number corresponding to the number of first hardware buffers (1020, 1025), and for obtaining a control value based on the maximum number; and a gate logic to activate less than all of the multitude of secondary hardware buffers in response to the control value. [22] System (1000) according to claim 21, wherein the first device (1010) comprises a first configuration memory (1070) comprising a first maximum count of virtual channels, and the second device comprises a second configuration memory comprising a second maximum count of virtual channels. [23] System (1000) according to claim 22, wherein the controller is for determining the maximum number of jointly supported virtual channels using the first maximum count of virtual channels and the second maximum count of virtual channels. [24] System (1000) according to claim 21, further comprising a non-volatile memory (1080) comprising an allocation table (1085) with a plurality of entries, each of which is for associating a maximum number of jointly supported virtual channels. [25] System (1000) according to claim 24, wherein the controller is for obtaining the control value from an entry of the mapping table (1085) which is accessed using the specified maximum number of jointly supported virtual channels. [26] System (1000) according to claim 25, wherein the control value comprises a plurality of bits, each of which is associated with one of the second plurality of hardware buffers, wherein each of the bits of a first state is to indicate that the associated second hardware buffer is to be activated, and of a second state is to indicate that the associated second hardware buffer is to be deactivated. [27] System (1000) according to claim 26, wherein the gate logic comprises a plurality of logic circuits, each of which is for receiving a bit of the control value and an operating voltage and for providing the operating voltage to one of the plurality of second hardware buffers based on a value of the bit.
Citation Information
Patent Citations
Low power arbiters in interconnection routers
US7472299B2
Performance and traffic aware heterogeneous interconnection network
US8379659B2