System, device and method for adaptive operating voltage in field programmable gate array (FPGA)
By dynamically controlling the operating voltage and frequency of the FPGA, and using the self-test circuit and power controller to identify the minimum operating voltage, the power consumption and performance problems of FPGA in specific functional configurations are solved, achieving more efficient energy efficiency management.
Patent Information
- Application Number
- CN201780094148.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2017-08-23
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2037-08-23
AI Technical Summary
Existing FPGAs have problems with excessive supply voltages in specific functional configurations resulting in increased power consumption and reduced performance.
By dynamically controlling the operating voltage and frequency of the FPGA, the minimum operating voltage is identified using a self-test circuit and power controller to achieve appropriate voltage/frequency values to ensure correct operation and reduce power consumption.
It effectively reduces the power consumption of FPGAs while ensuring their correct operation and life, achieving more efficient energy efficiency management.
Smart Images

Figure CN110998487B_ABST
Abstract
Description
Technical Field
[0001] Embodiments relate to power management in integrated circuits. Background Art
[0002] Integrated circuits called field-programmable gate arrays (FPGAs) have gained widespread popularity in recent years because these devices can be programmed by a given end user (such as the designer of the logic functions to be configured into the FPGA) to efficiently perform specific computing tasks. The FPGA's programmable logic configuration (i.e., the functionality implemented by the FPGA) is typically determined in the field by the end user. Unlike fixed-function designs such as general-purpose microprocessors, a given product FPGA can operate at varying frequencies from one FPGA logic configuration to another. This is due to the type of logic functions programmed into the FPGA. Some functions are more complex than others and have more speed-limiting paths. Typical FPGAs are tested and set to a certain voltage based on the FPGA's worst-case functionalities, as specified in the device specification. Even when the FPGA's specific functional configuration can operate at a lower frequency, supplying the same operating voltage to the FPGA can result in a higher voltage than is actually required. This situation results in increased power consumption and / or reduced performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] Figure 1 is a block diagram of a portion of a system according to an embodiment of the present invention.
[0004] Figure 2 is a block diagram of a processor according to an embodiment of the present invention.
[0005] Figure 3 is a block diagram of a multi-domain processor according to another embodiment of the present invention.
[0006] Figure 4 is an embodiment of a processor comprising multiple cores.
[0007] Figure 5 is a block diagram of a microarchitecture of a processor core according to one embodiment of the present invention.
[0008] Figure 6 is a block diagram of a microarchitecture of a processor core according to another embodiment.
[0009] Figure 7 is a block diagram of a microarchitecture of a processor core according to yet another embodiment.
[0010] Figure 8 is a block diagram of a microarchitecture of a processor core according to yet another embodiment.
[0011] Figure 9is a block diagram of a processor according to another embodiment of the present invention.
[0012] Figure 10 is a block diagram of a representative SoC according to an embodiment of the present invention.
[0013] Figure 11 is a block diagram of another example SoC according to an embodiment of the present invention.
[0014] Figure 12 is a block diagram of an example system with which embodiments may be used.
[0015] Figure 13 is a block diagram of another example system with which embodiments may be used.
[0016] Figure 14 is a block diagram of a representative computer system.
[0017] Figure 15 is a block diagram of a system according to an embodiment of the present invention.
[0018] Figure 16 is a block diagram illustrating an IP core development system used to fabricate integrated circuits to perform operations according to an embodiment.
[0019] Figure 17 is a block diagram of a system according to an embodiment.
[0020] Figure 18 is a flowchart of a method according to an embodiment of the present invention.
[0021] Figure 19 is a graphical illustration of a voltage-frequency curve according to an embodiment.
[0022] Figure 20 is a block diagram of a system according to another embodiment of the present invention. DETAILED DESCRIPTION
[0023] In various embodiments, an FPGA or other programmable control logic can be dynamically controlled to operate at appropriate voltage / frequency (VF) values for specific functions configured into the FPGA. More specifically, embodiments provide techniques that enable identification of the minimum operating frequency (Vmin) at which specific functions configured into the FPGA can operate correctly and safely. In this way, FPGAs according to embodiments can reduce power consumption while ensuring proper operation and longevity of the FPGA.
[0024] In some embodiments, a semi-static configuration can be provided, where the FPGA is tested offline with respect to specific functions loaded into the FPGA. In some cases, this testing can be one or more high-volume manufacturing (HVM) test routines used to test the FPGA during manufacturing. In other cases, this testing can be provided by a synthesis tool, such as an electronic design automation (EDA) synthesis tool that generates the configuration for the FPGA. Based on this testing, a VF curve is generated that provides a set of operating points, each of which includes a given operating frequency and operating voltage (either a full curve or a simplified offset to a specified worst-case scenario). This VF curve can then be provided to the FPGA along with the logic functions that configure the device. In this semi-static configuration arrangement, the responsibility for testing and verification lies with the designer who developed the logic functions.
[0025] In other embodiments, dynamic testing can be performed through FPGA self-testing during configuration loading. In such cases, this self-testing, which can be defined by the designer of the FPGA function, is executed to enable dynamic determination of the operating voltage. For example, this can be accomplished by loading a design for test (DFT) function into the FPGA for execution on the FPGA's self-test circuitry. This DFT function can represent the worst-case speed path for the function. As an example, for DFT creation, the synthesis tool can automatically generate logic cones associated with the critical path, which can itself be used as a Vmin search self-test pattern.
[0026] In an embodiment, the VF curve generated based on the self-test is used throughout the runtime of the FPGA until a new function is configured. Note that in some cases, if a given logic function does not have associated configured VF curve information or self-test information, a default voltage value as determined during HVM may be used (and which may be stored in non-volatile storage included in or associated with the FPGA). Note that this default voltage value may be at a relatively high level to ensure correct operation in worst-case corner cases.
[0027] As such, in embodiments herein, bitstream information loaded into the FPGA during the boot routine can be used to detect specific FPGA configuration scenarios and identify appropriate operating voltages. In embodiments, the appropriate operating voltage can be the minimum operating voltage at which logic functions correctly execute. To this end, the FPGA can include control circuitry or other control logic to control (internal or external) voltage regulators based on specific cells and specific scenarios. The FPGA boot routine occurs when the FPGA is reset, during which a bitstream (created offline using an EDA synthesis tool) is loaded. In addition to this bitstream, metadata is appended to the bitstream to provide information used to determine the appropriate operating voltage for a specific FPGA. In embodiments, this metadata can include critical path lengths and one or more diagnostic tests to be performed within the FPGA (after reset) to identify the appropriate Vmin for a specific configuration.
[0028] Although the following embodiments are described with reference to energy conservation and efficiency in specific integrated circuits (such as in computing platforms or processors), other embodiments may also be applied to other types of integrated circuits and logic devices. Similar techniques and teachings of the embodiments described herein can be applied to other types of circuits or semiconductor devices that can also benefit from improved energy efficiency and energy conservation. For example, the disclosed embodiments are not limited to any particular type of computer system. That is, the disclosed embodiments can be used in many different system types, ranging from server computers (e.g., tower servers, rack servers, blade servers, microservers, etc.), communication systems, storage systems, desktop computers of any configuration, laptop computers, notebook computers, and tablet computers (including 2:1 tablet devices, phablets, etc.), and can also be used in other devices (such as handheld devices, system-on-chips (SoCs), and embedded applications). Some examples of handheld devices include cellular phones such as smartphones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and handheld PCs. Embedded applications may typically include microcontrollers, digital signal processors (DSPs), network computers (NetPCs), set-top boxes, network hubs, wide area network (WAN) switches, wearable devices, or any other system that can implement the functions and operations taught below. In particular, embodiments may be implemented in mobile terminals with standard voice functionality (such as mobile phones, smartphones, and tablet phones) and / or in non-mobile terminals without standard wireless voice communication capabilities (such as many wearable devices, tablet devices, laptops, desktop computers, microservers, servers, etc.). In addition, the apparatus, methods, and systems described herein are not limited to physical computing devices, but may also involve software optimization for energy conservation and efficiency. As will become apparent in the description below, embodiments of the methods, apparatus, and systems described herein (whether referring to hardware, firmware, software, or a combination thereof) are critical to the "green technology" future, such as power conservation and energy efficiency in products that encompass a large portion of the U.S. economy.
[0029] Now refer to Figure 1 , which is a block diagram of a part of a system according to an embodiment of the present invention. Figure 1 As shown in FIG, system 100 may include various components, including a processor 110, which is a multi-core processor as shown. Processor 110 may be coupled to power supply 150 via external voltage regulator 160, which may perform a first voltage conversion to provide a main regulated voltage V to processor 110. reg .
[0030] As can be seen, the processor 110 may include multiple cores 120 a -120 nThe processor 110 may also include multiple independent FPGA functional units 122 a -122 n (each of which may be a given programmable logic circuit), each of which may be generated in a programmable manner via a bitstream and controlled to operate at a given minimum operating voltage as described herein. In addition, each core (and FPGA) may be coupled with an integrated voltage regulator (IVR) 125 a -125 n In association with an integrated voltage regulator, the integrated voltage regulator receives a main regulated voltage and generates an operating voltage to be provided to one or more agents of the processor associated with the IVR. Accordingly, an IVR implementation can be provided that allows fine-grained control over the voltage, and therefore the power and performance, of each individual core. Thus, each core can operate at an independent voltage and frequency, enabling significant flexibility and providing broad opportunities to balance power consumption with performance. In some embodiments, the use of multiple IVRs enables grouping components into separate power planes, such that power is regulated by the IVRs and supplied to only those components in the group. During power management, when a processor is placed into a low-power state, a given power plane of one IVR can be powered down or removed, while another power plane of another IVR remains active or fully powered. Similarly, cores 120 can include or be associated with independent clock generation circuitry, such as one or more phase-locked loops (PLLs), to independently control the operating frequency of each core 120.
[0031] Still refer to Figure 1 , additional components may be present within the processor, including an input / output interface (IF) 132, another interface 134, and an integrated memory controller (IMC) 136. As seen, each of these components may be powered by another integrated voltage regulator 125 x In one embodiment, the interface 132 may enable operation for an Intel® Quick Path Interconnect (QPI) interconnect that provides a point-to-point (PtP) link in a cache coherence protocol that includes multiple layers, including a physical layer, a link layer, and a protocol layer. Furthermore, the interface 134 may be connected via a peripheral component interconnect express (PCIe) TM ) protocol to communicate.
[0032] Also shown is a power control unit (PCU) 138, which may include circuitry, including hardware, software, and / or firmware, to implement power management operations for processor 110. As seen, PCU 138 provides control information to external voltage regulator 160 via digital interface 162 to cause the voltage regulator to generate an appropriate regulated voltage. PCU 138 also provides control information to IVR 125 via another digital interface 163 to control the generated operating voltage (or to disable the corresponding IVR in a low-voltage mode). In various embodiments, PCU 138 may include various power management logic or circuitry to implement hardware-based power management. Such power management may be entirely processor-controlled (e.g., through various processor hardware, and may be triggered by workload and / or power, thermal, or other processor constraints) and / or may be implemented in response to external sources (e.g., platform or power management sources or system software). PCU 138 may be configured to initiate a self-test of one or more FPGA programmable logic circuits to identify the minimum operating voltage to be used during normal operation of processor 110.
[0033] exist Figure 1 , PCU 138 is illustrated as being presented as a separate circuit from the processor. In other cases, PCU 138 may execute on a given core or cores 120. In some cases, PCU 138 may be implemented as a microcontroller (dedicated or general purpose) or other control logic configured to execute its own dedicated power management code, sometimes referred to as P-code. In yet another embodiment, the power management operations to be performed by PCU 138 may be implemented external to the processor, such as by a separate power management integrated circuit (PMIC) or other component external to the processor. In yet another embodiment, the power management operations to be performed by PCU 138 may be implemented within BIOS or other system software.
[0034] Embodiments may be particularly well-suited for multi-core processors, where each of the multiple cores may operate at an independent voltage and frequency point. As used herein, the term "domain" is used to refer to a collection of hardware and / or logic that operates at the same voltage and frequency point. In addition, a multi-core processor may further include other non-core processing engines, such as fixed function units, graphics engines, and the like. Such a processor may include independent domains rather than cores, such as one or more domains associated with a graphics engine (referred to herein as graphics domains) and one or more domains associated with non-core circuits (referred to herein as non-core or system agents). While many implementations of a multi-domain processor may be formed on a single semiconductor die, other implementations may also be implemented through multi-chip packaging, where different domains may exist on different semiconductor dies in a single package.
[0035] Although not shown for ease of illustration, it is understood that additional components may be present within the processor 110 such as uncore logic and other components such as internal memory (e.g., one or more levels of cache memory hierarchy, etc.). Figure 1 For example, other regulated voltages may be provided to the on-chip source from external voltage regulator 160 or one or more additional external regulated voltage sources.
[0036] Note that the power management techniques described herein can be independent of and complementary to operating system (OS)-based power management (OSPM) mechanisms. According to one example OSPM technique, a processor can operate in various performance states or levels (so-called P-states, i.e., from P0 to Pn). Typically, the P1 performance state may correspond to the highest guaranteed performance state that can be requested by the OS. Beyond this P1 state, the OS may further request a higher performance state, namely the P0 state. The P0 state can thus be an opportunistic, overclocked, or turbo mode state, in which the processor hardware can configure the processor, or at least portions thereof, to operate at a higher frequency than the guaranteed frequency when power and / or thermal budget is available. In many implementations, a processor may include multiple so-called bin frequencies above the P1 guaranteed maximum frequency, exceeding the maximum peak frequency of a particular processor, as incorporated into the processor during manufacturing or otherwise programmed into the processor. Furthermore, according to an OSPM mechanism, a processor can operate in various power states or levels. Regarding power states, the OSPM mechanism may specify different power consumption states, commonly referred to as C-states, C0, C1, through Cn. When a core is active, it runs in the C0 state, and when a core is idle, it can be placed in a core low-power state, also referred to as a core non-zero C-state (e.g., C1-C6 states), where each C-state is at a lower power consumption level (such that C6 is a deeper low-power state than C1, etc.). Note that using the degradation-based techniques herein, C-state requests from the OS can be overridden based, at least in part, on guest adjustment information provided via the interface as described herein.
[0037] It is understood that many different types of power management techniques may be used, alone or in combination with different embodiments. As a representative example, a power controller may control a processor to be power managed by some form of dynamic voltage frequency scaling (DVFS), in which the operating frequency and / or operating voltage of one or more cores or other processor logic may be dynamically controlled to reduce power consumption in certain situations. In an example, Enhanced Intel SpeedStep available from Intel Corporation (Santa Clara, CA) may be used. TM In another example, Intel TurboBoost can be used to implement DVFS to provide optimal performance at the lowest power consumption level. TM DVFS is implemented using techniques to enable one or more cores or other compute engines to operate at a higher frequency than the guaranteed operating frequency based on conditions (eg, workload and availability).
[0038] Another power management technique that can be used in some examples is the dynamic swapping of workloads between different compute engines. For example, a processor may include asymmetric cores or other processing engines that operate at different power consumption levels, so that in a power-constrained situation, one or more workloads can be dynamically swapped to execute on a lower-power core or other compute engine. Another exemplary power management technique is hardware duty cycling (HDC), which can cause cores and / or other compute engines to be periodically enabled and disabled according to a duty cycle, so that one or more cores can be rendered inactive during an inactive period of the duty cycle and rendered active during an active period of the duty cycle.
[0039] Power management techniques may also be used when constraints exist in the operating environment. For example, when power and / or thermal constraints are encountered, power may be reduced by reducing the operating frequency and / or voltage. Other power management techniques include throttling the instruction execution rate or limiting instruction scheduling. Furthermore, it may be possible for the instructions of a given instruction set architecture to include explicit or implicit indications regarding power management operations. Although described using these specific examples, it will be understood that many other power management techniques may also be used in certain embodiments.
[0040] Embodiments may be implemented in processors targeted at various markets including server processors, desktop processors, mobile processors, etc. Referring now to Figure 2 , which shows a block diagram of a processor according to an embodiment of the present invention. Figure 2 As shown in FIG, the processor 200 may include multiple cores 210 a -210 n and multiple FPGA programmable logic circuits 212 a -212n In one embodiment, each such core may have an independent power domain and may be configured to enter and exit an active state and / or a maximum performance state based on workload. One or more cores 210 may be heterogeneous with the other cores, e.g., having different microarchitectures, instruction set architectures, pipeline depths, power and performance capabilities. The various cores may be coupled to a system agent or uncore 220 comprising various components via interconnect 215. As seen, the uncore 220 may include a shared cache 230 which may be a last level cache. Additionally, the uncore may include an integrated memory controller 240 to communicate with system memory (e.g., via a memory bus) for example. Figure 2 Uncore 220 also includes various interfaces 250 and a power control unit 255, which may include logic for implementing power management techniques. Furthermore, PCU 255 may include a BIST circuit 256 configured to identify the minimum operating voltage of each FPGA programmable logic circuit 212.
[0041] In addition, through the interfaces 250a-250n, connections can be made to various off-chip components such as peripheral devices, mass storage devices, etc. Figure 2 The embodiments are shown with this particular implementation, but the scope of the present invention is not limited in this regard.
[0042] Now refer to Figure 3 , shows a block diagram of a multi-domain processor according to another embodiment of the present invention. Figure 3 In the embodiment shown in FIG. 3 , the processor 300 includes multiple domains. Specifically, the core domain 310 may include multiple cores 310a-310n, multiple FPGA programmable logic circuits 312 a -312 n , the graphics domain 320 may include one or more graphics engines and a system agent domain 350 may further exist. In some embodiments, the system agent domain 350 may execute at an independent frequency compared to the core domain and may always remain powered on to handle power control events and power management so that domains 310 and 320 can be controlled to dynamically enter and exit high-power states and low-power states. Each of domains 310 and 320 can operate at different voltages and / or powers. Note that although shown with only three domains, it is to be understood that the scope of the present invention is not limited to this point and additional domains may exist in other embodiments. For example, multiple core domains may each be presented as including at least one core.
[0043] Typically, in addition to the various execution units and additional processing elements, each core 310 may further include a low-level cache. Furthermore, the various cores may be coupled to each other and to a shared cache memory formed by multiple units of last-level cache (LLC) 340a-340n. In various embodiments, LLC 340 may be shared among the cores and the graphics engine, as well as various media processing circuits. As can be seen, ring interconnect 330 thus couples the cores together and provides interconnection between the cores, graphics domain 320, and system agent circuitry 350. In one embodiment, interconnect 330 may be part of the core domain. However, in other embodiments, the ring interconnect may have its own domain.
[0044] As further seen, the system agent domain 350 may include a display controller 352 that can provide control of and interface with an associated display. As further seen, the system agent domain 350 may include a power control unit 355 that may include logic for implementing power management techniques, including identifying a minimum operating voltage as described herein. To this end, the PCU 355 includes a BIST circuit 356 that is configured to identify the minimum operating voltage of each FPGA programmable logic circuit 312, as described herein.
[0045] As in Figure 3 As further seen in FIG, the processor 300 may further include an integrated memory controller (IMC) 370, which may provide an interface to system memory, such as dynamic random access memory (DRAM). A plurality of interfaces 380a-380n may be present to enable interconnection between the processor and other circuits. For example, in one embodiment, at least one direct media interface (DMI) interface and one or more PCIe interfaces may be provided. TM Furthermore, to provide communication between other agents (such as additional processors) or other circuits, one or more QPI interfaces may also be provided. Figure 3 The embodiments are shown at this high level, but understand the scope of the present invention is not limited in this regard.
[0046] Reference Figure 4, illustrates an embodiment of a processor including multiple cores. Processor 400 comprises any processor or processing device, such as a microprocessor, an embedded processor, a digital signal processor (DSP), a network processor, a handheld processor, an application processor, a coprocessor, a system on a chip (SoC), or other device for executing code. In one embodiment, processor 400 comprises at least two cores—cores 401 and 402—which may include asymmetric cores or symmetric cores (the illustrated embodiment). However, processor 400 may include any number of processing elements that may be symmetric or asymmetric.
[0047] In one embodiment, a processing element refers to hardware or logic used to support a software thread. Examples of hardware processing elements include: thread units, thread slots, threads, processing units, contexts, context units, logical processors, hardware threads, cores, and / or any other element that can maintain state (such as execution state or architectural state) for a processor. In other words, in one embodiment, a processing element refers to any hardware that can be independently associated with code, such as a software thread, an operating system, an application, or other code. A physical processor typically refers to an integrated circuit that potentially includes any number of other processing elements, such as cores or hardware threads.
[0048] A core generally refers to logic located on an integrated circuit that is capable of maintaining an independent architectural state, where each independently maintained architectural state is associated with at least some dedicated execution resources. In contrast to a core, a hardware thread generally refers to any logic located on an integrated circuit that is capable of maintaining an independent architectural state, where the independently maintained architectural states share access to execution resources. As can be seen, the line between the nomenclature of a core and a hardware thread overlaps when some resources are shared and other resources are dedicated to an architectural state. It is also common for cores and hardware threads to be viewed by an operating system as individual logical processors, where the operating system can schedule operations on each logical processor separately.
[0049] As in Figure 4 As illustrated in FIG, physical processor 400 includes two cores, core 401 and 402. Herein, cores 401 and 402 are considered symmetric cores, i.e., cores having the same configuration, functional units, and / or logic. In another embodiment, core 401 includes an out-of-order processor core, and core 402 includes an in-order processor core. However, cores 401 and 402 may be individually selected from any type of core, such as a native core, a software managed core, a core adapted to execute a native instruction set architecture (ISA), a core adapted to execute a translated ISA, a co-designed core, or other known core. Still further discussed, the functional units illustrated in core 401 are described in further detail below, with the units in core 402 operating in a similar manner.
[0050] As depicted, core 401 includes two hardware threads 401a and 401b, which may also be referred to as hardware thread slots 401a and 401b. Thus, in one embodiment, a software entity (such as an operating system) potentially views processor 400 as four separate processors, i.e., four logical processors or processing elements capable of executing four software threads simultaneously. As implied above, a first thread is associated with architecture state register 401a, a second thread is associated with architecture state register 401b, a third thread may be associated with architecture state register 402a, and a fourth thread may be associated with architecture state register 402b. Each of the architecture state registers (401a, 401b, 402a, and 402b) may be referred to herein as a processing element, thread slot, or thread unit, as described above. As illustrated, architecture state register 401a is replicated in architecture state register 401b so that separate architecture states / contexts can be stored for logical processors 401a and 401b. In core 401, other smaller resources (such as the instruction pointer and renaming logic in the allocator and renamer block 430) can also be replicated for threads 401a and 401b. Some resources (such as the reorder buffer, branch target buffer, and instruction translation lookaside buffer (BTB and I-TLB) 420 in the reorder / retirement unit 435, load / store buffers, and queues) can be shared through partitioning. Other resources (such as general internal registers, page table base register(s), low-level data cache and data TLB 450, execution unit(s) 440, and portions of the out-of-order unit 435) are potentially fully shared.
[0051] Processor 400 typically includes other resources that may be fully shared, shared by partitioning, or dedicated by / to processing elements. Figure 4 , a purely exemplary processor embodiment utilizing illustrative logic units / resources of the processor is illustrated. Note that the processor may include or omit any of these functional units, and may include any other known functional units, logic, or firmware not depicted. As illustrated, core 401 comprises a simplified, representative out-of-order (OOO) processor core. However, in-order processors may be utilized in different embodiments. The OOO core includes a branch target buffer 420 for predicting branches to be executed / taken, and an instruction translation buffer (I-TLB) 420 for storing address translation entries for instructions.
[0052] Core 401 further includes a decode module 425 coupled to the instruction fetch unit to decode fetched elements. In one embodiment, the instruction fetch logic includes sequencers associated with thread slots 401a and 401b, respectively. Typically, core 401 is associated with a first ISA, which defines / specifies instructions executable on processor 400. Typically, machine code instructions that are part of the first ISA include a portion of the instruction (called an opcode) that references / specifies the instruction or operation to be performed. Decode logic 425 includes circuitry that distinguishes these instructions from their opcodes and passes the decoded instructions through the pipeline for processing as defined by the first ISA. For example, in one embodiment, decoder 425 includes logic designed / adapted to recognize specific instructions (such as transactional instructions). As a result of the recognition by decoder 425, the architecture or core 401 takes specific, predefined actions to perform the tasks associated with the appropriate instructions. It is important to note that any of the tasks, blocks, operations, and methods described herein can be performed in response to a single or multiple instructions; some of which may be new or legacy instructions.
[0053] In one example, allocator and renamer block 430 includes an allocator to reserve resources, such as a register file for storing instruction processing results. However, threads 401a and 401b are potentially capable of out-of-order execution, with allocator and renamer block 430 also reserving other resources, such as a reorder buffer, to track instruction results. Unit 430 may also include a register renamer to rename program / instruction reference registers to other registers within processor 400. Reorder / retirement unit 435 includes components, such as the reorder buffer, load buffer, and store buffer mentioned above, to support out-of-order execution and subsequent in-order retirement of instructions executed out-of-order.
[0054] In one embodiment, scheduler and execution unit(s) block 440 includes a scheduler unit for scheduling instructions / operations on execution units. For example, floating-point instructions are scheduled on ports of execution units that have available floating-point execution units. Register files associated with the execution units are also included to store information about instruction processing results. Exemplary execution units include floating-point execution units, integer execution units, jump execution units, load execution units, store execution units, and other known execution units.
[0055] A lower-level data cache and data translation lookaside buffer (D-TLB) 450 are coupled to the execution unit(s) 440. The data cache is to store recently used / operated data (such as data operands) on the unit, which is potentially kept in a memory coherent state. The D-TLB is to store the most recent virtual / linear to physical address translations. As a specific example, the processor may include a page table structure to decompose physical memory into multiple virtual pages.
[0056] Here, cores 401 and 402 share access to a higher-level or further-out cache 410 that is to cache recently fetched elements. Note that higher-level or further-out refers to cache levels that increase or become further away from the execution unit(s). In one embodiment, higher-level cache 410 is a last-level data cache—the last level in the memory hierarchy on processor 400—such as a level 2 or level 3 data cache. However, higher-level cache 410 is not so limited, as it can be associated with or include an instruction cache. A trace cache (a type of instruction cache) can alternatively be coupled after decoder 425 to store recently decoded traces.
[0057] In the depicted configuration, processor 400 also includes a bus interface module 405 and a power control unit 460, which can implement power management according to embodiments of the present invention. In this case, bus interface 405 is to communicate with devices external to processor 400, such as system memory and other components.
[0058] Memory controller 470 can interface with other devices such as one or more memories. In one example, bus interface 405 includes a ring interconnect having a memory controller for interfacing with the memories and a graphics controller for interfacing with the graphics processor. In an SoC environment, even more devices such as network interfaces, coprocessors, memories, graphics processors, and any other known computer devices / interfaces can be integrated on a single die or integrated circuit to provide a small form factor with high functionality and low power consumption.
[0059] Now refer to Figure 5 , which shows a block diagram of a micro-architecture of a processor core according to an embodiment of the present invention. Figure 5 As shown in FIG, processor core 500 may be a multi-stage pipelined out-of-order processor. Core 500 may operate at various voltages based on a received operating voltage, which may be received from an integrated voltage regulator or an external voltage regulator.
[0060] As in Figure 5As seen in FIG, core 500 includes a front end unit 510 that can be used to fetch instructions to be executed and prepare them for later use in the processor pipeline. For example, the front end unit 510 can include an instruction fetch unit 501, an instruction cache 503, and an instruction decoder 505. In some implementations, the front end unit 510 can further include a trace cache, along with a microcode storage device and a micro-operation storage device. The instruction fetch unit 501 can fetch macroinstructions, for example, from memory or the instruction cache 503 and feed them to the instruction decoder 505 to decode them into primitives (i.e., micro-operations for execution by the processor).
[0061] Coupled between the front end unit 510 and the execution unit 520 is an out-of-order (OOO) engine 515 that can be used to receive microinstructions and prepare them for execution. More specifically, the OOO engine 515 may include various buffers to reorder the microinstruction stream and allocate various resources required for execution, as well as provide renaming of logical registers to storage locations within various register files such as the register file 530 and the extended register file 535. The register file 530 may include separate register files for integer and floating-point operations. For configuration, control, and additional operation purposes, a set of machine-specific registers (MSRs) 538 may also exist and be accessible to various logic within the core 500 (and outside the core).
[0062] Various resources may be present in execution units 520, including, for example, various integer, floating point, and single instruction multiple data (SIMD) logic units, in addition to other specific hardware. For example, such execution units may include one or more arithmetic logic units (ALUs) 522 and one or more vector execution units 524, in addition to other such execution units.
[0063] Results from the execution units can be provided to retirement logic, namely, a re-order buffer (ROB) 540. More specifically, ROB 540 can include various arrays and logic to receive information associated with executed instructions. This information is then examined by ROB 540 to determine whether the instruction can be effectively retired and the result data committed to the processor's architectural state, or whether one or more exceptions have occurred that prevent the instruction from being properly retired. Of course, ROB 540 can also handle other operations associated with retirement.
[0064] As in Figure 5As shown in FIG, ROB 540 is coupled to cache 550, which in one embodiment may be a low level cache (e.g., L1 cache), although the scope of the present invention is not limited in this respect. Furthermore, execution unit 520 may be directly coupled to cache 550. From cache 550, data communication may occur with higher level caches, system memory, and so forth. Although in Figure 5 The embodiments are shown at this high level, it is to be understood that the scope of the present invention is not limited in this regard. For example, although Figure 5 The implementations described herein are with respect to out-of-order machines such as those having the Intel® x86 Instruction Set Architecture (ISA), but the scope of the invention is not limited in this respect. That is, other embodiments may be implemented on an in-order processor, a reduced instruction set computing (RISC) processor (such as an ARM-based processor), or a processor of another type of ISA that can emulate instructions and operations of a different ISA via an emulation engine and associated logic circuitry.
[0065] Now refer to Figure 6 , which shows a block diagram of a micro-architecture of a processor core according to another embodiment. Figure 6 In an embodiment, core 600 may be a low-power core of a different microarchitecture, such as an Intel® Atom-based core with a relatively limited pipeline depth designed to reduce power consumption. TM As seen, core 600 includes an instruction cache 610 coupled to provide instructions to an instruction decoder 615. Branch predictor 605 may be coupled to instruction cache 610. Note that instruction cache 610 may be further coupled to another level of cache memory, such as an L2 cache (in Figure 6 In turn, instruction decoder 615 provides decoded instructions to issue queue (IQ) 620 for storage and delivery to a given execution pipeline. Microcode ROM 618 is coupled to instruction decoder 615.
[0066] The floating-point pipeline 630 includes a floating-point (FP) register file 632, which may include a plurality of architectural registers of a given bit width (such as 128, 256, or 512 bits). The pipeline 630 includes a floating-point scheduler 634 to schedule instructions for execution on one of the pipeline's multiple execution units. In the illustrated embodiment, such execution units include an ALU 635, a shuffle unit 636, and a floating-point adder 638. In turn, results generated in these execution units may be provided back to registers and / or buffers of the register file 632. It is understood that, while illustrated using these example execution units, additional or different floating-point execution units may be present in alternative embodiments.
[0067] An integer pipeline 640 may also be provided. In the illustrated embodiment, pipeline 640 includes an integer (INT) register file 642, which may include a plurality of architectural registers of a given bit width (such as 128 or 256 bits). Pipeline 640 includes an integer execution (IE) scheduler 644, which is used to schedule instructions for execution on one of the pipeline's multiple execution units. In the illustrated embodiment, such execution units include an ALU 645, a shifter unit 646, and a jump execution unit (JEU) 648. Furthermore, results generated in these execution units may be provided back to buffers and / or registers of register file 642. It is understood that, while illustrated using these few example execution units, additional or different integer execution units may be present in alternative embodiments.
[0068] A memory execution (ME) scheduler 650 may schedule memory operations for execution in an address generation unit (AGU) 652, which is also coupled to a TLB 654. As seen, these structures may be coupled to a data cache 660, which may be an L0 and / or L1 data cache that is in turn coupled to additional levels of the cache memory hierarchy, including an L2 cache memory.
[0069] To provide support for out-of-order execution, an allocator / renamer 670 may be provided in addition to the reorder buffer 680 and configured to reorder instructions executed out-of-order for in-order retirement. Figure 6 The diagrams of FIG are shown using this particular pipeline architecture, but it will be understood that many variations and alternatives are possible.
[0070] Note that in cases with asymmetric kernels (such as those according to Figure 5 and 6 In processors with a 1.1-GHz microarchitecture, workloads can be dynamically swapped between cores for power management reasons, since these cores can have the same or related ISAs, albeit with different pipeline designs and depths. This dynamic core swapping can be performed in a manner that is transparent to user applications (and possibly the kernel).
[0071] Reference Figure 7 , shows a block diagram of a micro-architecture of a processor core according to another embodiment. Figure 7As illustrated in FIG, core 700 may include a multi-stage in-order pipeline to execute at very low power consumption levels. As one such example, processor 700 may have a microarchitecture based on the ARM Cortex A53 design available from ARM Holdings, Inc. (Sunnyvale, CA). In an implementation, an 8-stage pipeline configured to execute 32-bit code and 64-bit code may be provided. Core 700 includes an instruction fetch unit 710 configured to fetch instructions and provide them to a decode unit 715, which may decode instructions, such as macroinstructions of a given ISA such as the ARMv8 ISA. It is further noted that a queue 730 may be coupled to the decode unit 715 to store decoded instructions. The decoded instructions are provided to issue logic 725, where the decoded instructions may be issued to a given one of a plurality of execution units.
[0072] Further references Figure 7 , issue logic 725 may issue instructions to one of a plurality of execution units. In the illustrated embodiment, these execution units include an integer unit 735, a multiplication unit 740, a floating point / vector unit 750, a dual issue unit 760, and a load / store unit 770. The results of these different execution units may be provided to a write-back (WB) unit 780. It is to be understood that while a single write-back unit is shown for ease of illustration, in some implementations a separate write-back unit may be associated with each execution unit. Furthermore, it is to be understood that while in Figure 7 Each of the units and logic shown in FIG is represented at a high level, but certain embodiments may include more or different structures. Figure 7 A processor designed with one or more pipelined cores can be implemented in many different end products, ranging from mobile devices to server systems.
[0073] Reference Figure 8 , shows a block diagram of a micro-architecture of a processor core according to yet another embodiment. Figure 8 As shown in FIG, core 800 may include a multi-stage, multi-issue, out-of-order pipeline to operate at very high performance levels (which may be Figure 7As one such example, the processor 800 may have a microarchitecture designed according to the ARM Cortex A57. In an implementation, a 15 (or more) stage pipeline configured to execute both 32-bit and 64-bit code may be provided. In addition, the pipeline may provide for 3 (or more) wide and 3 (or more) issue operations. The core 800 includes an instruction fetch unit 810 configured to fetch instructions and provide them to a decoder / renamer / dispatcher unit 815 coupled to a cache 820. The unit 815 may decode instructions (e.g., macroinstructions of the ARMv8 instruction set architecture), rename register references within the instructions, and (ultimately) dispatch the instructions to a selected execution unit. The decoded instructions may be stored in a queue 825. Note that although the instructions are not shown in the queue 825 for ease of illustration, the instructions are not shown in the queue 825. Figure 8 A single queue structure is shown in , but it will be appreciated that separate queues may be provided for each of a number of different types of execution units.
[0074] exist Figure 8 Also shown is issue logic 830, according to which decoded instructions stored in queue 825 can be issued to selected execution units. In certain embodiments, issue logic 830 can also be implemented using separate issue logic for each of the multiple different types of execution units to which issue logic 830 is coupled.
[0075] The decoded instruction may be issued to a given one of a plurality of execution units. In the illustrated embodiment, these execution units include one or more integer units 835, multiplication units 840, floating point / vector units 850, branch units 860, and load / store units 870. In an embodiment, the floating point / vector unit 850 may be configured to handle 128 or 256 bits of vector data or SIMD. Further, the floating point / vector execution unit 850 may perform IEEE-754 double precision floating point operations. The results of these different execution units may be provided to a write back unit 880. Note that in some implementations, a separate write back unit may be associated with each of the execution units. Furthermore, it is to be understood that while in Figure 8 Each of the elements and logic shown in FIG are represented at a high level, but specific implementations may include more or different structures.
[0076] Note that in cases with asymmetric kernels (such as those according to Figure 7 and 8In processors with a microarchitecture (e.g., a 1.1-GHz microarchitecture), workloads can be dynamically swapped for power management reasons because the cores can have the same or related ISAs despite having different pipeline designs and depths. Such dynamic core swapping can be performed in a manner that is transparent to user applications (and possibly the kernel).
[0077] Use as in Figure 5-8 Any one or more of the processors designed with one or more pipeline cores can be implemented in many different end products, ranging from mobile devices to server systems. Figure 9 , which shows a block diagram of a processor according to another embodiment of the present invention. Figure 9 In an embodiment of the present invention, the processor 900 may be a SoC including multiple domains, each of which may be controlled to operate at an independent operating voltage and operating frequency. As a specific illustrative example, the processor 900 may be an Intel® Architecture Core i7 processor. TM processor, such as an i3, i5, i7, or another such processor available from Intel Corporation. However, other low-power processors (such as those available from Advanced Micro Devices, Inc. (Sunnyvale, CA), an ARM-based design from ARM Holdings plc or its licensees, or a MIPS-based design from MIPS Technologies, Inc. (Sunnyvale, CA) or its licensors or purchasers) may instead be present in other embodiments, such as an Apple A7 processor, a Qualcomm Snapdragon processor, or a Texas Instruments OMAP processor. Such a SoC may be used in low-power systems such as smartphones, tablets, phablets, ultrabooks, and similar systems. TM A computer or other portable computing device that may incorporate a heterogeneous system architecture having a processor design based on the heterogeneous system architecture.
[0078] exist Figure 9In the high-level view shown in FIG, processor 900 includes multiple core units 910a-910n. Each core unit may include one or more processor cores, one or more cache memories, and other circuitry. Each core unit 910 may support one or more instruction sets (e.g., the x86 instruction set (with some extensions added by newer versions)), the MIPS instruction set, the ARM instruction set (with optional additional extensions such as NEON), or other instruction sets, or a combination thereof. Note that some of the core units may be heterogeneous resources (e.g., having different designs). Furthermore, each such core may be coupled to a cache memory (not shown), which in one embodiment may be a shared level 2 (L2) cache memory. Non-volatile storage device 930 may be used to store various programs and other data. For example, this storage device may be used to store at least a portion of microcode, boot information (such as BIOS), other system software, and so on.
[0079] Each core unit 910 may also include an interface such as a bus interface unit to enable interconnection with additional circuitry of the processor. In an embodiment, each core unit 910 is coupled to a coherence fabric that may serve as a primary cache coherence on-die interconnect, which in turn is coupled to a memory controller 935. In turn, the memory controller 935 controls communication with memory such as DRAM (in Figure 9 for ease of illustration (not shown).
[0080] In addition to the core units, there are additional processing engines within the processor, including at least one graphics unit 920, which may include one or more graphics processing units (GPUs) to perform graphics processing and possibly general-purpose operations on the graphics processor (so-called GPGPU operations). In addition, there may be at least one image signal processor 925. The signal processor 925 can be configured to process incoming image data received from one or more capture devices (internal to the SoC or external to the chip).
[0081] Other accelerators may also be present. Figure 9 In the illustration, a video encoder 950 can perform encoding operations, including encoding and decoding video information, for example, providing hardware acceleration support for high-definition video content. A display controller 955 can further be provided to accelerate display operations, including providing support for internal and external displays of the system. In addition, a security processor 945 can be present to perform security operations such as secure boot operations, various encryption operations, etc.
[0082] Each unit may have its power consumption controlled via a power manager 940 , which may include control logic to implement the various power management techniques described herein.
[0083] In some embodiments, the SoC 900 may further include a non-coherent fabric coupled to a coherent fabric to which various peripheral devices may be coupled. One or more interfaces 960a-960d enable communication with one or more off-chip devices. Such communication may be via various communication protocols such as PCIe TM ,GPIO,USB,I 2 C, UART, MIPI, SDIO, DDR, SPI, HDMI), in addition to other types of communication protocols. Figure 9 While the embodiments are shown at this high level, it will be understood that the scope of the invention is not limited in this regard.
[0084] Now refer to Figure 10 , a block diagram of a representative SoC is shown. In the illustrated embodiment, SoC 1000 may be a multi-core SoC configured for low-power operation that is optimized for incorporation into a smartphone or other low-power device such as a tablet or other portable computing device. As an example, SoC 1000 may be implemented using asymmetric or different types of cores, such as a combination of higher power cores and / or low power cores (e.g., out-of-order and in-order cores), and / or one or more FPGA programmable logic circuits. In various embodiments, these cores may be based on Intel® Architecture TM In another embodiment, a mixture of Intel cores and ARM cores can be implemented in a given SoC.
[0085] As in Figure 10 As seen in FIG, SoC 1000 includes a first core domain 1010 having a plurality of first cores 1012a-1012d. In an example, these cores may be low-power cores, such as in-order cores. In one embodiment, these first cores may be implemented as ARM Cortex A53 cores. In turn, these cores are coupled to a cache memory 1015 of the core domain 1010. In addition, SoC 1000 includes a second core domain 1020. Figure 10 In the illustration of FIG, the second core domain 1020 has a number of second cores 1022a-1022d. In an example, these cores may be cores with higher power consumption than the first core 1012. In an embodiment, the second cores may be out-of-order cores, which may be implemented as ARM Cortex A57 cores. In turn, these cores are coupled to the cache memory 1025 of the core domain 1020. Note that although Figure 10 The example shown in includes 4 cores in each domain, but it is understood that in other examples there may be more or fewer cores in a given domain.
[0086] Further references Figure 10 , a graphics domain 1030 is also provided, which may include one or more graphics processing units (GPUs) configured to independently execute graphics workloads (e.g., provided by one or more cores of core domains 1010 and 1020). As an example, GPU domain 1030 may be used to provide display support for various screen sizes in addition to providing graphics and display rendering operations.
[0087] As seen, the various domains are coupled to a coherent interconnect 1040, which in an embodiment may be a cache coherent interconnect fabric that is in turn coupled to an integrated memory controller 1050. In some examples, the coherent interconnect 1040 may include a shared cache memory, such as an L3 cache. In an embodiment, the memory controller 1050 may be a processor that provides multiple channels of communication with off-chip memory, such as multiple channels of DRAM (in an embodiment). Figure 10 A direct memory controller is not shown in the figure for simplicity of illustration).
[0088] In different examples, the number of core domains may vary. For example, for a low-power SoC suitable for incorporation into a mobile computing device, there may be a plurality of core domains such as Figure 10 1020. Furthermore, in such a low-power SoC, a core domain 1020 comprising higher-power cores may have a smaller number of such cores. For example, in one implementation, two cores 1022 may be provided to enable operation at a reduced power consumption level. Furthermore, the different core domains may also be coupled to an interrupt controller to enable dynamic swapping of workloads between the different domains.
[0089] In still other embodiments, there may be a larger number of core domains and additional optional IP logic, where the SoC can be scaled to higher performance (and power) levels for incorporation into other computing devices such as desktop computers, servers, high-performance computing systems, base stations, etc. As one such example, four core domains may be provided, each with a given number of out-of-order cores. Furthermore, in addition to optional GPU support (which may take the form of a GPGPU, for example), one or more accelerators may be provided to provide optimized hardware support for specific functions (e.g., network serving, network processing, switching, etc.). In addition, there may be input / output interfaces to couple these accelerators to off-chip components.
[0090] Now refer to Figure 11 , shows a block diagram of another example SoC. Figure 11In embodiments, SoC 1100 may include various circuits to enable high performance for multimedia applications, communications, and other functions. As such, SoC 1100 is suitable for incorporation into a wide variety of portable and other devices, such as smartphones, tablet computers, smart TVs, and the like. In the illustrated example, SoC 1100 includes a central processing unit (CPU) domain 1110. In embodiments, multiple individual processor cores may reside within CPU domain 1110. As an example, CPU domain 1110 may be a quad-core processor with four multi-threaded cores. Such processors may be homogeneous or heterogeneous, for example, a mix of low-power and high-power processor cores.
[0091] Furthermore, a GPU domain 1120 is provided to perform high-level graphics processing in one or more GPUs to handle graphics and compute APIs. A DSP unit 1130 may provide one or more low-power DSPs for handling low-power multimedia applications such as music playback, audio / video, etc., in addition to high-level computations that may occur during the execution of multimedia instructions. Furthermore, a communication unit 1140 may include various components to communicate via various wireless protocols such as cellular communications (including 3G / 4G LTE), wireless LAN protocols (such as Bluetooth), and the like. TM , IEEE 802.11, etc.) to provide connectivity.
[0092] Furthermore, the multimedia processor 1150 can be used to perform the capture and playback of high-definition video and audio content, including processing of user gestures. The sensor unit 1160 can include multiple sensors and / or sensor controllers for interfacing with various off-chip sensors present in a given platform. An image signal processor 1170 with one or more separate ISPs can be provided to perform image processing on content captured from one or more cameras of the platform (including still cameras and video cameras).
[0093] The display processor 1180 may provide support for connection to high-definition displays of a given pixel density, including the ability to wirelessly transmit content for playback on such displays. Further, the location unit 1190 may include a GPS receiver with support for multiple GPS satellite constellations to provide applications using such GPS receivers with highly accurate positioning information. It is understood that although in Figure 11 The examples shown in FIG. 5 are shown with this particular set of components, but many variations and alternatives are possible.
[0094] Now refer to Figure 12, shown is a block diagram of an example system with which embodiments may be used. As can be seen, system 1200 may be a smartphone or other wireless communicator. Baseband processor 1205 is configured to perform various signal processing with respect to communication signals to be transmitted from or received by the system. In turn, baseband processor 1205 is coupled to application processor 1210, which may be the system's main CPU to execute the OS and other system software, in addition to user applications such as many well-known social media and multimedia applications. Application processor 1210 may further be configured to perform various other computing operations for the device.
[0095] In turn, the application processor 1210 can be coupled to a user interface / display 1220, such as a touch screen display. Furthermore, the application processor 1210 can be coupled to a memory system, which includes non-volatile memory (i.e., flash memory 1230) and system memory (i.e., dynamic random access memory (DRAM) 1235). As further seen, the application processor 1210 is further coupled to a capture device 1240, such as one or more image capture devices that can record video and / or still images.
[0096] Still refer to Figure 12 A Universal Integrated Circuit Card (UICC) 1240, including a subscriber identity module and possibly a secure storage device and a cryptographic processor, is also coupled to the application processor 1210. System 1200 may further include a security processor 1250, which may be coupled to the application processor 1210. Multiple sensors 1225 may be coupled to the application processor 1210 to enable the input of various sensory information, such as accelerometer information and other environmental information. An audio output device 1295 may provide an interface for outputting sound, such as in the form of voice communications, played or streamed audio data, and the like.
[0097] As further illustrated, a near field communication (NFC) contactless interface 1260 is provided for transmitting in the NFC near field via an NFC antenna 1265. Figure 12 A single antenna is shown in FIG, but it is understood that in some implementations, one antenna or a different set of antennas may be provided to enable various wireless functionalities.
[0098] A power management integrated circuit (PMIC) 1215 is coupled to the application processor 1210 to implement platform-level power management. To this end, the PMIC 1215 can issue power management requests to the application processor 1210 to enter certain low-power states as desired. In addition, based on platform constraints, the PMIC 1215 also controls the power levels of other components of the system 1200.
[0099] To enable communications to be transmitted and received, various circuits may be coupled between baseband processor 1205 and antenna 1290. Specifically, a radio frequency (RF) transceiver 1270 and a wireless local area network (WLAN) transceiver 1275 may be present. Generally, RF transceiver 1270 may be used to receive and transmit wireless data and calls according to a given wireless communication protocol, such as a 3G or 4G wireless communication protocol, such as Code Division Multiple Access (CDMA), Global System for Mobile Communications (GSM), Long Term Evolution (LTE), or other protocols. Furthermore, a GPS sensor 1280 may be present. Other wireless communications, such as receiving or transmitting radio signals (e.g., AM / FM) and other signals, may also be provided. Furthermore, local wireless communications may also be enabled via WLAN transceiver 1275.
[0100] Now refer to Figure 13 , shown is a block diagram of another example system with which embodiments may be used. Figure 13 In the illustration of , system 1300 can be a mobile low-power system such as a tablet computer, a 2:1 tablet device, a tablet phone or other convertible or standalone tablet system. As shown, there is a SoC 1310 and it can be configured to operate as an application processor for the device.
[0101] Various devices can be coupled to the SoC 1310. In the illustrated diagram, the memory subsystem includes DRAM 1345 and flash memory 1340 coupled to the SoC 1310. Additionally, a touchpad 1320 is coupled to the SoC 1310 to provide display capabilities and user input via touch (including providing a virtual keyboard on the display of the touchpad 1320). To provide wired network connectivity, the SoC 1310 is coupled to an Ethernet interface 1330. A peripheral hub 1325 is coupled to the SoC 1310 to enable interfacing with various peripheral devices, such as may be coupled to the system 1300 through any of a variety of ports or other connectors.
[0102] In addition to the internal power management circuitry and functionality within SoC 1310, PMIC 1380 is coupled to SoC 1310 to provide platform-based power management, for example, based on whether the system is powered by battery 1390 or AC power via AC adapter 1395. In addition to this power source-based power management, PMIC 1380 can further perform platform power management activities based on environmental and usage conditions. Furthermore, PMIC 1380 can communicate control and status information to SoC 1310 to cause various power management actions within SoC 1310.
[0103] Still refer to Figure 13To provide wireless capabilities, a WLAN unit 1350 is coupled to the SoC 1310 and, in turn, to an antenna 1355. In various embodiments, the WLAN unit 1350 may provide communications according to one or more wireless protocols.
[0104] As further illustrated, a plurality of sensors 1360 may be coupled to the SoC 1310. These sensors may include various accelerometer sensors, environmental sensors, and other sensors, including user gesture sensors. Finally, an audio codec 1365 is coupled to the SoC 1310 to provide an interface with an audio output device 1370. Of course, it is understood that while Figure 13 Although shown with this particular implementation, many variations and alternatives are possible.
[0105] Now refer to Figure 14 , showing a notebook computer, an ultrabook TM , or other small form factor systems. In one embodiment, processor 1410 comprises a microprocessor, a multi-core processor, a multi-threaded processor, an ultra-low voltage processor, an embedded processor, or other known processing element. In the illustrated implementation, processor 1410 serves as the main processing unit and central hub for communicating with many of the various components of system 1400 and may include power management circuitry as described herein. As an example, processor 1410 is implemented as a SoC.
[0106] In one embodiment, processor 1410 is in communication with system memory 1415. As an illustrative example, system memory 1415 is implemented via multiple memory devices or modules to provide a fixed amount of system memory.
[0107] To provide persistent storage of information such as data, applications, one or more operating systems, and the like, a mass storage device 1420 may also be coupled to the processor 1410. In various embodiments, to enable thinner and lighter system designs and to improve system responsiveness, this mass storage device may be implemented via an SSD or the mass storage device may be implemented primarily by using a hard disk drive (HDD) with a smaller amount of SSD storage acting as an SSD cache to enable non-volatile storage of context state and other such information during a power loss event so that a fast power-up can occur upon reinitiation of system activity. Also in Figure 14 As shown in FIG, a flash memory device 1422 may be coupled to the processor 1410 (eg, via a serial peripheral interface (SPI)). The flash memory device may provide non-volatile storage for system software (including basic input / output software (BIOS)) and other firmware for the system.
[0108] Various input / output (I / O) devices may be present within the system 1400. Figure 14 Specifically shown in the embodiment of FIG. 1 is a display 1424, which may be a high-definition LCD or LED panel that further provides a touch screen 1425. In one embodiment, the display 1424 may be coupled to the processor 1410 via a display interconnect, which may be implemented as a high-performance graphics interconnect. The touch screen 1425 may be coupled to the processor 1410 via another interconnect, which in one embodiment may be an I 2 C interconnection. As in Figure 14 As further shown in FIG, in addition to the touch screen 1425, user input by touch can also occur via a touch pad 1430, which can be configured within the chassis and can also be coupled to the same I as the touch screen 1425. 2 C interconnection.
[0109] For sensory computing and other purposes, various sensors may be present in the system and may be coupled to the processor 1410 in different ways. Certain inertial sensors and environmental sensors may be coupled to the processor 1410 through the sensor hub 1440 (e.g., via I 2 C interconnect) and coupled to processor 1410. Figure 14 In the embodiment shown in FIG, these sensors may include an accelerometer 1441, an ambient light sensor (ALS) 1442, a compass 1443, and a gyroscope 1444. Other environmental sensors may include one or more thermal sensors 1446, which in some embodiments are coupled to processor 1410 via a system management bus (SMBus) bus.
[0110] Still Figure 14 As seen in FIG14 , various peripheral devices can be coupled to processor 1410 via a low pin count (LPC) interconnect. In the embodiment shown, various components can be coupled through an embedded controller 1435. Such components can include a keyboard 1436 (e.g., coupled via a PS2 interface), a fan 1437, and a thermal sensor 1439. In some embodiments, a touchpad 1430 can also be coupled to EC 1435 via a PS2 interface. In addition, a security processor (such as a Trusted Platform Module (TPM) 1438) can also be coupled to processor 1410 via the LPC interconnect.
[0111] System 1400 can communicate with external devices in various ways, including wirelessly. Figure 14In the embodiment shown in FIG, there are various wireless modules, each of which may correspond to a radio configured for a specific wireless communication protocol. One way to wirelessly communicate over short distances, such as near fields, may be via NFC unit 1445, which, in one embodiment, communicates with processor 1410 via SMBus. It should be noted that via this NFC unit 1445, devices that are very close to each other can communicate.
[0112] As in Figure 14 As further seen in the figure, additional wireless units may include other short-range wireless engines, including WLAN unit 1450 and Bluetooth TM Unit 1452. Using WLAN unit 1450, Wi-Fi can be implemented TM Communication via Bluetooth TM Unit 1452, short-range Bluetooth TM Communication. These units can communicate with the processor 1410 via given links.
[0113] Furthermore, wireless wide area communications (e.g., according to a cellular protocol or other wireless wide area protocols) may occur via a WWAN unit 1456, which in turn may be coupled to a subscriber identity module (SIM) 1457. Additionally, to enable the receipt and use of location information, a GPS module 1455 may also be present. Note that in Figure 14 In the embodiment shown in , WWAN unit 1456 and an integrated capture device (such as camera module 1454) can communicate via a given link.
[0114] To provide audio input and output, the audio processor may be implemented via a digital signal processor (DSP) 1460, which may be coupled to the processor 1410 via a high-definition audio (HDA) link. Similarly, the DSP 1460 may communicate with an integrated coder / decoder (CODEC) and amplifier 1462, which in turn may be coupled to an output speaker 1463, which may be implemented within the chassis. Similarly, the amplifier and CODEC 1462 may be coupled to receive audio input from a microphone 1465, which in one embodiment may be implemented via a dual array microphone (such as a digital microphone array) to provide high quality audio input to enable voice activated control of various operations within the system. Note also that audio output may be provided from the amplifier / CODEC 1462 to a headphone jack 1464. While in Figure 14 While the embodiments are shown using these specific components, it will be understood that the scope of the present invention is not limited in this regard.
[0115] The embodiments may be implemented in many different system types. Figure 15 , which is a block diagram of a system according to an embodiment of the present invention. Figure 15 As shown in FIG, multiprocessor system 1500 is a point-to-point interconnect system and includes a first processor 1570 and a second processor 1580 coupled via a point-to-point interconnect 1550. Figure 15 As shown in FIG, each of processors 1570 and 1580 can be a multi-core processor including first and second processor cores (i.e., processor cores 1574a and 1574b and processor cores 1584a and 1584b), although potentially more cores can be present in the processor. In addition, each processor can include at least one FPGA programmable logic circuit 1577, 1587. Each processor includes a PCU 1575, 1585 or other power management logic to implement processor-based power management, including dynamically identifying a minimum operating voltage to be applied to each FPGA programmable logic function based at least in part on a self-test performed on the critical path length of the function, as described herein.
[0116] Still refer to Figure 15 , the first processor 1570 further includes a memory controller hub (MCH) 1572 and point-to-point (PP) interfaces 1576 and 1578. Similarly, the second processor 1580 includes an MCH 1582 and PP interfaces 1586 and 1588. Figure 15 As shown in FIG, MCHs 1572 and 1582 couple the processors to respective memories (i.e., memory 1532 and memory 1534), which may be portions of system memory (e.g., DRAM) locally attached to the respective processors. First processor 1570 and second processor 1580 may be coupled to chipset 1590 via PP interconnects 1562 and 1564, respectively. As shown in FIG Figure 15 As shown in , chipset 1590 includes PP interfaces 1594 and 1598 .
[0117] In addition, the chipset 1590 includes an interface 1592 to couple the chipset 1590 with the high performance graphics engine 1538 via the PP interconnect 1539. In turn, the chipset 1590 can be coupled to the first bus 1516 via an interface 1596. Figure 15As shown in FIG1 , various input / output (I / O) devices 1514 may be coupled to first bus 1516 along with a bus bridge 1518 that couples first bus 1516 to a second bus 1520. In one embodiment, various devices may be coupled to second bus 1520 including, for example, a keyboard / mouse 1522, communication devices 1526, and a data storage unit 1528 (such as a disk drive or other mass storage device which may include code 1530). Further, an audio I / O 1524 may be coupled to second bus 1520. Embodiments may be incorporated into computers including, for example, smart cell phones, tablet computers, netbook computers, ultrabook computers, and the like. TM and so on in other types of systems of mobile devices.
[0118] One or more aspects of at least one embodiment may be implemented by representative code stored on a machine-readable medium that represents and / or defines logic within an integrated circuit, such as a processor. For example, a machine-readable medium may include instructions that represent various logic within a processor. When read by a machine, the instructions may cause the machine to fabricate logic to implement the techniques described herein. Such representations (referred to as "IP cores") are reusable logic units for integrated circuits that may be stored on a tangible, machine-readable medium as a hardware model that describes the structure of the integrated circuit. The hardware model may be provided to various customers or manufacturing facilities that load the hardware model on a manufacturing machine that manufactures the integrated circuit. The integrated circuit may be manufactured so that the circuitry implements the operations described in association with any of the embodiments described herein.
[0119] Figure 16 The block diagram illustrates an IP core development system 1600 that can be used to manufacture integrated circuits for implementation, according to an embodiment. The IP core development system 1600 can be used to generate modular, reusable designs that can be incorporated into larger designs or used to build entire integrated circuits (e.g., SoC integrated circuits). A design facility 1630 can generate a software simulation 1610 of the IP core design in a high-level programming language (e.g., C / C++). The software simulation 1610 can be used to design, test, and verify the behavior of the IP core. A register transfer level (RTL) design can then be created or synthesized based on the simulation model. The RTL design 1615 is an abstraction of the integrated circuit's behavior that models the flow of digital signals between hardware registers, including the associated logic implemented using the modeled digital signals. In addition to the RTL design 1615, lower-level designs below the logic or transistor level can also be created, designed, or synthesized. Therefore, the specific details of the initial design and simulation may vary.
[0120] The RTL design 1615 or equivalent can be further synthesized into a hardware model 1620 by a design facility, which can be in a hardware description language (HDL), or some other representation of physical design data. The HDL can be further simulated or tested to verify the IP core design. The IP core design can be stored for delivery to a third-party manufacturing facility using non-volatile memory 1640 (e.g., a hard disk, flash memory, or any non-volatile storage medium). Alternatively, the IP core design can be transmitted via a wired connection 1650 or a wireless connection 1660 (e.g., via the Internet). The manufacturing facility 1665 can then manufacture an integrated circuit based at least in part on the IP core design. The manufactured integrated circuit can be configured to perform operations according to at least one embodiment described herein.
[0121] Now refer to Figure 17 , shows a block diagram of a system according to an embodiment. System 1700 can be any type of computing system, ranging from a small device to a large server computer. In the particular embodiment shown, system 1700 provides an environment for generating bitstreams that provide one or more programmable logic functions to be implemented within a field programmable gate array (FPGA). To this end, system 1700 includes an EDA synthesis tool 1720. Using the embodiments described herein, synthesis tool 1720 can generate one or more bitstreams that provide the programmable functions, along with metadata that is used to dynamically control the operating voltage of a particular FPGA based on a priori knowledge of the programmable logic function(s) to be executed within the FPGA. As shown, synthesis tool 1720 provides the bitstreams, along with the associated metadata, to non-volatile memory 1730. In an embodiment, non-volatile memory 1730 can be implemented as flash memory.
[0122] During normal system operation (in the absence of synthesis tool 1720, such as when implementing FPGA 1710 in the field), and more specifically during boot activities, one or more bitstreams and metadata may be downloaded to FPGA 1710. The downloaded bitstreams may be used to program one or more programmable logic circuits 17120-1712 N Programming. Furthermore, the metadata may include information for performing a dynamic characterization self-test within FPGA 1710 to identify the appropriate operating voltage at which FPGA 1710 can operate. Such a self-test can be used to identify and dynamically test one or more critical path functions within programmable logic circuit 1712 to identify the appropriate operating voltage. After completing such a dynamic diagnostic self-test, operating parameters, including the minimum operating voltage, can be identified and stored locally for use during normal operation.
[0123] Still refer to Figure 17, the FPGA 1710 includes a built-in self-test (BIST) circuit 1714 (also referred to herein as a "self-test circuit"). In embodiments herein, the BIST circuit 1714 may perform a diagnostic self-test as described herein based on information within the metadata to identify one or more appropriate operating voltages. In turn, the operating voltage(s) so identified may be provided to a power controller 1716, which in turn may control a voltage regulator 1718 to provide such operating voltage(s) to various circuits of the FPGA 1710, including one or more programmable logic circuits 1712. In embodiments, the voltage regulator 1718 may be an internal or external voltage regulator. In embodiments, the power controller 1716 may control the voltage regulator 1718 based on a determined minimum operating voltage obtained via a Vmin search performed by the BIST circuit 1714, and maintain that same voltage level until the next boot sequence. Although Figure 17 The embodiment shown in FIG. 1 shows a hardware embodiment of the self-test circuit 1714, but it is understood that other implementations are possible. In an embodiment, the self-test circuit 1714 can be executed during the boot routine to perform a VMIN search using a predefined model for a specific FPGA configuration to identify the optimal minimum voltage. This offline generated metadata can enable faster and more accurate modeling of specific scenarios.
[0124] Although Figure 17 The embodiments of FIG. 1 show a standalone FPGA, but it is understood that the embodiments are not limited in this respect. For example, in other embodiments, a single integrated circuit may include a combination of an FPGA and a SoC on the same die or on multiple dies in the same package. Although FIG. Figure 17 The embodiments are shown at this high level, but many variations and alternatives are possible.
[0125] Now refer to Figure 18 , shows a flow chart of a method according to an embodiment of the present invention. Figure 18 As shown in , method 1800 can be performed by circuitry within an FPGA to identify one or more minimum operating voltages at which the programmable logic circuitry of the FPGA is to operate during normal operation. Method 1800 can be performed when a system including at least one FPGA is restarted. At block 1810, a bitstream is loaded from a storage device into the FPGA. More specifically, the bitstream can be loaded from a source device, such as a given non-volatile memory system, into the FPGA. Thereafter, at block 1820, a boot routine is entered to perform startup activities within the FPGA.
[0126] Still refer to Figure 18, control next passes to block 1830, where metadata can be loaded from a storage device. More specifically, the metadata can be loaded from the storage device to the BIST circuitry of the FPGA, and the metadata can include one or more diagnostic self-tests and an estimated minimum operating voltage. Next, control passes to block 1840, where the BIST circuitry can communicate the estimated minimum operating voltage to a power controller, for example, so that the power controller can set the operating voltage to the estimated minimum operating voltage value, for example, via communication with one or more voltage regulators.
[0127] At this point, the FPGA is fully configured to perform diagnostic self-tests. Control therefore passes to block 1850, where diagnostic tests are performed based on the estimated minimum voltage and metadata, which may include one or more sets of diagnostic self-tests to fully exercise one or more critical logic paths of the programmable logic circuit. Note that in some cases, in addition to using metadata to control voltage regulators, the FPGA's power controller can also read metadata including a given operating frequency and voltage and can trigger an exercise of logic within the FPGA to perform BIST functionality. Note also that, in different embodiments, the critical path length of a function can be a specific circuit or a given functional circuit of the FPGA to be accessed during self-test. For example, if the FPGA critical circuit is a multiplier, DFT logic can be used to activate the multiplier at BIST to test it in situ, and then use the same multiplier at runtime.
[0128] Next, at the conclusion of the diagnostic self-test, a determination is made at diamond 1860 as to whether the self-test met a performance threshold. For example, a determination may be made as to whether the test results calculated during the self-test met expected test results. Alternatively, a determination may be made as to whether the self-test met appropriate operating criteria, such as instruction execution rate. If it is determined that the test met one or more such performance thresholds, control passes to diamond 1870, where the voltage level is stored as a minimum operating voltage value, for example, in a given configuration storage device. Furthermore, a control signal may be sent to one or more voltage regulators to cause the one or more voltage regulators to operate at the operating voltage.
[0129] Note that if the diagnostic test does not meet the performance threshold, then control is transferred from diamond 1860 to block 1880 instead. There, the given minimum voltage at which the diagnostic self-test is performed can be updated, for example, by a certain step value. In an embodiment, the step value can be a predetermined step value, for example, between about 5 and 50 millivolts. Thereafter, control is transferred back to block 1850 for further diagnostic self-tests. Although in Figure 18While the embodiment of the present invention is shown at this high level, many variations and alternatives are possible. For example, where there are multiple programmable logic circuits, each of which is to perform at least one logic function, the flow of method 1800 can be iteratively performed to identify a separate minimum operating voltage for each such circuit. In this way, a single FPGA can be operated such that different programmable logic circuits can execute simultaneously at independent voltages.
[0130] With this minimum voltage, a lower voltage guard band can be applied based on die manufacturing statistical variation parameters and further based on real-time field testing. In this way, embodiments provide a single-part solution that can deliver reduced power consumption across a wide range of potential frequencies, providing users with significant flexibility for a wide range of design configurations. In contrast, conventional FPGAs use only a single predefined voltage-to-frequency modeling that adapts to a specific design architecture and worst-case critical path. In such conventional FPGAs, while the voltage can be varied for lower frequencies as characterized by the predefined curve, there is no ability to adjust the voltage for a specific configuration of the FPGA.
[0131] By using embodiments, multiple predefined models of voltage to frequency are supported. Additionally, changes from any of these models can be accommodated based on actual device testing using self-test circuits as described herein.
[0132] Now refer to Figure 19 , shows a graphical illustration of a voltage-frequency curve according to an embodiment. Figure 19 As shown in FIG. 1 , example 1900 provides a plurality of voltage-frequency curves 19100-1910 n In embodiments as described herein, a single FPGA can be programmably controlled to operate at operating points associated with one or more of the voltage-frequency curves 1910, for example, based at least in part on one or more functions to be executed on one or more programmable logic circuits of the FPGA. Figure 19 A given FPGA can dynamically operate according to different ones of these models depending on the specific design logic configuration.
[0133] Now refer to Figure 20 , shows a block diagram of a system according to another embodiment of the present invention. More specifically, Figure 20A computing system 2000 is shown that includes an FPGA 2010 that is programmably controlled to perform one or more logic functions. In addition, as described herein, the FPGA 2010 can be controlled to operate at a separate, dynamically determined minimum operating voltage (or multiple independent minimum operating voltages) sufficient to ensure that the one or more logic functions are properly executed. To this end, the FPGA 2010 is coupled to a storage device 2080, which can be any type of non-volatile storage device. Although various information is stored in such a storage device, it is of interest here that the storage device 2080 includes a configuration table 2085 that is used to store various configuration information that can be downloaded to the FPGA 2010 during power-up and normal operation within the system 2000. In order to enable the FPGA 2010 to perform one or more programmable logic functions, a plurality of embedded bitstreams 20860-2086 N It can be stored in the configuration table 2085. Although various information exists within such a bitstream, what is of interest here is that each bitstream 2086 can include one or more logic functions and metadata. Such metadata can include one or more voltage / frequency tables that provide estimated values of voltage / frequency for the corresponding logic function. The FPGA 2010 can use such estimated voltage information within these V / F tables to enable dynamic determination of the appropriate minimum operating voltage at which the corresponding logic function can operate during normal operation. To this end, the bitstream 2086 can further include one or more self-tests to enable determination of the minimum operating voltage. It is understood that although the present embodiment contemplates that the bitstream itself includes metadata, in other cases, although metadata is associated with a given bitstream, the metadata can also be stored, accessed, and downloaded separately.
[0134] As can be seen, FPGA 2010 also includes a loader 2015. Loader 2015 can be a boot loader or other controller that receives an incoming bitstream and corresponding metadata. Loader 2015 can use the incoming bitstream to program logic circuits 20120-2012 N2012. The metadata within the associated bitstream can then be used to enable BIST circuitry 2020 to perform a self-test of programmable logic circuitry 2012. Loader 2015 can further receive the configuration data and transmit it to voltage and frequency control circuitry 2060. Control circuitry 2060 can issue control signals to power control circuitry 2030, which in turn initiates control signals transmitted via voltage / frequency table 2040 to generate one or more operating frequencies within clock generator 2050. Additionally, similar control signals can be provided to voltage regulator 2070 (illustrated as an external voltage regulator in this embodiment) to generate and provide one or more operating voltages to FPGA 2010.
[0135] In various embodiments, an EDA synthesis tool can be configured to generate metadata associated with a given bitstream. This metadata can include various information, including the critical path length to be exercised by the self-test circuit to determine specific design constraints. This metadata can be stored in a non-volatile storage device in association with the corresponding bitstream so that it can be downloaded to the FPGA at boot time. At boot time, the metadata stored in the storage device is transferred to the FPGA, and more specifically, to the self-test unit. In turn, the bitstream is sent to the FPGA's fabric to configure the FPGA logic elements to the specific design configuration.
[0136] Embodiments may therefore enable power consumption reduction for FPGA standalone devices and / or systems on a chip or other subsystems including a central processing unit (CPU) (e.g., including one or more general-purpose processing cores) and one or more accelerators implemented as FPGAs.
[0137] The following examples relate to further embodiments.
[0138] In one example, an apparatus includes an FPGA, comprising: at least one programmable logic circuit to perform a function programmed with a bitstream; a self-test circuit to perform a self-test at a first voltage, the self-test and the first voltage being programmed with first metadata associated with the bitstream, the self-test including at least one critical path length of the function; and a power controller to identify an operating voltage for the at least one programmable logic circuit based at least in part on the performance of the self-test at the first voltage.
[0139] In an example, the first voltage comprises an estimated minimum voltage at which the function is to be performed correctly.
[0140] In an example, the estimated minimum voltage is determined by a synthesis tool based on at least one critical path length of the function.
[0141] In an example, the power controller is to send a voltage command to a voltage regulator to cause the voltage regulator to provide the operating voltage to the at least one programmable logic circuit.
[0142] In an example, the power controller is to identify a first voltage as the operating voltage in response to execution of the self-test satisfying a performance threshold.
[0143] In an example, in response to the execution of the self-test at the first voltage not satisfying the performance threshold, the power controller is to increase the first voltage to a second voltage and cause the self-test circuit to perform the self-test at the second voltage.
[0144] In an example, the apparatus further includes a non-volatile memory coupled to the FPGA, the non-volatile memory to store the bitstream and the first metadata, wherein, when the FPGA is powered on, the non-volatile memory is to provide the bitstream and the first metadata to the FPGA so that the at least one programmable logic circuit is programmed for the function.
[0145] In an example, the operating voltage includes a minimum voltage sufficient to support the at least one critical path length at an operating frequency for the function and is further based on process variations of the FPGA.
[0146] In an example, the minimum voltage is lower than a worst-case minimum voltage for the FPGA based on a fastest supported frequency and a slowest process variation.
[0147] In an example, the FPGA includes a plurality of programmable logic circuits, each programmable logic circuit to perform one of a plurality of functions programmed with one of a plurality of bitstreams.
[0148] In an example, at least some of the plurality of programmable logic circuits are to operate at independent voltages based on self-tests performed by the self-test circuitry for corresponding programmable logic circuits and functions.
[0149] In an example, the apparatus comprises a system on a chip including the FPGA and one or more processor cores of a central processing unit.
[0150] In an example, the apparatus further includes a semiconductor die including the FPGA and the central processing unit.
[0151] In another example, a method includes receiving a bitstream and metadata associated with the bitstream in an FPGA having at least one programmable logic circuit; programming the at least one programmable logic circuit with the bitstream to enable the at least one programmable logic circuit to perform a function; performing at least one self-test for the function in the FPGA at a first voltage, the metadata identifying the at least one self-test and the first voltage; determining whether the execution of the at least one self-test satisfies a performance threshold; and operating the FPGA at the first voltage in response to determining that the execution of the at least one self-test satisfies the performance threshold.
[0152] In an example, the method further includes, in response to determining that execution of the at least one self-test does not satisfy the performance threshold, executing the at least one self-test for the function in the FPGA at a second voltage, the second voltage being greater than the first voltage.
[0153] In an example, the first voltage comprises an estimated minimum operating voltage to perform the function, and the second voltage comprises a sum of the estimated minimum operating voltage and a step value.
[0154] In an example, the method further includes performing the at least one self-test via a self-test circuit of the FPGA, the self-test circuit being programmed with at least a portion of the metadata.
[0155] In another example, a computer-readable medium includes instructions to perform the method of any of the above examples.
[0156] In another example, a computer-readable medium includes data to be used by at least one machine to fabricate at least one integrated circuit to perform the method of any of the above examples.
[0157] In another example, an apparatus includes means for performing the method of any of the above examples.
[0158] In another example, an apparatus includes a system on a chip including at least one core and at least one FPGA. The at least one FPGA may include: a plurality of programmable logic circuits, each programmable logic circuit to perform a function programmed with a corresponding bitstream; at least one self-test circuit to perform one or more self-tests at one or more first voltages, the at least one self-test circuit to be programmed with metadata associated with the corresponding bitstream; and a power controller to identify a minimum operating voltage for the plurality of programmable logic circuits based on the execution of the one or more self-tests.
[0159] In an example, the apparatus includes a semiconductor die including the at least one core and the at least one FPGA.
[0160] In an example, the apparatus further comprises a non-volatile storage device coupled to the at least one FPGA, the non-volatile storage device to store the corresponding bitstream and the metadata, and wherein the first voltage comprises an estimated minimum voltage at which the function is to be performed correctly.
[0161] In another example, an apparatus for controlling the voltage of an FPGA includes: a component for receiving a bitstream and metadata associated with the bitstream; a component for programming at least one programmable logic component of the FPGA with the bitstream to enable the at least one programmable logic component to perform a function; a component for performing at least one self-test for the function in the FPGA at a first voltage, the metadata identifying the at least one self-test and the first voltage; a component for determining whether the execution of the at least one self-test satisfies a performance threshold; and a component for causing the at least one programmable logic component of the FPGA to operate at the first voltage in response to determining that the execution of the at least one self-test satisfies the performance threshold.
[0162] In an example, the apparatus further includes means for performing the at least one self-test for the function in the FPGA at a second voltage greater than the first voltage in response to determining that execution of the at least one self-test does not satisfy the performance threshold.
[0163] In an example, the first voltage comprises an estimated minimum operating voltage to perform the function, and the second voltage comprises a sum of the estimated minimum operating voltage and a step value.
[0164] In an example, the apparatus further includes means for programming a self-test component of the FPGA with at least a portion of the metadata.
[0165] It is understood that various combinations of the above examples are possible.
[0166] Embodiments may be used in many different types of systems. For example, in one embodiment, a communication device may be configured to implement the various methods and techniques described herein. Of course, the scope of the present invention is not limited to communication devices, but other embodiments may involve other types of apparatuses for processing instructions, or one or more machine-readable media containing instructions, which, in response to being executed on a computing device, cause the device to implement one or more of the methods and techniques described herein.
[0167] Embodiments may be implemented in code and stored on a non-transitory storage medium having stored thereon instructions that can be used to program a system to perform the instructions. Embodiments may also be implemented in data and stored on a non-transitory storage medium that, if used by at least one machine, causes the at least one machine to fabricate at least one integrated circuit to perform one or more operations. The storage medium may include, but is not limited to, any type of magnetic disk (including floppy disks, optical disks, solid-state drives (SSDs), compact disk read-only memories (CD-ROMs), compact disk rewritable memories (CD-RWs), and magneto-optical disks), semiconductor devices (such as read-only memories (ROMs), random access memories (RAMs) (such as dynamic random access memories (DRAMs), static random access memories (SRAMs), erasable programmable read-only memories (EPROMs)), flash memories, electrically erasable programmable read-only memories (EEPROMs)), magnetic or optical cards, or any other type of medium suitable for storing electronic instructions.
[0168] While the present invention has been described with respect to a limited number of embodiments, those skilled in the art will appreciate numerous modifications and variations therefrom. It is intended that the appended claims cover all such modifications and variations as fall within the true spirit and scope of the present invention.
Claims
1. A device for identifying an operating voltage, comprising: Field-programmable gate arrays (FPGAs), including: at least one programmable logic circuit to perform the functions programmed by the bitstream; self-test circuitry to perform a self-test at a first voltage, the self-test and the first voltage being programmed with first metadata associated with the bitstream, the self-test including at least one critical path length of the function; and A power controller is to identify an operating voltage for the at least one programmable logic circuit based at least in part on execution of a self-test at a first voltage.
2. The device according to claim 1, wherein The first voltage comprises an estimated minimum voltage at which the function is to be performed correctly.
3. The device according to claim 2, wherein The estimated minimum voltage is determined by a synthesis tool based on at least one critical path length of the function.
4. The device according to claim 1, wherein The power controller is to send a voltage command to a voltage regulator so that the voltage regulator provides the operating voltage to the at least one programmable logic circuit.
5. The device according to claim 1, wherein The power controller is to identify a first voltage as the operating voltage in response to execution of the self-test satisfying a performance threshold.
6. The device according to claim 5, wherein In response to the execution of the self-test at the first voltage not satisfying the performance threshold, the power controller is to increase the first voltage to a second voltage and cause the self-test circuit to perform the self-test at the second voltage.
7. The device according to claim 1, wherein The apparatus further includes a non-volatile memory coupled to the FPGA, the non-volatile memory to store the bitstream and the first metadata, wherein, when the FPGA is powered on, the non-volatile memory is to provide the bitstream and the first metadata to the FPGA so that the at least one programmable logic circuit is programmed for the function.
8. The device according to claim 1, wherein The operating voltage includes a minimum voltage sufficient to support the at least one critical path length at an operating frequency for the function and is further based on process variations of the FPGA.
9. The device according to claim 8, wherein The minimum voltage is lower than a worst-case minimum voltage for the FPGA based on a fastest supported frequency and a slowest process variation.
10. The device according to claim 1, wherein The FPGA includes a plurality of programmable logic circuits, each programmable logic circuit to perform one of a plurality of functions programmed with one of a plurality of bitstreams.
11. The device according to claim 10, wherein At least some of the plurality of programmable logic circuits are to operate at independent voltages based on self-tests performed by the self-test circuitry for corresponding programmable logic circuits and functions.
12. The device according to claim 1, wherein The apparatus includes a system on a chip, which includes the FPGA and one or more processor cores of a central processing unit.
13. The apparatus of claim 12, further comprising a semiconductor die including the FPGA and the central processing unit.
14. A method for controlling voltage, comprising: receiving a bitstream and metadata associated with the bitstream in a field programmable gate array (FPGA) having at least one programmable logic circuit; programming the at least one programmable logic circuit with the bitstream to enable the at least one programmable logic circuit to perform a function; performing at least one self-test for the functionality in the FPGA at a first voltage, the metadata identifying the at least one self-test and the first voltage; determining whether execution of the at least one self-test satisfies a performance threshold; as well as In response to determining that execution of the at least one self-test satisfies the performance threshold, the FPGA is operated at a first voltage. 15 . The method of claim 14 , further comprising, in response to determining that execution of the at least one self-test does not satisfy the performance threshold, executing the at least one self-test for the function in the FPGA at a second voltage, the second voltage being greater than the first voltage.
16. The method according to claim 15, wherein The first voltage includes an estimated minimum operating voltage to perform the function, and the second voltage includes a sum of the estimated minimum operating voltage and a step value.
17. The method of claim 14, further comprising performing the at least one self-test via self-test circuitry of the FPGA, the self-test circuitry being programmed with at least a portion of the metadata.
18. A computer-readable storage medium comprising computer-readable instructions, which when executed are to implement the method according to any one of claims 14 to 17.
19. A device for identifying a minimum operating voltage, comprising: A system on a chip comprising at least one core and at least one field programmable gate array (FPGA), the at least one FPGA comprising: a plurality of programmable logic circuits, each programmable logic circuit to perform a function programmed by a corresponding bitstream; at least one self-test circuit to perform one or more self-tests at one or more first voltages, the at least one self-test circuit to be programmed with metadata associated with a corresponding bitstream; and A power controller is to identify a minimum operating voltage for the plurality of programmable logic circuits based on execution of the one or more self-tests.
20. The apparatus of claim 19, further comprising a semiconductor die comprising the at least one core and the at least one FPGA.
21. The apparatus of claim 19, further comprising a non-volatile storage device coupled to the at least one FPGA, the non-volatile storage device to store the corresponding bitstream and the metadata, and wherein The first voltage comprises an estimated minimum voltage at which the function is to be performed correctly.
22. A device for controlling the voltage of a field programmable gate array (FPGA), comprising: means for receiving a bitstream and metadata associated with the bitstream; means for programming at least one programmable logic component of the FPGA with the bitstream to enable the at least one programmable logic component to perform a function; means for performing at least one self-test for the functionality in the FPGA at a first voltage, the metadata identifying the at least one self-test and the first voltage; means for determining whether execution of the at least one self-test satisfies a performance threshold; as well as Means for causing at least one programmable logic component of the FPGA to operate at a first voltage in response to determining that execution of the at least one self-test satisfies the performance threshold.
23. The apparatus of claim 22, further comprising means for performing the at least one self-test for the function in the FPGA at a second voltage greater than the first voltage in response to determining that execution of the at least one self-test does not satisfy the performance threshold.
24. The device according to claim 23, wherein The first voltage includes an estimated minimum operating voltage to perform the function, and the second voltage includes a sum of the estimated minimum operating voltage and a step value.
25. The apparatus of claim 22, further comprising means for programming a self-test component of the FPGA with at least a portion of the metadata.
Citation Information
Patent Citations
Self-test circuitry to determine minimum operating voltage
CN101176009A
Device specific configuration of operating voltage
CN103003816A