System, device, and method for power estimation

The hardware-based power estimator addresses inefficiencies in power consumption estimation by using scaling and MADD circuits to quickly and accurately calculate power in complex electronic systems, enhancing dynamic power management and optimization.

JP2025112284APending Publication Date: 2025-07-31INFINEON TECHNOLOGIES AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025006804
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-18
Filing Date
2025-01-17
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Current methods for estimating power consumption in complex electronic systems, such as microcontrollers, are inefficient and impractical, requiring gate-level simulations and pre-silicon measurements, which are time-consuming and complex, especially for systems with numerous configurations and dynamic power variations.

Method used

A hardware-based power estimator (HPE) using scaling circuits and multiply-accumulate (MADD) circuits to estimate dynamic power consumption in real-time by applying activation functions to input signals representing the status and activity of IP blocks, reducing computational complexity through recursive calculations.

Benefits of technology

Enables fast and accurate power consumption estimation within a few clock cycles, facilitating dynamic power management and optimization across various configurations, supporting both pre-silicon and post-silicon phases with low latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025112284000001_ABST
    Figure 2025112284000001_ABST
Patent Text Reader

Abstract

To provide a system, a device, and a method for power estimation.SOLUTION: A hardware power estimator (HPE) configured to provide power estimation for an electronic device includes a scaling circuit configured to receive an input and generate a scaled, nonlinear output of the input, the scaling circuit applying an approximation of an activation function to the input. The HPE further includes a multiply-add (MADD) circuit device configured to use the scaled, nonlinear output to generate an output indicative of mathematical power estimation for the electronic device. The input to the scaling circuit includes data values representing state information or activity information of one or more components or intellectual property (IP) blocks of the electronic device.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various embodiments generally relate to power estimation of electronic devices.

Background Art

[0002] Managing power consumption in embedded devices is important for extending battery life and minimizing the environmental footprint of the system. As the importance of this aspect in system design increases, estimating power usage at the system level is becoming increasingly challenging. This complexity applies to a wide range of scenarios, from industrial setups with sensors to IoT applications with numerous embedded devices, to automotive applications such as engine control modules and power transmission system modules. For example, automotive microcontrollers incorporate accelerators that exhibit large variations in dynamic power consumption based on their configuration and data pipeline.

[0003] There has been a paradigm shift towards power-aware methods for integrating system-on-chip (SoC). Power-aware designs and techniques with inputs based on accurate pre-silicon power consumption are widely used in modern SOCs. These techniques are also utilized at the silicon / hardware level, where it is necessary to manage system power more efficiently. By estimating dynamic power consumption at runtime, faster and better power management schemes (e.g., DVC, DVFS, etc.) can be utilized in a given system. Furthermore, estimating power for the entire system (i.e., the sum of all individual SoC / ASICs) while real-time OS / software is operating on complex hardware leads to global optimization of power / energy consumption.

[0004] Currently, estimating the average power consumption of ASICs and microcontrollers for specific use cases requires gate-level, RTL simulation, or timing-based activity simulation. The simulation time constraints and the complexity of converting user code to vectors prevent an immediate estimation of the average power consumption for iterative power performance optimization of complex applications / use cases. Further, due to implementation constraints, it is impractical to implement a hardware-based power estimator using the same technology. Therefore, modern SoCs reach the overall system power state (deep sleep, sleep, idle, standby, etc.) using power management IP that aggregates various IP logic states. These states are then used by a power management controller to optimize the overall power.

[0005] Generally, a microcontroller includes several processing units, multiple examples of different types of memory cells, and peripheral devices. Consequently, each of the IPs supports different configurations. For example, a CPU core can operate at different frequencies, with or without lockstep between cores, and with different activity factors (IPC rate). An IO interface such as CAN supports different baud rates, various power management states, etc. Therefore, theoretically, there are millions of combinations of unique configurations for this type of microcontroller. Furthermore, typical applications dynamically switch the microcontroller between some of these configurations. For example, an accelerator used in an ADAS or visual computing system supports several configurations with variable data and bus load sizes. This further increases the complexity of estimating or modeling the dynamic power behavior of this type of system. In a specific example such as a microcontroller, the digital logic power consumption depends on its configuration compared to a general-purpose microprocessor where the size and complexity of the software can affect the power consumption. Currently, estimating the average dynamic current consumption of this type of microcontroller user software requires current measurements in silicon or gate-level silicon simulation (in the pre-silicon phase).

[0006] In the drawings, like reference numerals generally refer to the same parts throughout different drawings. The drawings are not necessarily to scale; rather, they are emphasized generally to illustrate the principles of the present invention. In the following description, various embodiments of the present invention are described with reference to the accompanying drawings.

Brief Description of the Drawings

[0007]

Figure 1A

Figure 1B

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

[0008] The following detailed description refers to the accompanying drawings which illustrate specific details and embodiments by way of example in which the invention may be practiced.

[0009] As used in this specification, the term "exemplary" is used to mean "serving as an example, instance, or illustration". Any embodiment or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or designs.

[0010] The terms "plurality" and "multiple" in the specification or claims explicitly mean a quantity greater than one. Terms such as "(the) group", "(the) set", "(the) collection", "(the) series", "(the) sequence", "(the) grouping", etc. in the specification or claims mean a quantity of one or more, i.e., one or more. Any term expressed in the plural form without explicitly stating "plurality" or "multiple" similarly means a quantity of one or more. The terms "suitable subset", "reduced subset", and "smaller subset" mean a subset of a set that is not equal to the set, i.e., a subset of the set that contains fewer elements than the set.

[0011] The terms "at least one" and "one or more" may be understood to include a quantity of one or more (e.g., 1, 2, 3, 4,... etc.).

[0012] When used in this specification, unless otherwise specified, the use of ordinal adjectives such as "first", "second", "third", etc. for describing common objects is merely to indicate that different examples of similar objects are being referred to, and it is not intended to imply that the objects so described must be in a predetermined order, whether temporally, spatially, in ranking order, or in any other way.

[0013] As used herein, the term "data" includes any suitable analog or digital form of information, and may be understood, for example, as provided as a file, a portion of a file, a set of files, a signal or stream, a portion of a signal or stream, a set of signals or streams, etc. Further, the term "data" may be used to mean a reference to information (e.g., in the form of a pointer). However, the term "data" is not limited to the examples described above, may take various forms, and may represent any information as understood in the prior art.

[0014] As used herein, the term "processor" or "controller" may be understood as any type of entity capable of handling data, signals, etc. The data, signals, etc. may be handled in accordance with one or more specific functions executed by the processor or controller.

[0015] Thus, the processor or controller may be an analog circuit, a digital circuit, a mixed-signal circuit, a logic circuit, a processor, a microprocessor, a central processing unit (CPU), a neuromorphic computer unit (NCU), a graphics processing unit (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA), an integrated circuit, an application specific integrated circuit (ASIC), etc. or any combination thereof, or may include these. Any other type of implementation of each function described in detail below may also be understood as a processor, controller or logic circuit. It should be understood that any two (or more) of the processors, controllers or logic circuits detailed herein may be implemented as a single entity having equivalent functionality, etc., and conversely, any single processor, controller or logic circuit detailed herein may be implemented as two (or more) separate entities having equivalent functionality, etc.

[0016] As used in this specification, "circuit" is understood as any kind of logic execution entity and may include a processor that executes special-purpose hardware or software. Thus, a circuit may be an analog circuit, a digital circuit, a mixed-signal circuit, a logic circuit, a processor, a microprocessor, a signal processor, a central processing unit ("CPU"), a graphics processing unit ("GPU"), a neuromorphic computer unit (NCU), a digital signal processor ("DSP"), a field-programmable gate array ("FPGA"), an integrated circuit, an application-specific integrated circuit ("ASIC"), etc. or any combination thereof. Any other kind of implementation of each function detailed below may also be understood as a "circuit". It should be understood that any two (or more) of the circuits detailed in this specification may be implemented as a single circuit having substantially equivalent functionality. Conversely, it should be understood that any single circuit detailed in this specification may be implemented as two (or more) separate circuits having substantially equivalent functionality. In addition, a reference to a "circuit" may also mean two or more circuits that collectively form a single circuit.

[0017] As used herein, terms such as "module," "component," "system," "circuit," "element," "interface," "slice," "circuitry," etc. are intended to mean a set of one or more electronic components, computer-related entities, hardware, software (e.g., as in an implementation note), and / or firmware. For example, a circuit or similar term can be a computer having a processor, a process running on the processor, a controller, an object, an executable program, a storage device, and / or a processing device. By way of example, an application running on a server and the server can also be a circuit. One or more circuits can be present within the same circuit, a circuit can be located on one computer, and / or can be distributed between two or more computers. A set of elements or other set of circuits can be described herein, and the term "set" can be interpreted as "one or more."

[0018] As used herein, a "signal" may be transmitted or conducted through a signal chain, where the signal is processed to change characteristics such as phase, amplitude, frequency, etc. Even if this type of characteristic is adapted, the signal may still be referred to as the same signal. Generally, a signal may be considered the same signal as long as it continues to encode the same information.

[0019] As used herein, a signal that "represents" a value or other information may be a digital or analog signal that is decodable and / or encodes or conveys the value or other information such that it can cause a response operation of a component that receives the signal. The signal may be stored or buffered in a computer-readable storage medium prior to its receipt by the receiving component. The receiving component may read the signal from the storage medium. Further, a "value" that "represents" some quantity, state, or parameter may be physically implemented as a digital signal, an analog signal, or stored bits that encode or convey the value.

[0020] When an element is said to be "connected" or "coupled" to another element, it should be understood that the element can be physically connected or coupled to the other element and that current and / or electromagnetic radiation (e.g., a signal) can flow along a conductive path formed by the element. When it is described that an element is coupled or connected to another element, there may be intervening conductive, inductive, or capacitive elements between the element and the other element. Further, when one element is coupled or connected to another element, without physical contact or intervening components, one element may be able to induce a voltage or current flow or the propagation of an electromagnetic wave in the other element. Further, when a voltage, current, or signal is said to be "applied" to an element, the voltage, current, or signal may be connected to the element by a physical connection or by a capacitive, electromagnetic, or inductive coupling that does not include a physical connection.

[0021] As used herein, "memory" is understood as a non-transitory computer-readable medium capable of storing data or information for reading. Thus, references to "memory" included in this specification may be understood to mean volatile or non-volatile memory including random access memory (RAM), read-only memory (ROM), flash memory, solid state storage devices, magnetic tape, hard disk drives, optical drives, etc. or any combination thereof. Further, registers, shift registers, processor registers, data buffers, etc. are also embraced by the term "memory" in this specification. A single component referred to as "memory" or "a memory" may be composed of multiple different types of memory, and thus may mean a component of an aggregate comprising one or more types of memory. Any single memory component may be separated into a plurality of collectively equivalent memory components, and vice versa. Further, memory may be depicted as separate from one or more other components (as in the drawings), but memory may also be integrated into other components, e.g., on a common integrated chip or controller, as embedded memory.

[0022] The term "software" means any type of executable instructions including firmware.

[0023] Unless otherwise specified, in this specification, discussions using terms such as "process", "compute", "calculate", "determine", "present", "display", etc. may refer to actions or processes of a machine (e.g., a computer / processor / etc.) that manipulates or transforms data represented as a physical (e.g., electronic, magnetic, or optical) quantity within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.

[0024] Exemplary embodiments of the present disclosure may be implemented by one or more computers (or computing devices) that read and execute computer-executable instructions recorded on a storage medium (e.g., a non-transitory computer-readable storage medium) to perform one or more functions of the present disclosure described herein. The computer may include one or more of a central processing unit (CPU), a microprocessing unit (MPU), or other circuitry, and may include a network of separate computers or separate computer processors. The computer-executable instructions may be provided to the computer, for example, from a network or a non-volatile computer-readable storage medium. The storage medium may include, for example, one or more of a hard disk, random access memory (RAM), read-only memory (ROM), storage devices of a distributed computing system, an optical drive (compact disc (CD), digital versatile disc (DVD), or Blu-ray disc (BD)), a flash memory device, a memory card, and the like. Specific details and embodiments in which the present invention may be practiced are described by way of example.

[0025] As used herein, unless otherwise specified, the use of ordinal adjectives such as "first," "second," "third," etc. to describe a common object merely indicates that different examples of similar objects are being referred to, and is not intended to imply that the objects so described must be in a given order, whether temporally, spatially, in ranking order, or in any other way.

[0026] Other embodiments may be utilized and structural or logical changes may be made without departing from the scope of the present disclosure. Therefore, the following detailed description should not be taken in a limiting sense. For the present disclosure, the phrase "A and / or B" means (A), (B), or (A and B). For the present disclosure, the phrase "A, B, and / or C" means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C). References to "an embodiment" or "embodiments" in the present disclosure mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The appearances of the phrase "in an embodiment" or "in embodiments" are not necessarily all referring to the same embodiment. The appearances of the phrases "for example," "in one example," or "in some examples" are not necessarily all referring to the same example.

[0027] FIG. 1A includes a block diagram of a microcontroller or microcontroller unit (MCU) 100 according to one or more exemplary embodiments of the present disclosure.

[0028] The MCU 100 or microcontroller 100 includes one or more cores 10, a memory or memory circuit 20, one or more intellectual property (IP) blocks 30, and one or more other applications 40. Other components may be included but are not shown. Although not depicted in FIG. 1, connections between the components of the microcontroller 100 may be assumed. The MCU 100 may be implemented as part of a SoC.

[0029] Referring to FIG. 1A, the one or more cores 10 may be processor or central processing unit (CPU) cores. The one or more cores 10 can execute one or more operations by executing program instructions or software. For example, the application 40 may be in the form of instructions executable by the one or more cores. The instructions or software described herein may be stored or located in a (non-transitory) computer-readable storage medium located within the microcontroller 100 or SoC, or otherwise accessible to the core 10. For example, the MCU 100 can include a memory circuit 20 that may include a computer storage medium for storing instructions executable by the core 10.

[0030] The intellectual property (IP) block 30 may mean a reusable element of logic, circuitry, software, or chip layout. The IP block or IP may, in some cases, support multiple functions implemented by one or more devices included within the IP block and / or may be at least partially implemented by one or more processor cores.

[0031] In addition, one or more peripheral devices 60 can be operably coupled to the MCU 100.

[0032] Generally, in an electronic device including a digital logic device, power consumption includes a leakage component and a dynamic component. Leakage power is generally a function of manufacturing process parameters (threshold voltage, mobility, etc.) and voltage. However, dynamic power depends on the switching activity of logic gates within the integrated circuit logic. The consumption of dynamic power can be written as Equation (1). Equation (1):

Number

[0033] Typically, modern silicon designs use several similar cells called standard cells throughout the logic design. For example, many state machines are implemented using a specific type of flip-flop, and each flip-flop consists of standard logic gates with specific threshold voltages. Each of the standard cells can be summarized, for example, for the calculation of dynamic power. In that case, the total dynamic power can be written as Equation (2). Equation (2):

Number

[0034] The above two equations for calculating dynamic power can be used together for different parts of the digital logic. For example, Equation (1) can be used to estimate the power consumption of subsystems and IPs with high operating frequencies, and Equation (2) can be used for other logics. Therefore, the total dynamic power can be calculated as Equation (3). Equation (3):

Number

[0035] Modern microcontrollers or microprocessors can contain millions of logic gates. Calculating the dynamic power for these millions of logic nodes / networks increases the computational complexity. If a finite set of applications / use cases (and thus chip configurations) is known, the configuration-based dynamic power for a given cluster of logic can be correlated to the aggregate gate-based dynamic power for other sets / clusters of logic. Thus, the total dynamic power can be expressed as Equation (4). Equation (4):

Number

[0036] Using this method recursively significantly reduces the computational complexity of calculating the dynamic power. Next, these nested equations can generally be expressed as Equation (5). Equation (5):

Number

[0037] Equation (5) shows that the dynamic power calculation can be expressed as a series of multiplication and addition operations. Thus, the dynamic power can be represented as a network of operations, e.g., every node is a MADD circuit. Input A k can be the activity α or the number of cells N i and so on. Next, the set of weights (W, X, Y) substantially represents the remaining terms (e.g., the capacitance, voltage in Equation (4)). The activity of a particular node and the number of logic gates of a given cell type are supplied as inputs.

[0038] Figure 1B shows Table 150 that illustrates various configurations for, for example, an MCU or an SOC. Configurations 1 through 4 (CFG1-CFG4) can be for specific IPs of the MCU / SOC. Each of the configurations CFG1-CFG4 can execute a fast Fourier transform (FFT) using 1 MB of memory at different resolutions. Table 150 shows the associated power estimates for each configuration, i.e., the active or dynamic power (P dyn )), the active load jump current, and the active load jump duration. In one or more examples of the present disclosure, the power estimates can be provided by hardware in real time.

[0039] Figure 2 shows a block diagram of at least one example of a hardware power estimator (HPE) 200 according to one or more aspects of the present disclosure.

[0040] HPE 200 can be implemented in hardware, for example, by hardware components. Further, the hardware components can also be wired. In other words, HPE 200 can include hardware components or circuits configured to perform one or more functions without having a processor and without using software code or executing instructions.

[0041] HPE 200 can be implemented as part of a microcontroller / SoC or can be external to it. HPE 200 can also be formed as part of the SoC.

[0042] HPE 200 is configured to generate an output 250 that indicates an approximation or estimate of the power for an electronic device, this type of microcontroller, or an MCU (see, for example, MCU 100 in FIG. 1). In some examples, the output 250 can indicate or identify one or more power consumption related values, such as current, voltage, or power (watts).

[0043] In other examples, the output 250 may function as an indicator that indicates, for example, one of several potential power results or an item in a table or list that indicates a range of potential results.

[0044] Furthermore, in at least one example, the HPE 200 can output a power consumption estimate or an approximation of power consumption in real time.

[0045] As shown in FIG. 2, the HPE 200 includes a scaling circuit 230 and a multiply-accumulate (MADD) circuit or circuitry 240.

[0046] In one or more examples, the scaling circuit 230 can be configured to obtain or receive the input 210 and, in response, generate a scaled non-linear output 235 of the input 210. For example, the scaling circuit 230 can apply an activation function or an approximation of an activation function to the obtained input. An activation function is a mathematical function that can introduce non-linearity in various processes. For example, an activation function can transform or map input data so that the input data becomes non-linear in a useful way. Some well-known activation functions include the sigmoid function, the hyperbolic tangent (tanh) function, the rectified linear unit function (ReLU or Relumax), and the exponential linear unit (ELU).

[0047] For example, the tanh function maps input values to a range between -1 and 1. When the input is positive, the ReLU function outputs the input value directly, and when the input is negative, it outputs zero, thus introducing non-linearity and sparsity.

[0048] The input 210 to the HPE200 can be a signal (e.g., a digital signal) containing data values representing the status information or activity information of one or more components of an electronic device or intellectual property (IP) blocks. For example, various types of signals of such data can be used as inputs when indicating the activity or status information regarding one or more IP blocks or functional components of the MCU100. For example, the input 210 can include one or more of the following signals, or can include data from one or more of the following signals. - IP enable or disable signal, - IP configuration register or logic signal (the IP activity may be stored or captured as data in a configuration (CFG) register), - Clock configuration register or logic signal, - IP performance monitoring signal, - Bus utilization equivalent signal, - State transition signal of a hardware finite state machine, or - Power state signal.

[0049] These signals can include data indicating the activity or status of IP blocks, components (functional blocks) of the MCU. That is, these signals can include data or values indicating ~.

[0050] FIG. 3 shows a block diagram of a scaling circuit 300 according to an aspect of the present disclosure. The scaling circuit 230 of FIG. 2 can be implemented or carried out in a manner similar to the configuration of the scaling circuit 300. As depicted, various hardware components designed to perform specific functions may be present within the circuit 300. In other words, the scaling circuit 300 can include tasks such as performing shift operations, multiplications, addition operations, and summations. In addition, the scaling circuit 300 incorporates a comparator circuit and handles the task of comparing values.

[0051] In at least one example, the components of the scaling circuit 300 may be wire-connected and / or implemented using logic devices without using instructions executed by a processor.

[0052] The scaling circuit 300 can execute an approximation of an activation function, such as a tanh activation function or a Relu activation function, on the received input. That is, the scaling circuit 300 can scale or transform the received input according to the approximation of the activation function.

[0053] The scaling circuit 300 in the example of FIG. 3 can execute an approximation of the activation function of the input 310 in a piecewise manner. For example, from the received input, a plurality of processes or operations can be executed separately or in parallel on the input to generate a plurality of outputs. These outputs (outputs 370a - 370g in FIG. 3) generated from the individual piecewise operations or processes are summed by the summing circuit 380 to generate the total or final output 390 for the scaling circuit 300.

[0054] To generate the piecewise output, the scaling circuit 300 can execute one or any combination of the following functions or operations, namely, multiply - add operations, shift operations, and comparison operations. Since the input 310 for the scaling circuit can be digital, binary operations such as shift operations can be executed. Shifting a digital value to the left can be considered substantially a multiplication by 2. Similarly, a right - shift operation or righting is considered a division by 2. For this type of operation, the least - significant bit is assumed to be on the right and the most - significant bit is located on the left.

[0055] As depicted in FIG. 3, one partition output 370a can be generated or formed by performing a multiply-accumulate operation (using the multiply-accumulate circuit 320), which uses the input 310 (first operand) and the input 310 (second operand) shifted left twice (using the left shifter circuits 325 and 335).

[0056] Similarly, the output 370b can be generated by performing a multiply-accumulate operation (via the MADD circuit 345) from the input 310 (first operand) and the input 310 (second operand) after being shifted left three times (using the left shifter circuits 325, 335, and 340).

[0057] Other partition outputs 370c can be generated simply by shifting the input 310 left three times using the left shifters 325, 335, and 340. Other partition outputs 370d can be generated by shifting the input 310 left twice using the left shifters 325 and 335. The partition output 370e can be generated by shifting the input 310 left once using the left shifter 325.

[0058] Furthermore, one partition output 370f can be simply the input 310 itself without performing any operation.

[0059] In addition, other partition outputs 370g can be generated by using the comparator circuit 350, which compares the input 310 with a predetermined value or threshold to generate the output 370g.

[0060] Next, the partition outputs 370a - 370g can be summed by the sum circuit 380 to generate the output 390. Depending on a predetermined threshold, the output 390 can approximate the result of an activation function, such as the hyperbolic tangent (tanh) or Relu activation function, applied to the input 310.

[0061] FIG. 4 shows a graph diagram 400 depicting an ideal hyperbolic tangent plot 410 and a plot 420 of an approximation of the hyperbolic tangent function achieved through a circuit such as, for example, scaling circuit 300.

[0062] In the HPE200 of FIG. 2, after the input 210 is scaled or transformed by the scaling circuit 230, the MADD circuit 240 operates on this scaled or transformed input. The MADD circuit 240 can include one or more stages.

[0063] FIG. 5 shows an example of a single-stage multiply-accumulate (MADD) or multiply-accumulate (MAC) circuit 500. FIG. 6 shows an example of a two-stage MADD circuit 600. FIG. 7 shows a three-stage MADD circuit 700. The MADD circuit 240 depicted in FIG. 2 can be implemented in a manner very similar to the configurations of the MADD circuits 500 - 700.

[0064] The MADD circuits 500 - 700 are configured to perform multiplication and addition / accumulation operations and output an estimated power consumption value. Thus, the MADD circuits 500 - 700 can execute or calculate a linear equation and generate an output that reflects an estimate of the dynamic power consumption.

[0065] In at least one example, the MADD circuits 500 - 700 may be wire-connected and / or implemented using logic devices (e.g., gates) without using instructions executed by a processor.

[0066] As mentioned, the single-stage MADD circuit 500 is configured to perform multiplication and accumulation of two sets of numbers. In other words, the MADD circuit 500 can be regarded as a simple two-input digital circuit for performing a 4-bit multiplication in a single clock cycle.

[0067] The MADD circuit 500 includes an input or input layer 510 coupled to an output or output layer 550 via a single stage 530. In the example of FIG. 5, the input or input layer 510 can have or include two (different) inputs or input vectors, namely, a[0...3] and b[0...3]. The input vector or input 510a can represent a scaled version of the input, i.e., data indicating the number and / or activity of active cells or circuits of at least one IP or function block (e.g., of an MCU). This data can be found in the signals described herein. The input vector b or input 510b can be multiplied and used as the weight when multiplying and adding the input vector a, and can generate a result indicating power estimation. The weight of the input 510b can be derived from the training and optimization process.

[0068] The weights used herein for the MADD circuit may be stored in or provided from any suitable storage device or non-volatile memory circuit or device operably coupled to the MADD circuit. For example, the HPE200 of FIG. 2 may be implemented in the MCU100 of FIG. 1A. The memory circuit 20 may include the weights. In other cases, the registers of the MCU may include these weights.

[0069] The single stage 530 of the MADD circuit 500 includes (four) multipliers 540 and (four) adders 545 connected to determine the output 550. The output or output layer 550 can include a value indicating real-time dynamic power consumption, e.g., the dynamic current of an IP of an MCU or SoC or at least one individual component.

[0070] FIG. 6 shows another MADD circuit 600 according to at least one exemplary embodiment of the present disclosure. The MADD circuit 600 includes two stages, namely, a first stage 630a and a second stage 630b. In one example, the MADD circuit 240 of the HPE200 may be implemented similar to the MADD circuit 600. The MADD circuit 600 includes a multiplier 640 and an adder 645 as shown.

[0071] The MADD circuit 600 includes an input or input layer 610 and an output or output layer 650. For the first stage 630a, the input layer 610 has two inputs or input vectors, namely, A[0…3] or input 610a and B[0…3] or input 610b.

[0072] The input or input vector 610a can be the scaled input described herein. For example, the input 610a can be a scaled version of an input indicating the activity and / or number of active cells of an aspect of the MCU (e.g., IP or functional block). The input vector 610b is a weight or weight value. Also, these weights or weight values may be trained or optimized. The first stage 630a generates a first output, namely output 620.

[0073] For the second stage 630b, another input, input vector C or input 610c is provided for the second stage 630b. The input 610 can be another set of weights or weight values. These weights or weight values may also be optimized or trained according to the aspects of the present disclosure.

[0074] The first input 610a and the second input 610b are used for the first stage 630a. The first output or output 620 of the first stage 630a and the input 610c can be used as inputs for the second stage 630b. The output 650 of the second stage 630b is also the output for the MADD circuit 600. Thus, the output 650 can correspond to the dynamic power consumption for at least one component or IP of a microcontroller (e.g., MCU100) or SoC. The MADD circuit 600 can generate the output 650 in two clock cycles.

[0075] That is, the MADD circuit 600 includes two sequential stages and implements a three-input multiplication that can be generated or achieved in two clock cycles. Expanding on this concept, a multi-stage MADD circuit that can process a four-input multiplication achievable in three clock cycles is further realized. This is shown in FIG. 7, which features the circuit 700. This circuit 700 can be considered basically the same as the other MADD circuits 500 and 600, the only difference being that it includes three stages (730a - 730c).

[0076] Assuming that each stage n is represented by the equation X×kn + Cn, then, according to at least one example, the three-stage network or three-stage MADD circuit 700 may be represented by Equation (6). Equation (6): ((X×k1 + c1)×k2 + c2)×k3 + c3 = X×k1×k2×k3 + C1×k2×k3 + C2×k3 + C3

[0077] Here, X represents the number of instances of IP (scaled by a scaling circuit), kn represents the average current consumed by the IP (a predetermined value), Cn represents the leakage contribution of the IP.

[0078] In one or more examples, the values for k1, k2, k3 and C1, C2 and C3 are different values for the same IP.

[0079] The input value X is a digital value that can represent the following signals (after scaling). - Hardware IP enable / disable logic signals (e.g., Core_en signal; ip_inst_en signal) - IP configuration register / logic signals (e.g., Lockstep_en signal) - Clock configuration register / logic signals (e.g., PLL, DLL, and divider configuration signals) - IP performance monitoring signals (e.g., IPC for a computing core, baud rate equivalent for a communication model) - Bus usage equivalent signals (e.g., bus interconnect usage duty ratio rate) - State transition signals of a hardware finite state machine (e.g., init_to_active signal for an FSM) - Device and / or IP power state signals (e.g., Sx signal)

[0080] This simple three-stage MADD circuit 700 can perform a total of four-input multiplications, three-input multiplications, and two-input multiplications. The MADD circuit 700 includes inputs 710a - 710d with intermediate outputs 720a and 720b and a final output 750.

[0081] FIG. 8 shows a diagram of an HPE network 800 or an aspect thereof. The HPE network 800 can be configured to estimate, for example, dynamic power consumption in real time for an MCU or SoC. More specifically, the HPE network 800 can be implemented similar to the HPE200 of FIG. 2. Further, the HPE network 800 can provide an immediate or real-time estimate of the dynamic power consumption for the functional blocks or IP components of an MCU or SoC 850. The HPE 800 may be implemented in an MCU that includes a plurality of MCUs described herein.

[0082] In the HPE network 800, each node 810a - 810N can represent at least an HPE or an HPE unit that can correspond to a specific IP block. Each node 810a - 810N can represent an HPE unit and, thus, can include, for example, a scaling circuit and a MADD circuit as shown in FIG. 2. The MADD circuits of the HPE units 810a - 810N can be implemented as described herein, except that each MADD circuit is adjustable with respect to its weight to a specific corresponding one of the IPs of the MCU / SoC 850.

[0083] The HPE network 800 may be implemented by a hardware (wired-connected) network including hardware nodes in the form of HPE units (e.g., HPE200) described herein. The MADD circuits of the nodes 810a - 810N may be implemented in the form of a single stage or a multi-stage (e.g., refer to the MADD circuits 500, 600, or 700).

[0084] In some examples, the nodes 810a - 810N may share one or more common scaling circuits (such as the scaling circuit 230 in FIG. 2), so the nodes 810a - 810N may simply represent MADD circuits. In this kind of case, a single scaling circuit may, for example, scale the input from the MCU or SoC 850 and provide the scaled input to each of the nodes 810a - 810N. The HPE network 800 may include a circuit for directing the scaled input to the appropriate nodes. The output of each node can represent, for example, an estimation of the dynamic power consumption (e.g., current, voltage, or power in watts) for the corresponding IP or associated IP of an electronic device (MCU / SoC 850).

[0085] Generally, as long as HPE can support a larger number of gates (such as multipliers and adders) required for the MADD stage for all IPs, current estimation for the entire SoC cannot be done in the order of a few clock cycles. Thus, substantially, the entire HPE network 800 can be implemented to estimate the power consumption for the MCU or SoC within 10 to 16 clock cycles. This is because power estimations from different MADD circuits for all IPs or components can be performed simultaneously or in parallel.

[0086] Various embodiments herein depict and describe a hardware-based power estimator (HPE) that can estimate the power consumption of an SoC within a few clock cycles, thereby enabling faster dynamic power management.

[0087] The various HPEs described herein may be used in both pre-silicon and post-silicon phases using simple multiplication and addition circuits that can estimate dynamic power consumption at high speed in real time.

[0088] In the post-silicon phase, an HPE implemented as part of the silicon can be used to estimate the power consumption of an application / IP block. Since the HPE can estimate power consumption with very low latency, it can estimate faster transients / changes in power consumption. This estimated power can be compared with the actual power consumption measured using the power supply (powering the silicon / DUT). As shown in FIG. 9, a learning algorithm is used to minimize the least mean square (LMS) loss function and reach the optimized weights for the HPE.

[0089] Various embodiments relate to training a power estimator based on the configuration of each IP, subsystem, SoC, or overall system. Learning algorithms, LMS (Least Mean Square) multivariate curve fitting (for a particular subsystem or IP), and neural networks (for other IPs and complete system-on-chip (SoC)) are used to arrive at the coefficients for this estimator (HPE). For example, the model for a learning algorithm (LA) can be trained by pre-silicon current measurements (using a simulator such as PrimePower(R)) or post-silicon (using a power supply). Next, the model can be enabled by a separate set of application code / software. Further, HPE embodiments are scalable to accommodate an increase in the number of IP instances between derivative products (or different architectures). Further, HPE includes both the leakage and dynamic components of silicon power consumption, thereby explaining PVTF variations.

[0090] Referring back to Equations (4) and (5), the linear equations work well for each node. Therefore, the following function can be used to train the model for HPE. Equation (7): y = max(0, x)

[0091] For simple linear regression, it can be shown that the minimum set of training data sets required to arrive at one potential set of weights (W, X, Y) is Equation (8). Equation (8): (n + m) 5 / 3 +(n + m)+(n + m) 1 / 3

[0092] However, executing the training and arriving at an appropriate set of weights is achieved using a loss function and can maximize accuracy. For example, a loss function can be used that represents the least mean square (LMS) error of power estimation for a set of leaf cells. Since Equation (8) is a multivariate non-linear polynomial, typically, the training dataset is 5 - 10 times more than necessary, and a modified gradient descent algorithm is used to optimize the weights.

[0093] FIG. 9 shows an exemplary flow diagram and environment 900 that represents an exemplary post-silicon training process for determining a weight value or set of weights used for the HPE or a portion thereof (e.g., MADD circuit) described in this specification.

[0094] For post-silicon training, a power estimator (PET) or HPE 930 is already implemented as part of the silicon or MCU 915. Thus, the HPE 930 is operable to estimate the power consumption of an MCU IP block or application, such as any IP which is IPX920 in this example.

[0095] The MCU or SoC 915 can be provided with an input, e.g., code or input pattern 905 that operates or functions at least one IPx920. The corresponding parameters generated by the IPx920, e.g., the number of active cells, can be captured and used as an input to the HPE 930, e.g., as a signal described in this specification. Thus, the HPE 930 can generate an output, e.g., an estimated power consumption. This can be, for example, in the form of an estimated current I est 935. The HPE 930 may already be configured or set with initial values for a set of weights 945 (Wi,Xj,Yl) that can be stored within the registers of the MCU. Also, the HPE 930 can provide an estimation of power consumption with very low latency. It can dynamically estimate relatively fast or faster transient phenomena or changes in power consumption.

[0096] FIG. 9 shows the estimated power 935 compared to the actual power consumption 945 measured using the power supply 940. The power supply 940 can be the supply power to the device under test (DUT), e.g., an SoC or MCU 915.

[0097] The difference between the measured power 945 and the estimated power 935 from the HPE 930 can be used as an input to the learning algorithm (LA) 950. The learning algorithm is executable as instructions (e.g., stored on a non - transitory computer - readable medium) and can be executed by one or more processors on separate computing devices, for example. The LA 950 uses the current weights 955 and the difference in power measurement values between the directly measured power consumption and the estimated power consumption to determine the optimized weights 955 used by the HPE 930. In particular, the LA 950 can be configured to determine an optimized set of weights 955. Any suitable (machine) learning algorithm or technique can be used to find the optimized weights 955. In one example, the LA 950 can use the least - mean - square (LMS) loss function and can reach the optimized weights for the HPE 930.

[0098] FIG. 10 shows an exemplary flow diagram and environment 1000 representing an exemplary pre - silicon training process for determining the weight value or set of weights for the HPE described herein. For the pre - silicon training process, the SoC / MCU 1010 and its components, e.g., IPx 1020 and HPE 1030, are not physically realized and implemented. Instead, the SoC / MCU 1010 can be represented as an abstract data form for simulation data, e.g., data used for RTL simulation or a similar type of simulation.

[0099] During the simulation, different inputs, such as code or pattern 1005, can be input into the simulated SoC / MCU 1010, and thus one or more operations or tasks from IPx 1020 can be performed. Similar to post-silicon training, the simulated HPE 1030 can generate an output of estimated power consumption. That is, parameters generated by IPx 1020, such as the number of active cells, can be captured and used as an input to HPE 1030. Also, the power consumption can be, for example, in the form of the estimated current I est 1035. Similarly, HPE 1030 may be simulated with an initial value for the set of weights 1055 (Wi,Xj,Yl) used for HPE / PET 1030, which can be updated for later simulations based on the weights determined by LA 1050.

[0100] Also, the estimated power consumption 1035 can be comparable to the power consumption 1045 generated by the simulated power estimator 1040, and the simulated power estimator 1040 can also be the simulated power supply for SoC / MCU 1010.

[0101] Similar to the post-silicon phase, the difference between the simulated measured power 1045 and the simulated estimated power 1035 from HPE 1030 can be used by the learning algorithm (LA) 1050 to find optimized weight values.

[0102] Also, LA1050 can be implemented as instructions (stored, for example, on a non-transitory computer-readable medium) and executable by one or more processors (on separate computing devices, for example). LA1050 receives a difference in power measurements, similar to the current weight set. Using this type of input, LA1050 can be configured to determine an optimized weight set 1055. LA1050 can use a least mean squares (LMS) loss function and can reach an optimized weight for HPE1030. The LA can be applied repeatedly or iteratively to multiple simulations to update or find the best or optimized set of weights that produces a minimum error in power consumption measured by the simulated HPE1030. Further, since the HPE is not physically realized, HPE1030 itself can be repeatedly optimized, or its design or configuration can be updated accordingly to determine appropriate results.

[0103] Note that it is not necessary for all application code used in this training set to be functional (with respect to functionality). During pre-silicon analysis, a vector-driven approach can be used to achieve higher coverage of internal nodes. Therefore, a set of patterns (or code) is used to estimate the logical activity using simulation or emulation. A Fast Signal Data Base (FSDB) captures all node / signal activity and is used as input to industry-standard power estimation tools (such as PrimePower(TM) from Synopsys(R), Voltus(TM) from Cadence(R), etc.). As described later, the power consumption (I meas ) obtained from these tools is comparable to the estimated power consumption I est from the realized or implemented HPE. A learning algorithm can be used to minimize the LMS loss function and reach an optimized weight for the HPE.

[0104] Without loss of generality, the HPE described in this specification can be regarded as a hardware accelerator for accurate power estimation. Further, the HPE can be used in both pre-silicon and post-silicon phases and can assist in improving the overall system power performance as follows.

[0105] - Dynamic power management scheme In actual applications / usage cases, the HPE can be used to estimate that the power is at a minimum (a few clock cycles) latency. This feature of the HPE is available for several power management schemes such as dynamic voltage control (DVC), dynamic voltage and frequency scaling (DVFS), envelope tracking, etc.

[0106] - Predictive software-based performance optimization The application can be designed such that the estimated power using the HPE is available to the processor within a few clock cycles before the actual code is implemented. In this case, the core can prepare itself by instructing the voltage regulator circuit about the need to sink (draw in) or source more current.

[0107] - Architecture investigation The RTL implementation of the HPE can be used for comparative analysis of architectures. For example, different bus / bridge topologies can be evaluated using the previously trained HPE RTL that models the slave IP.

[0108] - On-chip and system diagnosis and debugging Using real-time data from HPE, problems within the SoC can be diagnosed. For example, using tracking data from HPE, the previous history of activities within the SoC can be evaluated prior to an unexpected alarm / reset / event. Similarly, by combining HPE trace data from different silicon within a given system, the overall state of the system at a given point in time can be reached.

[0109] FIG. 11 shows a method 1100 according to an aspect of the present disclosure. Method 1100 includes, at 1110, obtaining an input comprising data values representing state information or activity information of one or more components or intellectual property (IP) blocks of an electronic device.

[0110] At 1120, method 1100 includes generating a scaled non-linear output using a scaling circuit by applying an approximation of an activation function to the input.

[0111] At 1130, method 1100 includes generating an output indicative of a mathematical power estimate of the electronic device using the generated scaled non-linear output, using a multiply-accumulate (MADD) circuitry.

[0112] FIG. 12 shows a method 1200 according to an aspect of the present disclosure. Method 1200 includes, at 1210, providing an input from one version of an electronic device to a prototype of a hardware power estimator.

[0113] At 1220, method 1200 includes determining an output from the prototype of the hardware power estimator based on the provided input.

[0114] At 1230, method 1200 includes determining a power measurement value of the electronic device corresponding to the mathematical power estimate indicated by the output of the prototype of the hardware power estimator.

[0115] At 1240, method 1200 includes applying a learning algorithm to the difference between the determined output from a prototype of a hardware power estimator and the determined power measurement value to derive optimized weight values for the MADD circuitry of the prototype of the hardware power estimator.

[0116] The HPE described herein can be used with other software and debugging tools to predict the power consumption for a given application code, thereby assisting in optimizing the application code in an efficient manner. Thus, the initial information is available to the design and architecture teams for planning silicon parameters.

[0117] The following examples relate to further aspects of this disclosure.

[0118] Example 1 is a hardware power estimator (HPE) configured to provide power estimation for an electronic device, the hardware power estimator including a scaling circuit configured to receive an input and generate a scaled non - linear output of the input by applying an approximation of an activation function to the input, and a multiply - accumulate (MADD) circuitry configured to generate an output indicative of a mathematical power estimation of the electronic device using the scaled non - linear output, wherein the input to the scaling circuit comprises data values representing state information or activity information of one or more components or intellectual property (IP) blocks of the electronic device.

[0119] Example 2 is the subject matter of Example 1, and the scaling circuit may include a plurality of hardware components configured to execute an approximation of the activation function on the received input.

[0120] Example 3 is the subject matter of Example 2, where the scaling circuit is configured to apply an approximation of the activation function in a piecewise manner and generate a plurality of piecewise outputs in parallel, and the scaling circuit may be further configured to sum the generated piecewise outputs to generate the output of the scaling circuit. [[ID=*]] [[ID=*]]

[0121] [[ID=*]] Example 4 is the subject matter of Example 3, and in order to generate at least one piecewise output, the scaling circuit may be configured to perform a multiply-accumulate operation on the input. [[ID=*]] [[ID=*]]

[0122] [[ID=*]] Example 5 is the subject matter of Example 3, and in order to generate at least one piecewise output, the scaling circuit is configured to apply one or more shift operations to the input. [[ID=*]] [[ID=*]]

[0123] [[ID=*]] Example 6 is the subject matter of Example 3, and in order to generate at least one piecewise output, the scaling circuit may be configured to apply at least one or more shift operations and at least one multiply-accumulate operation to the input. [[ID=*]] [[ID=*]]

[0124] [[ID=*]] Example 7 is the subject matter of Example 3, and in order to generate at least one piecewise output, the scaling circuit may be configured to apply a comparator having a predetermined threshold value to the input. [[ID=*]] [[ID=*]]

[0125] [[ID=*]] Example 8 is the subject matter of any one of Examples 1 to 7, and the scaling circuit configured to apply an approximation of the activation function may include a scaling circuit that applies an approximation of the hyperbolic tangent function. [[ID=*]] [[ID=*]]

[0126] [[ID=*]] Example 9 is the subject matter of any one of Examples 1 to 8, and the scaling circuit configured to apply an approximation of the activation function may include a scaling circuit that applies an approximation of the rectified linear unit function. [[ID=*]] [[ID=*]]

[0127] [[ID=*]] Example 10 is the subject matter of any one of Examples 1 to 9, and the MADD circuit device may include a plurality of MADD circuits arranged in a plurality of stages such that the output of one stage is input to the next stage. [[ID=*]] [[ID=*]]

[0128] Example 11 is the subject of Example 10, and each stage may include a MADD circuit configured to calculate an accumulated sum of values of a product of a first input and a second input to the stage, the first input comprising an output from a previous stage or an output from a scaling circuit, and the second input comprising a predetermined weight value.

[0129] Example 12 is the subject of Example 11, and the predetermined weight value may correspond to one or more electrical parameters and / or states of one or more components of the electronic device.

[0130] Example 13 is the subject of Example 11, and the predetermined weight value may correspond to one or more electrical parameters and / or states of one or more components of the electronic device.

[0131] Example 14 is the subject of Example 11, and the predetermined weight value may correspond to one or more parameters and / or states of one or more hardware accelerators of the electronic device.

[0132] Example 15 is the subject of Example 11, and the predetermined weight value may correspond to one or more parameters and / or states of the IP of the electronic device.

[0133] Example 16 is the subject of Example 11, and the predetermined weight value is a value determinable from a training process.

[0134] Example 17 is the subject of Example 16, and the training process may include providing simulated inputs from a simulated hardware electronic device to a prototype of a hardware power estimator, determining simulated power measurement values, applying a learning algorithm to a difference between a power estimate from the prototype of the hardware power estimator and the simulated power measurement values to derive a predetermined weight value.

[0135] Example 18 is any one of the themes of Examples 10 to 17, and the MADD circuit device may include three or fewer stages and generate an output from an input in three clock cycles or less.

[0136] Example 19 is any one of the themes of Examples 1 to 18, and the hardware power estimator may be configured to provide power estimation for an electronic device in real time.

[0137] Example 20 is any one of the themes of Examples 1 to 19, and the electronic device can be a microcontroller chip.

[0138] Example 21 is any one of the themes of Examples 1 to 20, and the MADD circuit device can be configured to generate an output indicating a mathematical power estimation according to the following formula:

Equation

[0139] Example 22 is any one of the themes of Examples 1 to 21, and the input to the scaling circuit can include an IP enable or disable signal.

[0140] Example 23 is any one of the themes of Examples 1 to 22, and the input to the scaling circuit can include an IP configuration register or a logic signal.

[0141] Example 24 is any one of the themes of Examples 1 to 23, and the input to the scaling circuit can include a clock configuration register or a logic signal.

[0142] Example 25 is any one of the themes of Examples 1 to 24, and the input to the scaling circuit can include an IP performance monitoring signal.

[0143] Example 26 is any one of the themes of Examples 1 to 25, and the input to the scaling circuit can include a bus usage equivalent signal.

[0144] Example 27 is any one of the themes of Examples 1 to 26, and the input to the scaling circuit can include a state transition signal of a hardware finite state machine.

[0145] Example 28 is any one of the themes of Examples 1 to 27, and the input to the scaling circuit can include a power state signal.

[0146] Example 29 is a microcontroller chip that can also include any one of the hardware power estimators of Examples 1 to 28, and the electronic device is a microcontroller chip.

[0147] Example 30 is the theme of Example 20, and the microcontroller can include a plurality of functional blocks or IP blocks, at least some of the functional blocks have at least one state variable associated therewith, and at least some of the functional blocks are connected to the scaling circuit, so at least some of the state variables function as the input to the scaling circuit.

[0148] Example 1A is a method for hardware power estimation for an electronic device, the method including obtaining an input including a data value representing state information or activity information of one or more components or intellectual property (IP) blocks of the electronic device, generating a scaled non-linear output using a scaling circuit by applying an approximation of an activation function to the input, and generating an output indicating a mathematical power estimation of the electronic device using the generated scaled non-linear output using a multiply-accumulate (MADD) circuit device.

[0149] Example 2A is the subject matter of Example 1A, and the scaling circuit can include a plurality of hardware components configured to execute an approximation of the activation function on the received input.

[0150] Example 3A is the subject matter of Example 1A or 2A, and applying the activation function to the input can include generating a plurality of segmented outputs in parallel by applying the activation function to the input in a segmented manner, and summing the generated segmented outputs to generate the called non-linear output.

[0151] Example 4A is the subject matter of Example 3A, and generating at least one segmented output can include the scaling circuit executing a multiply-accumulate operation on the input.

[0152] Example 5A is the subject matter of Example 3A or 4A, and generating at least one segmented output includes the scaling circuit applying one or more shift operations to the input.

[0153] Example 6A is the subject matter of any of Examples 3A to 5A, and generating at least one segmented output may include the scaling circuit applying at least one or more shift operations and at least one multiply-accumulate operation to the input.

[0154] Example 7A is the subject matter of any of Examples 3A to 6A, and generating at least one segmented output may include the scaling circuit applying a comparator having a predetermined threshold value to the input.

[0155] Example 8A is the subject matter of any of Examples 1A to 8A, and the scaling circuit applying an approximation of the activation function to the input can include a scaling circuit applying an approximation of the hyperbolic tangent function to the input.

[0156] Example 9A is the subject matter of any of Examples 1A to 8A, and the scaling circuit applying an approximation of the activation function to the input can include a scaling circuit applying an approximation of the rectified linear unit function to the input.

[0157] Example 10A is a subject of any one of Examples 1A to 10A, and the MADD circuit device may include a plurality of MADD circuits arranged in a plurality of stages such that the output of one stage is input to the next stage.

[0158] Example 11A is a subject of Example 10A, and each stage includes a MADD circuit configured to calculate the accumulated sum of the values of the product of the first input and the second input to the stage. The first input includes the output from the previous stage or the output from the scaling circuit, and the second input includes a predetermined weight value.

[0159] Example 12A is a subject of Example 11A, and the predetermined weight value can correspond to one or more electrical parameters and / or states of the electronic device.

[0160] Example  13A is a subject of Example 11A or 12A, and the predetermined weight value may correspond to one or more electrical parameters and / or states of the components of the electronic device.

[0161] Example 14A is a subject of any one of Examples 11A to 13A, and the predetermined weight value can correspond to one or more parameters and / or states of the hardware accelerator of the electronic device.

[0162] Example 15A is a subject of any one of Examples 12A to 15A, and the predetermined weight value may correspond to one or more parameters and / or states of the IP of the electronic device.

[0163] Example 16A is a subject of any one of Examples 11A to 15A, and the predetermined weight value may be a value determinable from the training process.

[0164] Example 17A is a subject of any of Examples 10A to 16A, the MADD circuit device can include three or fewer stages, and the step of generating an output using the MADD circuit device includes the step of generating an output from a non-linear output scaled in three clock cycles or less.

[0165] Example 18A is a subject of any of Examples 1A to 17A, and an output indicating a mathematical power estimation of an electronic device is generated in real time.

[0166] Example 1B is a method for training a hardware power estimator configured to provide power estimation for an electronic device. The hardware power estimator is a scaling circuit configured to receive an input and generate a scaled non-linear output of the input, the scaling circuit applying an approximation of an activation function to the input, and a multiply-accumulate (MADD) circuit device configured to generate an output indicating a mathematical power estimation of the electronic device using the scaled non-linear output. The training may include providing an input from one version of the electronic device to a prototype of the hardware power estimator, determining an output from the prototype of the hardware power estimator based on the provided input, determining a power measurement value of the electronic device corresponding to the mathematical power estimation indicated by the output of the prototype of the hardware power estimator, applying a learning algorithm to a difference between the determined output from the prototype of the hardware power estimator and the determined power measurement value, and deriving optimized weight values for the MADD circuit device of the prototype of the hardware power estimator.

[0167] Example 2B is a subject of Example 1B, and the prototype of the hardware power estimator may be a simulated version of the hardware power estimator.

[0168] Example 3B is the subject matter of Example 2B, where the electronic device is a simulated electronic device configured to provide its output as a simulated input to a simulated version of the hardware power estimator, and the step of determining the power measurement value includes performing a power measurement of the electronic device in simulation.

[0169] Example 4B is the subject matter of Example 1B, and the prototype of the hardware power estimator may be a physical version of the hardware power estimator.

[0170] Example 5B is the subject matter of Example 4B, where the electronic device is a physical electronic device configured to provide its output as an input to a physical version of the hardware power estimator, and the step of determining the power measurement value includes performing a power measurement of the physical version of the electronic device.

[0171] Note that one or more of the features of any of the above-described examples may be appropriately combined with any of the other examples or the embodiments disclosed in this specification.

[0172] Those skilled in the art will recognize that the above description is given merely by way of example and may be modified without departing from the broader spirit or scope of the invention as claimed. Therefore, the specification and drawings should be regarded as illustrative rather than restrictive.

[0173] Accordingly, the scope of the disclosure is indicated by the appended claims and is intended to include all changes within the meaning and scope of equivalents of the claims.

[0174] It should be recognized that the embodiments of the methods detailed in this specification are of an exemplary nature and are thus understood to be implementable in corresponding devices. Similarly, it should be recognized that the embodiments of the devices detailed in this specification are understood to be implementable as corresponding methods. Therefore, it should be understood that a device corresponding to the methods detailed in this specification may include one or more components configured to perform each aspect of the related methods.

[0175] In addition, all acronyms defined in the foregoing description are effective in all claims included in this specification.

Claims

1. A hardware power estimator configured to provide power estimation for an electronic device, the hardware power estimator comprising: A scaling circuit configured to receive an input and generate a scaled non-linear output of the input, the scaling circuit applying an approximation of an activation function to the input; A multiply-accumulate (MADD) circuit device configured to generate an output indicative of a mathematical power estimate of the electronic device using the scaled non-linear output; Comprising; The input to the scaling circuit comprises data values representing status information or activity information of one or more components or intellectual property (IP) blocks of the electronic device. Hardware power estimator.

2. The scaling circuit comprises a plurality of hardware components configured to execute the approximation of the activation function on the received input. The hardware power estimator according to claim 1.

3. The scaling circuit is configured to apply the approximation of the activation function in a segmented manner and generate a plurality of segmented outputs in parallel. The scaling circuit is further configured to sum the generated segmented outputs to generate an output of the scaling circuit. The hardware power estimator according to claim 2.

4. To generate at least one segmented output, the scaling circuit is configured to perform a multiply-accumulate operation on the input. The hardware power estimator according to claim 3.

5. To generate at least one segmented output, the scaling circuit is configured to apply one or more shift operations to the input. The hardware power estimator according to claim 3.

6. To generate at least one segmented output, the scaling circuit is configured to apply at least one or more shift operations and at least one multiply-accumulate operation to the input. The hardware power estimator according to claim 3.

7. To generate at least one segmented output, the scaling circuit is configured to apply a comparator having a predetermined threshold value to the input. The hardware power estimator according to claim 3.

8. The scaling circuit configured to apply an approximation of the activation function comprises a scaling circuit that applies an approximation of the hyperbolic tangent function. The hardware power estimator according to any one of claims 1 to 7.

9. The scaling circuit configured to apply an approximation value of an activation function includes a scaling circuit that applies an approximation value of a normalized linear unit function. The hardware power estimator according to any one of claims 1 to 7.

10. The MADD circuit device includes a plurality of MADD circuits arranged in a plurality of stages such that the output of one stage is input to the next stage. The hardware power estimator according to any one of claims 1 to 9.

11. Each stage includes a MADD circuit configured to calculate the accumulated sum of the values of the product of the first input and the second input to the stage, The first input includes an output from the previous stage or an output from the scaling circuit, The second input includes a predetermined weight value. The hardware power estimator according to claim 10.

12. The predetermined weight value corresponds to one or more electrical parameters and / or states of the electronic device, The predetermined weight value corresponds to one or more electrical parameters and / or states of one or more components of the electronic device, The predetermined weight value corresponds to one or more parameters and / or states of the hardware accelerator of the electronic device, and / or The predetermined weight value corresponds to one or more parameters and / or states of the IP of the electronic device. The hardware power estimator according to claim 11.

13. The MADD circuit device includes three or fewer stages and generates an output from an input in three clock cycles or less. The hardware power estimator according to claim 11 or 12.

14. The hardware power estimator is configured to provide the power estimation for the electronic device in real time. The hardware power estimator according to any one of claims 1 to 13.

15. The input to the scaling circuit includes an IP configuration register or logic signal, a clock configuration register or logic signal, an IP performance monitoring signal, a bus utilization equivalent signal, a state transition signal of a hardware finite state machine, and / or a power state signal. The hardware power estimator according to any one of claims 1 to 14.

16. A method for hardware power estimation for an electronic device, the method comprising Obtaining an input comprising data values representing status information or activity information of one or more components or intellectual property (IP) blocks of the electronic device; Generating a scaled non-linear output using a scaling circuit by applying an approximation of an activation function to the input; Generating an output indicating a mathematical power estimate of the electronic device using the generated scaled non-linear output by means of a multiply-accumulate (MADD) circuit device; A method comprising.

17. The scaling circuit comprises a plurality of hardware components configured to execute the approximation of the activation function on a received input, Applying the activation function to the input Generating a plurality of segmented outputs in parallel by applying the activation function to the input in a segmented manner; Summing the generated segmented outputs to generate the called non-linear output; Including, The method according to claim 16.

18. Generating at least one segmented output includes the scaling circuit performing a multiply-accumulate operation on the input, Generating at least one segmented output includes the scaling circuit applying one or more shift operations to the input, Generating at least one segmented output includes the scaling circuit applying at least one or more shift operations and at least one multiply-accumulate operation to the input, The method according to claim 17.

19. A method for training a hardware power estimator configured to provide a power estimate for an electronic device, the method comprising: The hardware power estimator is A scaling circuit configured to receive an input and generate a scaled non-linear output of the input, the scaling circuit applying an approximation of an activation function to the input; A multiply-accumulate (MADD) circuit device configured to generate an output indicating a mathematical power estimate of the electronic device using the scaled non-linear output; Comprising, The training is Providing an input from one version of the electronic device to a prototype of the hardware power estimator; Determining an output from the prototype of the hardware power estimator based on the provided input; Determining a power measurement value of the electronic device corresponding to the mathematical power estimate indicated by the output of the prototype of the hardware power estimator; Applying a learning algorithm to a difference between the determined output from the prototype of the hardware power estimator and the determined power measurement value to derive an optimized weight value for the MADD circuitry of the prototype of the hardware power estimator; A method comprising. **Claim 20** The prototype of the hardware power estimator is a simulated version of the hardware power estimator, The electronic device is a simulated electronic device configured to provide its output as a simulated input to the simulated version of the hardware power estimator, The step of determining the power measurement value includes performing a power measurement of the electronic device in simulation, The method according to claim 19.