Logical level metric count extraction from emulation hardware
By instantiating the logic level metric counting module in simulation hardware, the logic level switching and high state in circuit design are monitored in real time, and the problem of slow power estimation speed in the prior art is solved, achieving fast and efficient power analysis.
Patent Information
- Application Number
- CN202510221767.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-28
- Filing Date
- 2025-02-27
- Publication Date
- 2025-08-29
AI Technical Summary
Existing hardware simulation methods are slow in power estimation, making it difficult to complete power analysis of complex circuit designs within a reasonable time, especially long-term software workload simulation.
Instantiated the logic level metric counting module in simulation hardware, through components such as multiplexer, sampling register, previous value memory and switching detector, the logic level switching and high state in circuit design are monitored in real time, and the incrementer is used to calculate the switching count and time high count to achieve fast power estimation.
It significantly improves the power estimation speed, improves efficiency by 380 times compared to traditional methods, and can complete power analysis of complex circuit designs in a short time, supporting longer simulation workloads.
Smart Images

Figure CN120562360A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to circuit emulation and programmable logic devices. Background Art
[0002] The simulation system or simulator may include scalable hardware units, wherein each unit may include programmable logic blocks interconnected by high-speed links. The programmable logic blocks may include elements such as registers, gates, and memory primitives. These elements may be interconnected by switches, which may be configured to implement connection functions between the elements. Additionally, the gates may be programmed to simulate any Boolean logic function of finite variables. A simulation compiler may be assigned to map the input register transfer level (RTL) design to the target simulation system so that the target implementation is optimized for simulation throughput and capacity utilization of simulator system resources. Summary of the Invention
[0003] An embodiment of the present disclosure provides a method comprising: adding at least one logic level metric count module to a circuit design; loading the circuit design into a simulation system via a processing device; applying a simulation workload to the simulation system into which the circuit design is loaded via the processing device; obtaining at least one logic level metric count from the at least one logic level metric count module, the at least one logic level metric count being associated with at least a portion of the simulation workload; and presenting at least one power utilization estimate for the circuit design, wherein the at least one power utilization estimate for the circuit design is based on the at least one logic level metric count.
[0004] Another embodiment of the present disclosure provides a programmable logic device comprising: a multiplexer having: a plurality of inputs of a plurality of sampling registers associated with a plurality of nodes of a circuit design; and a select line for selecting an input of the plurality of inputs to pass to an output of the multiplexer; a previous value memory coupled to the output of the multiplexer to store a value of the output of the multiplexer from a current simulation clock cycle until a next simulation clock cycle in a plurality of simulation cycles of a simulation workload applied to the programmable logic device; a toggle detector having a first input for the output of the multiplexer and a second input for the output of the previous value memory, wherein the toggle detector is configured to output a logic high value when a first value on the output of the multiplexer is different from a second value on the output of the previous value memory; and an incrementer unit coupled to the output of the toggle detector, wherein the incrementer unit is configured to increment a count within each polling clock cycle in which the output of the toggle detector is a logic high value.
[0005] Yet another embodiment of the present disclosure provides a non-transitory computer-readable medium comprising a stored circuit design that, when loaded onto at least one programmable logic device, configures the at least one programmable logic device to include: a multiplexer having: a plurality of inputs to a plurality of sampling registers associated with a plurality of nodes of the stored circuit design; and a select line for selecting an input of the plurality of inputs to pass to an output of the multiplexer; and an incrementer unit coupled to the output of the multiplexer, wherein the incrementer unit is configured to increment a count within each polling clock cycle in which the output of the multiplexer is a logic high value. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] The present disclosure will be more fully understood from the detailed description given below and the accompanying drawings of embodiments of the present disclosure. The accompanying drawings are intended to provide knowledge and understanding of the embodiments of the present disclosure and are not intended to limit the scope of the present disclosure to these specific embodiments. In addition, the drawings are not necessarily drawn to scale.
[0007] Figure 1 An example logic level metric counting module / circuit of the present disclosure is described.
[0008] Figure 2 An example of using multiple logic level metric counting modules in a simulation system according to the present disclosure is described.
[0009] Figure 3 Two examples of an aggregated logic level metric counting module according to the present disclosure are illustrated.
[0010] Figure 4 An example simulation system including several programmable logic devices configured to simulate a circuit design according to the present disclosure is described.
[0011] Figure 5 A flow chart illustrating an example method for constructing a simulation system to include a merged memory implementation using memory primitives of a memory primitive type to simulate a selected logical memory.
[0012] Figure 6 Flowcharts depicting various processes used during the design and fabrication of integrated circuits according to some embodiments of the present disclosure.
[0013] Figure 7 A diagram depicting an example simulation system according to some embodiments of the present disclosure.
[0014] Figure 8 A diagram depicting an example computer system in which embodiments of the present disclosure may operate. DETAILED DESCRIPTION
[0015] Aspects of the present disclosure relate to, for example, circuit simulation within a circuit design process. For example, a circuit design can be simulated using programmable logic to verify the functionality of the circuit design. A simulation compiler can map an input circuit design (e.g., a register transfer level (RTL) design) to a target simulator (e.g., a programmable logic device / system) such that the target implementation is optimized for simulation throughput and capacity utilization of the simulator's resources.
[0016] In one example, the simulator may be composed of scalable hardware units (e.g., programmable gate arrays, such as field programmable gate arrays (FPGAs), etc.), where each unit is a collection of programmable logic blocks interconnected by high-speed interconnect links. The programmable logic blocks may include elements such as registers, gates, and memory primitives. In one example, these elements may be interconnected by switches that may be configured to implement any connection function between the elements. Similarly, gates may be programmed to emulate any Boolean logic function of finite variables. Memory primitives may include monolithic blocks of random access memory (referred to as Block-RAM) or lookup table RAM (LUTRAM). Hardware units may be added or removed from the simulator based on the quality of the RTL design logic and / or the mapping generated by the simulation compiler. When mapping the RTL design onto the simulator, the simulation compiler anticipates optimal use (e.g., expected use) of all available elements (registers, gates, and memory primitives).
[0017] One aspect of electronic circuit design, such as system-on-chip (SoC) development, is power estimation. In many cases, SoC designs have specific power usage targets to be achieved. Circuit developers can use simulations and emulations to run circuit designs using software programs that are identical or similar to those likely to be encountered based on the expected use of the electronic circuit design. These programs, which can be referred to as workloads, can be used to obtain more realistic power information during the power estimation process. To reduce the amount of time spent on power estimation, simulations can be used, in which the circuit design being analyzed is simulated in simulation hardware.
[0018] In one approach, simulation hardware is used to extract waveforms from the design. These waveforms can be used in software-based power analysis tools to extract switching information. Power consumption can be calculated from the switching information by power estimation software. However, despite the use of hardware simulation in such approaches, the actual time required to obtain power estimation information is relatively slow. For example, typical performance may be less than 1 kHz. In contrast, examples of the present disclosure may modify / supplement simulation circuit designs to use the capabilities of simulation hardware to collect and calculate logic level metric counts, such as switching counts (TC) or logic level high counts, which may be referred to as T1 counts (and logic level low counts) for nodes in the circuit design. For example, simulation hardware significantly increases the speed of calculating logic level metric counts (e.g., TC and / or T1 counts) compared to software-based approaches. This further enables circuit designers to run long software workloads on the simulated design. As a result, characterizing much longer runs for power estimation is possible.
[0019] In one example, the present disclosure may be instantiated within simulation hardware (e.g., within one or more programmable logic devices) to obtain logic level metric counts for nodes in a circuit design and then transmit such information to one or more circuits or modules at a workstation. In one example, a counter module / circuit may be configured to detect a toggle count (TC). In another example, a counter module / circuit may be configured to detect a signal that remains constant at a monitoring node, such as time-at-1 (T1). The module may perform a certain number of logic level metric counts within a simulation cycle. Additionally, the module may include shared circuitry for performing logic level metric counts across multiple nodes, thereby reducing module overhead. As mentioned above, the counting module / circuit may be added to a design under test (DUT) in a simulator / simulation system. For example, the DUT may be loaded into one or more FPGAs (or other PLDs). Along with the DUT, a counter module / circuit may be instantiated in the FPGA to collect logic level metric counts (toggle and / or T1) and communicate with a management system that can further process the results.
[0020] In one example, two modes of power analysis can proceed from the collected counts. For example, both TC and T1 counts can be used to estimate average-mode power. Alternatively, or in addition, only TC can be used to generate an estimated weighted power based on the influence cone of the flip-flop. To capture both TC and T1, the PLD's memory allocation may be doubled (e.g., with two separate counters for TC and T1 counts, respectively). However, other resources can be shared. To further illustrate, for average-mode power calculation, TC can be converted to an activity rate or switching rate (TR) = TC / Emulation_Duration. Similarly, T1 can be converted to a probability of 1 (Prob) = T1 / Emulation_Duration. The switching rate and probability can then be annotated onto various signals, such as sequential outputs, power input (PI), bounding box outputs, etc. In one example, statistical propagation can be applied to estimate TR and Prob for various signals, after which the average-mode power of the circuit design can be calculated.
[0021] For power window detection, the present disclosure can generate TC and T1 for various time slices / time windows, and can generate average mode power for each slice. Over many cycles (e.g., billions), the average power per slice can provide a high-level snapshot of slices with high power consumption (and / or slices with low power consumption, average power consumption, etc.). Thus, circuit designers can use power window detection to identify various power windows of interest, for example, to match various events in a software workload. To further illustrate, average power can include two components: dynamic power and static (leakage) power. TC can be used to calculate dynamic power, and T1 can be used to calculate static power. In one example, the dynamic power of a network can be determined using the formula 0.5*C*V^2*F, where C is the capacitance of the circuit design, which can be obtained from the design parasitic file, V is the operating voltage of the design, and F is the switching rate (TR). Similarly, the dynamic power of a cell's pins can be calculated using information present in a technology library, which can be in the form of a lookup table from each input pin to the output pin of the cell. The slope and capacitance of the input pin serve as indices into the lookup table. The slew rate can be calculated using the timing information available in the technology library and design structure. The capacitance can be obtained from the design parasitic file. Using this information, the internal power per switching can be calculated for each pin of the cell. The average power per pin can be calculated by multiplying the switching rate of the pin. Therefore, the average dynamic power of the entire circuit design can be calculated by summing all net and pin powers.
[0022] In addition, static (leakage) power data can be present in the technology library in the form of various input state conditions. For example, for a 2-input AND gate, there are four input combinations: A=B=1, A=B=0, A=1B=0, and A=0B=1. The leakage of each of the input state conditions can be specified. T1 can then be used to determine the probabilities of A and B. Next, using the input pin probabilities, the probability of each input state condition can be determined. Next, each individual state condition probability is multiplied by its corresponding leakage power and the results are added to provide the leakage power of the cell. Therefore, the leakage power of all cells can be added to calculate the average leakage (static) power of the entire circuit design.
[0023] Notably, the examples of the present disclosure can provide a performance rate of approximately 200KHz. In comparison, other hardware simulation methods extract waveforms from the simulator at a maximum rate of approximately 45 kilohertz (KHz) and convert the waveforms to switches at a maximum rate of approximately 650Hz (the total run time is the sum of the time to complete these two stages). Therefore, compared to previous hardware-based methods, the examples of the present disclosure provide a speed increase of approximately 380 times for the same power estimation workload. To further illustrate, a simulation run that takes one (1) hour using the examples of the present disclosure may take at least 380 hours (approximately 16 days) using previous methods, which may be an unreasonable amount of time to complete a power estimate for a single workload. Additionally, considering scenarios in which there may be several different workloads, the advantages over previous methods are even more significant.
[0024] It should be noted that a workload can be a software program used to measure the performance of a circuit design / model via simulated hardware. For example, different workloads can be applied in hardware simulation to simulate streaming video, engaging in social media, etc. Thus, workloads can be run to determine processor and memory requirements, for example, in addition to power requirements. For a 1GHz processor, one second of execution of the workload may span 1 billion cycles. To simulate playing a video game for one minute, this corresponds to 60 billion clock cycles. However, for one second at 1GHz, a 200KHz simulation may take 5,000 seconds (approximately 2 hours of simulation for only one second of real-time workload).
[0025] Technical advantages of the present disclosure include, but are not limited to, improved power estimation for circuit designs, including significantly faster power estimation compared to other methods and / or the ability to increase the amount of workload over the same duration. Examples of the present disclosure also provide an improved computing device or system for implementing examples of the present disclosure. For example, the computing device may add a counter module to a circuit design, and such circuit design may be compiled and loaded into a simulation system to provide a significantly faster hardware-based power estimation process. Additionally, the simulation process is improved to the extent that examples of the present disclosure may add a hardware-based logic level metric counting function that may be activated and used in conjunction with various test workloads. In one example, for example, the simulation system is improved by enhancing the ability of the simulation system to provide a hardware-based logic level metric counting function in addition to the other capabilities of the simulation system. The following is in conjunction with Figures 1 to 8 These and other aspects of the present disclosure are discussed in further detail in the Examples.
[0026] exist Figure 1 In the illustrative example of FIG, a first example module 100 (e.g., a counter module / circuit of the present disclosure) can include a plurality of sampling registers 110 associated with a plurality of nodes 190 of a circuit design or design under test (DUT). Each of the sampling registers 110 can sample the logic value of a corresponding one of the nodes 190 in the DUT once per simulation clock cycle and can place the sampled / recorded value on the output until the logic value is updated at the next sampling on the next simulation clock cycle.
[0027] The module 100 includes an incrementer 150, which in this example may also be referred to as a switching incrementer circuit (TIC). Figure 1 In an example, module 100 can sample and count the switching of up to 512 nodes (e.g., nodes 190) at a time. For example, module 100 can include a multiplexer 120 that can select between nodes 190 for counting the switching via sampling register 110. Notably, multiplexer 120 (e.g., a 512:1 multiplexer in one example) can operate at a significantly faster clock speed than the simulation clock (e.g., at least 512 times faster), such that multiplexer 120 can select all 512 signals one at a time within one simulation cycle. For example, multiplexer 120 can continuously select sampling register 110 to detect the switching of 512 nodes 190 per simulation clock cycle. To further illustrate, the module 100 may run at a native speed of 100 MHz, or at a polling speed to cycle through 512 samples in 5120 ns (195 KHz), where the DUT's simulation clock cycle may run up to 195 KHz.
[0028] Module 100 further includes a storage element, namely a previous value memory 130, to store previous values from each of sampling registers 110. For example, previous value memory 130 maintains logic values from each of sampling registers 110 from a previous simulation clock cycle so that these logic values can be compared with logic values from the current simulation clock cycle to detect whether and when a switch occurs at node 190. In this regard, for example, if the previous value is not equal to the current value of a given node 190 and / or associated one of sampling registers 110, a switch detection cell 140 (e.g., an edge detector or edge detector circuit) can be used to indicate that a selected one of nodes 190 has switched. For example, the previous value can be from previous value memory 130. In one example, the edge detector can be a two-input lookup table (LUT) to detect edges (or level signals, depending on the particular function). In one example, an address select input (ADDR) of the previous value memory 130 can be used to select a particular address (e.g., bit / element) storing a previous value corresponding to a given node 190 and / or an associated one of the sampling registers 110, wherein the previous value (e.g., a logical one or zero) stored at the selected address is passed to the switch detection cell 140. A second input of the switch detection cell 140 can be the output of the multiplexer 120. Notably, the multiplexer 120 can include a select input (SEL) to select a logical value to be passed from the corresponding one of the sampling registers 110 to the output of the multiplexer 120.
[0029] In addition, module 100 includes a counter, namely an incrementer 150, which increments a previous switching count if a selected one of nodes 190 has switched. In one example, the output (NEW_COUNT) of incrementer 150 at output port CNT_OUT is fed back as an input (PREV_COUNT) to incrementer 150 at input port (CNT_IN). When a switch is detected at switch detection cell 140, the enable input (EN) of incrementer 150 may be a logic value of one (1), in which case incrementer 150 is configured to increase the stored value in incrementer 150 by one. For example, if EN is logic high, the new count value (NEW_COUNT) may be equal to PREV_COUNT+1. Otherwise, if EN is logic low, NEW_COUNT may be equal to PREV_COUNT. In various examples, incrementer 150 may be configured for different use cases. For example, the width (in bits) of incrementer 150 can determine how many cycles of switching data it can store. In one example, incrementer 150 can be a saturating counter that can count up to a certain maximum value (e.g., without overflowing) and then reset. In one example, the output port CNT_OUT of incrementer 150 can have a width such that a sufficient number of bits can be used to store up to the maximum selected value to be stored in incrementer 150.
[0030] In one example, the NEW_COUNT value can be fed to an external memory element (not shown), where it can be stored as the current / previous toggle count for node 190. For example, the external memory element can be instantiated on the same programmable logic device (PLD) as module 100 and / or on a different PLD in the simulation system. Additionally, module 100 (e.g., the toggle detection / counter circuit) can be replicated multiple times with the simulation system, for example, within one or more programmable logic devices thereof, to account for all nodes in the DUT to be counted.
[0031] It should also be noted that Figure 1 An example module 100 is described for detecting and counting switches within a portion of a DUT (e.g., up to 512 nodes thereof). However, in another example, the present disclosure may include a similar module for counting logic highs at nodes within the DUT (e.g., node 190). For example, in this case, the module for detecting logic highs may omit the previous value memory 130, the switch detection cell 140, and the connections between these and other elements. Additionally, the output of the multiplexer 120 may be fed directly to the enable input EN of the incrementer 150. Likewise, in one example, the module may include shared elements for both switch detection / counting and logic high detection / counting. For example, in such an example, the module 100 may include a Figure 1, and may further include a second incrementer fed directly from the output of multiplexer 120. Additionally, in such an example, the output of multiplexer 120 may be split to feed previous value memory 130, toggle detection cell 140, and an additional incrementer for counting the number of logic high values exhibited by node 190 / sampling register 110. Thus, toggle detection / counting and logic high detection / counting operations may be performed in parallel.
[0032] It should also be noted that Figure 1 Only one example of a toggle detection / counting module / circuit is illustrated, and other, additional, and different examples of the present disclosure may provide modules / circuits for logic level metric counting having more or fewer elements, having different elements, etc. For example, as just one additional example, a circuit for logic level low detection / counting may be represented by module 100 without previous value memory 130 and toggle detection cell 140, wherein the output of multiplexer 120 may be fed to an inverter, and wherein the inverter output may be fed to enable input EN of incrementer 150, e.g., to count logic level lows (e.g., logic zeros (0)). Accordingly, these and other modifications are contemplated within the scope of the present disclosure.
[0033] Figure 2 An example of using a plurality of logic level metric counting modules 231 to 234 within a simulation system 200 is described. For example, the simulation system 200 may include at least a first programmable logic device (PLD) 210 (e.g., an FPGA, etc.). A design under test (DUT) 220 may be compiled and loaded into the PLD 210. In other words, the PLD 210 is programmed / configured to replicate the elements and operations of a circuit design as the DUT 220, for example, including a plurality of nodes. The PLD 210 may further include a plurality of logic level metric counting modules 231 to 234, each of which may correspond to a logic level metric counting module / circuit (e.g., Figure 1 Module 100, etc.). The synchronization unit 240 can provide a sampling / polling / counting clock, and all logic level metric counting modules 231 to 234 can operate with the sampling / polling / counting clock, while the DUT 220 can operate according to the lower frequency simulation clock. For example, the sampling / polling / counting clock can operate at a frequency 512 times or greater than the frequency of the simulation clock (e.g., using a 512:1 multiplexer within the logic level metric counting modules 231 to 234).
[0034] The accumulator module 250 may collect and aggregate logic level metric counts (e.g., switching counts (TC) and / or T1 counts) from the logic level metric count modules 231 through 234. In one example, the accumulator module 250 may further create a packet to send to a management station / workstation, i.e., the host 290. For example, the packet may include aggregated logic level metric counts for one or more workloads, one or more time slices within a workload, and the like.
[0035] The accumulator module 250 can map the data from the logic level metric count modules 231 to 234 to the corresponding emulated clock cycles. Using this information, the accumulator module 250 can collect the received logic level metric count information into a local buffer. This information can then be aggregated into a format that can be utilized by the host 290 (e.g., T1 count information and / or toggle counts per time slice, etc.). In one example, the accumulator module 250 can also group this data so that the message control module 260 can communicate the data to the host 290 via a channel, or optionally buffer the data locally in the local storage device 270 (e.g., a local cache) until the host 290 is ready to receive the data. In one example, the accumulator module 250 can perform additional processing on the logic level metric count information. For example, data can be marked as "interesting" or "not interesting." A set of "interesting" data (T1 counts and / or toggle counts) is that which has been indicated or marked as containing information to be stored for later processing. This data can be transmitted to the message control module 260 to ensure that this storage request is met. A set of "uninteresting" T1 counts and / or handover counts are those that do not have a special indication or tag associated with them. This data may be stored or discarded based on a predefined policy or the like.
[0036] Determining which data is "interesting" or "not interesting" can vary from simulation model to simulation model and / or workload to workload. Various external mechanisms exist for this purpose. For example, in one example, accumulator module 250 may have an event or trigger input signal that has one value when data received from logic level metric counter modules 231 through 234 is "interesting" and another value when the data is "not interesting." Accumulator module 250 may be further configured with one or more rules to determine how to handle "not interesting" data (e.g., discard it, send it out, send it out only under certain conditions, etc.). The event / trigger input can be used to mark "interesting" and "not interesting" data based on simple or complex settings and algorithms from one or more sources internal and / or external to PLD 210 (e.g., from host 290, e.g., from user input or other automated processes). To further illustrate, a user or other automated process may indicate that data is "interesting" when temporally correlated with a particular portion of a simulation model run, when a toggle count and / or T1 count exceeds a threshold, when an error indicator or other indicator is received from one or more debug modules or other components of the circuit design, and so forth. Alternatively, or in addition, a user or other automated process may indicate that data is "not interesting" for one or more time blocks within a simulation model run. For example, an initial cycle / time block of a simulation model run may place the circuit in a desired state, where the circuit may then be further tested through more rigorous processing tasks in subsequent cycles / time blocks of the simulation model run. Thus, an initialization portion of a simulation model run may be designated as "not interesting." Various other indicators of “interesting” and / or “not interesting” data of the same or similar nature may be programmed into accumulator module 250, may be programmed into one or more other modules of the circuit design (which may be configured to signal accumulator module 250 when an “interesting” condition (or a “not interesting” condition) is detected), and / or may be indicated during simulation model execution via external input (e.g., from a user or from another system external to PLD 210).
[0037] In one example, PLD 210 can be further configured to include a message control module 260 that can manage the communication channel to host 290, or optionally to local memory 270, which can be used as a cache. For example, message control module 260 can temporarily store a large amount of logic level metric count data, such as logic level metric count values, including TC and T1 counts, for various time slices of one or more workloads. Furthermore, message control module 260 can retrieve any such data and transfer it to host 290 or other entities upon request. For example, message control module 260 can use local storage 270 to save the data until host 290 is no longer busy and can receive the data. Optionally, if no local storage is available, message control module 260 can signal the remaining simulation to stop until host 290 can consume the data to be transferred.
[0038] In one example, the present disclosure may also include software aspects to facilitate the collection of logic level metric counts for power estimation, etc. For example, the host 290 may program the accumulator module 250 with desired settings. These may include indicating how many cycles to accumulate data, conditions for storing or sending logic level metric count data, etc. The host 290 may also manage data generated during the workload / simulation run. For example, the host 290 may receive data and may store and / or process the data so as not to stall the simulation run.
[0039] In one example, host 290 can also control the operation of the emulator / emulation system. For example, in a scenario where a large amount of logic level metric count data is to be moved from PLD 210 to another location (e.g., to a data storage device), host 290 can then manage the starting and stopping of PLD 210's emulation clock to allow for the data offloading process. To further illustrate, host 290 can configure logic level metric count logic, for example, by adding logic to the circuit design for the purpose of logic level metric count collection. Additionally, host 290 can indicate the number of cycles over which logic level metric counts are to be accumulated. Additionally, host 290 can indicate the filtering and triggering settings for accumulator module 250.
[0040] During the simulation run, the host 290 may further enable logic level metric counting logic, such as one or more logic level metric counting modules 231 to 234 instantiated on the simulation hardware of the PLD 210. The host 290 may then instruct the workload to run the requested number of simulation cycles on the simulated design DUT 220. At run time, the host 290 may aggregate information from all sources in the simulation system 200 and may perform power calculations / estimations. For example, the host 290 may use information about the circuit design and the received logic level metric counting data to calculate power information and provide it to a user and / or to one or more other automated systems for further analysis. In one example, the host 290 may also generate and display power consumption visualizations, such as graphs of power consumption over time. It should also be noted that Figure 2 Only one example of a simulation system 200 configured to include logic level metric counting modules 231 - 234 is illustrated, and other, additional, and different examples of the present disclosure may provide simulation systems with more or fewer modules, with different elements, etc.
[0041] Figure 3 Two examples of aggregated logic level metric count modules (aggregation modules 310 and 350) according to the present disclosure are illustrated. For example, aggregation module 310 may include logic level metric count modules 321-323, which may store logic level metric counts of respective groups of nodes 311-313, for example, in respective incrementers 331-333 (e.g., incrementer circuits). For example, each of logic level metric count modules 321-323 may represent Figure 1 For ease of illustration, not all components within the logic level metric counting modules 321 to 323 are shown. Figure 3 , the aggregation module 310 may also include a buffer 340, which may include two counter memories 345 and 346. In this case, the buffer 340 may be referred to as a ping-pong buffer, and one of the counter memories 345-346 may be utilized as a real-time counter to accumulate logic level metric counts from the corresponding incrementer 331-333, while using the other of the counter memories 345-346 as a transmission buffer, for example, to offload to a management station / host.
[0042] For example, the incrementers 331 to 333 may be independently operated to count the logic level metrics (e.g., TC and or T1 counts) of all nodes 311 to 313 attached to the corresponding logic level metric counting modules 321 to 323 over a number of simulation cycles. The logic level counts may be periodically unloaded to a real-time one of the counter memories 345 to 346. The real-time one of the counter memories 345 to 346 may be read to retrieve the previous count value of each node and then written to store the updated current count value of each of the modules 321 to 323. Once a certain number of simulation cycles have passed (which may be determined automatically by the circuitry or programmed by the user), the real-time counter memory may become a transmit counter memory. This transmit counter memory may then be read and the value transmitted to the management station / host. For example, in one example, the aggregation module 310 may represent Figure 2 234 and accumulator 250. Thus, while one of the counter memories 345-346 described above is sending data to a management station for analysis, another of the counter memories 345-346 can become a real-time counter memory, and so on for additional blocks of emulation clock cycles. Thus, buffer 340 continuously counts one of the counter memories 345-346 and transmits the count value to the other of the counter memories 345-346, thereby maintaining high system throughput. In one example, buffer 340 can include a selection module 349 having a selection line via which the real-time and transmit counter memories can be changed.
[0043] exist Figure 3 In the second example illustrated in FIG, aggregation module 350 may include a single memory buffer implementation. For example, aggregation module 350 may include the same or similar components as aggregation module 310, such as logic level metric counter modules 361-363, including incrementers 371-373, and distributed to corresponding groups of nodes 351-353. As with the previous example, not all components within logic level metric counter modules 361-363 are shown. In this case, aggregation module 350 includes a buffer 380 with a single counter memory 385. Notably, this implementation provides a more compact module / circuit and a less expensive implementation because only one memory is used in buffer 380. However, aggregation module 350 may provide slightly reduced performance compared to aggregation module 310, as the simulation clock may be stopped in order to offload counter memory 385 to, for example, a management station. Once the data in memory 385 has been offloaded, the simulation system's clock can be restored.
[0044] To further illustrate, with respect to the aggregation module 310, the counter memories 345 and 346 may be two block random access memories (BRAMs) (e.g., having a size of 72x512). The aggregation module 310 may support approximately 1,500 nodes using three 512:1 multiplexers (e.g., one for each of the logic level metric counting modules 321-323) and 24 lookup table random access memories (LUTRAMs) (e.g., 1x64 LUTRAMs, 8 for each of the logic level metric counting modules 321-323) to implement a previous value memory (e.g., Figure 1 ). Additionally, the aggregation module 310 may use three 18-bit counters for the incrementers 331-333.
[0045] The aggregation module 350 may utilize a similar set of programmable logic device (PLD) building blocks, such as three 512:1 multiplexers (e.g., one for each of the logic level metric counting modules 361-363) and 24 lookup table random access memories (LUTRAMs) (e.g., 1x64 LUTRAMs, 8 for each of the logic level metric counting modules 361-363) to implement the previous value memory (e.g., Figure 1 ). Additionally, the aggregation module 350 may use three 18-bit counters for the incrementers 371 to 373. However, the aggregation module may use a single BRAM (e.g., having size 72x512) for the buffer 380.
[0046] It should also be noted that Figure 3Only two example aggregation modules are described, and other, additional, and different examples of the present disclosure may provide aggregation modules with more or fewer elements, with different elements, etc. For example, in an example with four incrementers and one ping-pong buffer, the aggregation module may include four 512:1 multiplexers capable of calculating the count of approximately 2,000 nodes within 512 cycles, two BRAMs (72x512) (e.g., for the ping-pong buffer), 32 LUTRAMs (4x8=32 LUTRAMs of size 1x64), and four 18-bit counters. In another example, using four incrementers and a single buffer, the aggregation module may include four 512:1 multiplexers capable of calculating the count of approximately 2,000 nodes within 512 cycles, one BRAM (72x512) (e.g., for the ping-pong buffer), 32 LUTRAMs (4x8=32 LUTRAMs of size 1x64), and four 18-bit counters. In another example, the aggregation module 310 may include multiple incrementers within each of the logic level metric counting modules 321 to 323, for example, an incrementer for the switch count and an incrementer for the T1 count. In such an example, the buffer 340 may include two additional counter memories, for example, two for the switch count aggregation and two for the T1 count aggregation. In another example, the aggregation module 350 may be similarly modified to handle both the switch count and the T1 count. In each case, the required BRAM will be doubled for the TC and T1 counts, but other resources can be shared. For example, for 16K nodes being monitored and a message sent once every 256K cycles, the TC can utilize 16 BRAMs (8 x 2 = 16 BRAMs of size 72 x 512). For the T1 count, another 16 BRAMs of the same size can be used. Therefore, these and other modifications are considered within the scope of the present disclosure.
[0047] Figure 4 An example simulation system 400 is illustrated that includes several programmable logic devices (PLDs) 410, 420, and 430 (e.g., FPGAs, etc.) configured to simulate a circuit design or design under test (DUT). For example, aspects of the DUT (e.g., DUT logic 412, DUT logic 422, and DUT logic 432) can be assigned to different ones of the PLDs 410, 420, and 430. The DUT logic 412, DUT logic 422, and DUT logic 432 can each include multiple nodes that can be monitored by corresponding logic level metric counting modules 414, 424, and 434. For example, the logic level metric counting module 414 can include one or more logic level metric counting modules, such as Figure 1100, etc. (and similarly for logic level metric counting modules 424 and 434). Alternatively, or in addition, logic level metric counting modules 414, 424, and 434 may each correspond to an aggregation module, such as Figure 3 , etc. The logic level metric count modules 414, 424, and 434 may collect logic level metric counts (which, in one example, may include aggregated logic level metric counts from multiple modules) and transmit them to the management station 490. Thus, Figure 4 As illustrated in the example of FIG4 , as the DUT logic 412, 422, and 432 are distributed to different PLDs 410, 420, and 430, the corresponding logic level metric counting modules 414, 424, and 434 may be similarly distributed therewith. The management station 490 may then calculate various additional metrics, such as design-wide TC and / or T1 counts, design-wide power consumption estimates, and the like. Alternatively, or in addition, the management station 490 may compare different aspects of the circuit design that may be separated into the DUT logic 412, 422, and 432, such as simultaneously identifying different power utilization estimates in different aspects of the circuit design for the same workload. As in the previous example, it should also be noted that Figure 4 Only one example of a distributed simulation system 400 including multiple PLDs 410 , 420 , and 430 is illustrated, and other, additional, and different examples of the present disclosure may provide distributed simulation systems with more or fewer PLDs, with different components, etc.
[0048] Figure 5 A flow chart illustrating an example method 500 for presenting at least one power utilization estimate for a circuit design based on at least one logic level metric count obtained from at least one logic level metric count module added to the circuit design and loaded into a simulation system. In one example, the method 500 can be performed by a computing device or system, such as a processing system or processing device, including at least one processor, a memory storing instructions that, when executed by the at least one processor, cause the processing system to perform operations, etc. For example, the method 500 can be performed by a processing system including at least one processor, such as Figure 7 host system 707 and / or compiler 710, or a host system 707 and / or compiler 710 in combination with the emulation system 702, Figure 8 The method 500 may be performed by a computer system 800, and / or any one or more components thereof, such as the processing device 802, or by multiple instances of the computer system 800 communicating over one or more networks and operating collectively to perform one or more aspects of the method 500. For illustrative purposes, the method 500 is described with reference to an example performed by a processing system. The method 500 begins at 505 and may proceed to 510.
[0049] The processing system may obtain a circuit design at 510. For example, the circuit design may include a register transfer level (RTL) design (eg, an RTL netlist).
[0050] At 520, the processing system may add at least one logic level metric counting module to the circuit design. For example, the at least one logic level metric counting module may include one or more modules for switching counting and / or T1 counting, such as Figure 1 and described above. For example, at least one logic level metric counting module can be configured to count at least one of a toggle count (TC) of a signal at one or more nodes of the circuit design or a count of instances of the signal at a logic level high (T1 count) at one or more nodes of the circuit design. Alternatively, or in addition, at least one logic level metric counting module can be an aggregation module, such as Figure 3 Aggregation module 310 and / or aggregation module 350, and / or a group of aggregation logic level metric count value modules, such as Figure 4 4, 520. In one example, 520 may include creating and connecting one or more new netlist artifacts corresponding to at least one logic level metric count module. For example, the new netlist artifact may include memory primitives for one or more incrementers, one or more previous value memories (e.g., where the logic level metric count module is used for toggle counting), one or more buffers, etc., as well as multiplexers, gates, etc. that may facilitate the operation of the at least one logic level metric count module.
[0051] At 530, the processing system may load the circuit design into a simulation system (e.g., including at least one logic level metric counting module). For example, the processing system may include the functionality of a simulation compiler to map an input circuit design (e.g., a register transfer level (RTL) design) onto a target simulator (e.g., a programmable logic device / system), e.g., such that the target implementation is optimized for simulation throughput and capacity utilization of the simulator's resources. For example, the processing system may convert a specification written in a description language representing a device under test (DUT) to generate data (e.g., binary data) and information for constructing a simulation system to simulate the DUT. Additionally, the processing system may convert, modify, reconfigure, add new functionality to the DUT, and / or control the timing of the DUT.
[0052] As described above, at least one logic level metric counting module may include a multiplexer having a plurality of inputs for a plurality of sampling registers associated with a plurality of nodes of the circuit design and a select input for selecting an input from the plurality of inputs to pass to an output of the multiplexer. Furthermore, at least one logic level metric counting module may include an incrementer unit coupled to the output of the multiplexer, wherein the incrementer unit increments a count during each polling clock cycle (e.g., according to a sampling / polling / counting clock of the at least one logic level metric counting module) during which the output of the multiplexer is a logic high value. For example, in such an example, the at least one logic level metric counting module may be configured to collect and report T1 counts. In one example, each sampling register may store a value for a corresponding node of the circuit design during each of a plurality of simulation clock cycles of the simulation workload.
[0053] In one example, at least one logic level metric counting module may alternatively or additionally include a multiplexer having a plurality of inputs of a plurality of sampling registers associated with a plurality of nodes of the circuit design, and a select line for selecting an input from the plurality of inputs to be passed to an output of the at least one multiplexer within a polling clock cycle. At least one logic level metric counting module may alternatively or additionally include a previous value memory coupled to the output of the multiplexer to store the value of the output of the multiplexer from the current simulation clock cycle until the next simulation clock cycle in the plurality of simulation cycles of the simulation workload. At least one logic level metric counting module may also include a toggle detector having a first input for the output of the multiplexer and a second input for the output of the previous value memory, wherein the toggle detector outputs a logic high value when a first value at the output of the multiplexer is different from a second value at the output of the previous value memory. For example, in one example, the toggle detector may include an edge detector circuit. Furthermore, at least one logic level metric counting module may further include an incrementer unit coupled to the output of the at least one toggle detector. For example, the incrementer unit may increment a count within each polling clock cycle in which the output of the toggle detector is a logic high value. In one example, the previous value memory may include a number of storage bits at least as large as the number of the multiplexer's inputs. In one example, the previous value memory may include the same address selection input as the multiplexer's select line. As in the previous example, each of the plurality of sampling registers may store the value of a corresponding node of the plurality of nodes within each of a plurality of simulation clock cycles of the simulation workload.
[0054] The processing system may obtain a simulated workload at 540. For example, the simulated workload may include one or more programs, applications, and / or data processes, or a data set that simulates one or more programs, applications, and / or data processes.
[0055] At 550, the processing system may apply a simulation workload to a simulation system to which the circuit design is loaded, e.g., including at least one logic level metric counting module. For example, the simulation workload may be applied via input-output pins / ports of the simulation system (e.g., one or more programmable logic devices (PLDs) thereof). Alternatively, or in addition, at least a portion of the simulation workload may be preloaded into one or more memories of the one or more PLDs, for example, wherein the workload may be externally activated via one or more instructions, and wherein at least a portion of the workload data may be retrieved from the one or more memories in response to the one or more instructions.
[0056] At 560 , the processing system may obtain an indication of interest in a time slice that includes at least a portion of the workload.
[0057] At 570, the processing system may obtain at least one logic level metric count from at least one logic level metric count module, the at least one logic level metric count associated with at least a portion of the simulation workload. For example, in one example, the at least one logic level metric count may include one or both of: a toggle count of a signal at one or more nodes of the circuit design or a count of instances of the signal at a logic high level at one or more nodes of the circuit design. In one example, obtaining the at least one logic level metric count from the at least one logic level metric count module may include unloading the at least one logic level metric count in response to obtaining the indication of interest in the time slice at 560.
[0058] The processing system may calculate at least one power utilization estimate based on the at least one logic level metric count at 580. For example, as described above, the processing system may calculate a switching rate, a probability of a logic high value, an average mode power, and the like for at least a portion of the simulated workload (e.g., for a pre-specified time slice and / or for a time slice marked as of interest to an operator or another automated system), and so on.
[0059] At 590, the processing system may present at least one power utilization estimate for the circuit design, wherein the at least one power utilization estimate for the circuit design is based on at least one logic level metric count. For example, the at least one power utilization estimate may include a toggle rate, a probability of a logic high value, an average mode power, etc., as described above. Alternatively, or in addition, the at least one power utilization estimate may include the at least one logic level metric count itself. After 590, method 500 proceeds to 595, where method 500 ends.
[0060] It should be noted that method 500 can be expanded to include additional steps or modified to replace steps with different steps, combine steps, omit steps, perform steps in a different order, etc. For example, in one example, the processing system can repeat one or more steps of method 500, such as steps 510 through 590 for one or more updates to the circuit design and / or for one or more new, additional, and / or different circuit designs, steps 540 through 590 for one or more additional workloads, etc. In one example, step 520 can be a portion or sub-operation of step 530. In one example, for example, method 500 can further include combining count data from multiple logic level metric counting modules into a buffer (e.g., a ping-pong buffer) and / or offloading multiple logic level metric counting modules and / or multiple buffers (from different FPGAs / PLDs) to a workstation. In one example, aspects of method 500 can be performed by a processing system according to instructions from a non-transitory computer-readable medium storing such instructions, as described below. Additionally, in one example, the programmable logic device associated with method 500 may include a programmable logic device such as described in more detail below. In one example, method 500 may be expanded or modified to include steps, functions, and / or operations, or to combine Figures 1 to 4 or the example descriptions of 6 to 8, or other features as described elsewhere herein. Therefore, these and other modifications are considered within the scope of the present disclosure.
[0061] Thus, in one example, the present disclosure may include a processing device that may add at least one logic level metric count module to a circuit design, load the circuit design into a simulation system, and apply a simulation workload to the simulation system into which the circuit design is loaded. The processing device may then obtain at least one logic level metric count from the at least one logic level metric count module, the at least one logic level metric count associated with at least a portion of the simulation workload, and present at least one power utilization estimate for the circuit design, wherein the at least one power utilization estimate for the circuit design is based on the at least one logic level metric count.
[0062] In one example, the present disclosure may include a programmable logic device (PLD) having a multiplexer, the multiplexer including multiple inputs of a plurality of sampling registers associated with multiple nodes of a circuit design (e.g., compiled and loaded onto the PLD), and a select line for selecting an input from the plurality of inputs to be passed to an output of the multiplexer. The PLD may include a previous value memory coupled to the output of the multiplexer to store the value of the output of the multiplexer from a current simulation clock cycle until the next simulation clock cycle in a plurality of simulation cycles of a simulation workload applied to the programmable logic device. The PLD may also include a toggle detector having a first input for the output of the multiplexer and a second input for the output of the previous value memory, wherein the toggle detector is configured to output a logic high value when the first value at the output of the multiplexer is different from the second value at the output of the previous value memory. The PLD may also include an incrementer unit coupled to the output of the toggle detector, wherein the incrementer unit is configured to increment a count within each polling clock cycle in which the output of the toggle detector is a logic high value. In one example, a programmable logic device (PLD) may be configured to include a multiplexer, a previous value memory, a toggle detector, and an incrementer unit. For example, a circuit design may be modified to include design aspects of the multiplexer, the previous value memory, the toggle detector, and the incrementer unit. Then, based on a compilation of the modified circuit design, the PLD may be configured to include the multiplexer, the previous value memory, the toggle detector, and the incrementer unit.
[0063] In one example, the previous value memory may include a number of storage bits at least as large as the number of the multiplexer's inputs. In one example, the previous value memory may include an address select input that is the same as the multiplexer's select input. In one example, the toggle detector may include an edge detector circuit. In one example, each sampling register in the plurality of sampling registers stores the value of a corresponding node in the plurality of nodes within each of a plurality of simulation clock cycles of the simulation workload. Furthermore, in one example, the multiplexer, the previous value memory, the toggle detector, and the incrementer unit may be a first toggle counter circuit, wherein the PLD may include a plurality of toggle counter circuits including the first toggle counter circuit. In such an example, the PLD may further include a buffer coupled to the plurality of toggle counter circuits and configured to aggregate counts from a plurality of incrementer units in the plurality of toggle counter circuits.
[0064] In another example, the present disclosure may alternatively or additionally include a non-transitory computer-readable medium storing a circuit design that, when loaded onto at least one programmable logic device, configures the at least one programmable logic device to include a multiplexer having a plurality of sampling register inputs associated with a plurality of nodes of the circuit design, a select line for selecting an input from the plurality of inputs to pass to an output of the multiplexer, and an incrementer unit coupled to the output of the multiplexer, the incrementer unit configured to increment a count during each polling clock cycle in which the output of the multiplexer is a logic high value. In one example, each sampling register of the plurality of sampling registers may store a value of a corresponding node of the plurality of nodes during each of a plurality of simulation clock cycles of a simulated workload. In one example, when loaded onto the at least one programmable logic device, the stored circuit design further configures the at least one programmable logic device to include a plurality of logic-high count circuits / modules, and a buffer coupled to the plurality of logic-high count circuits and configured to aggregate counts from the plurality of incrementer units of the plurality of logic-high count circuits. In such an example, the multiplexer and incrementer unit can be a first logic high count circuit in a plurality of logic high count circuits. In one example, the aforementioned programmable logic device and non-transitory computer readable medium can be used with Figure 5 Any one or more aspects of the example method 500 may be used in combination.
[0065] Figure 6 An example set of processes 600 used during the design, verification, and fabrication of an article of manufacture, such as an integrated circuit, to transform and verify design data and instructions representing the integrated circuit is illustrated. Each of these processes can be structured and enabled as multiple modules or operations. The term 'EDA' stands for the term 'electronic design automation'. These processes begin with the creation of a product concept 610 using information supplied by a designer, which is transformed to create an article of manufacture using a set of EDA processes 612. When the design is complete, the design is taped out 634, which is when the artwork (e.g., geometric pattern) of the integrated circuit is sent to a fabrication facility to produce a mask set, which is then used to manufacture the integrated circuit. After tapeout, semiconductor die are fabricated 636, and packaging and assembly processes 638 are performed to produce a finished integrated circuit 640.
[0066] The specification of a circuit or electronic structure can range from low-level transistor material layout to a high-level description language. High-level representations can be used to design circuits and systems using hardware description languages ('HDL') such as VHDL, Verilog, SystemVerilog, SystemC, MyHDL or OpenVera. The HDL description can be converted to a logic level register transfer level ('RTL') description, a gate level description, a layout level description or a mask level description. Each lower level of representation that is a more detailed description adds more useful detail to the design description, for example, more detail of the modules that contain the description. The lower level of representation that is a more detailed description can be computer generated, derived from a design library or created by another design automation process. An example of a lower level specification language for specifying a more detailed description representation language is SPICE, which is used for detailed descriptions of circuits with many analog components. The description of each level of representation is made usable by the corresponding system for that level (for example, a formal verification system). The design process can use Figure 6 The sequence depicted in . The described process can be enabled by an EDA product (or EDA system).
[0067] During system design 614, the functionality of the integrated circuit to be manufactured is specified. The design may be optimized for desired characteristics such as power consumption, performance, area (physical and / or lines of code), and cost reduction. At this stage, the design may be divided into different types of modules or components.
[0068] During logic design and functional verification 616, modules or components within a circuit are specified using one or more descriptive languages and checked for functional accuracy. For example, components within a circuit may be verified to generate outputs that match the specifications of the designed circuit or system. Functional verification may utilize simulators and other programs, such as test bench generators, static HDL checkers, and formal checkers. In some embodiments, specialized systems, referred to as "simulators" or "prototyping systems," are used to accelerate functional verification.
[0069] During synthesis and test design 618, the HDL code is converted into a netlist. In some embodiments, the netlist can be a graph structure, where the edges of the graph structure represent components of the circuit, and where the nodes of the graph structure represent how the components are interconnected. Both the HDL code and the netlist are hierarchical artifacts that can be used by EDA products to verify that the integrated circuit performs according to the specified design when manufactured. The netlist can be optimized for the target semiconductor manufacturing technology. Additionally, the finished integrated circuit can be tested to verify that the integrated circuit meets the requirements of the specification.
[0070] During netlist verification 620, the netlist is checked for compliance with timing constraints and for correspondence with the HDL code. During design planning 622, the overall floor plan of the integrated circuit is constructed and analyzed for timing and top-level routing.
[0071] During layout or physical implementation 624, physical layout (positioning of circuit components such as transistors or capacitors) and routing (connecting circuit components through multiple conductors) are performed, and selection of cells from a library to implement a specific logic function may be performed. As used herein, the term 'cell' may specify a group of transistors, other components, and interconnects that provide a Boolean logic function (e.g., AND, OR, NOT, XOR) or a storage function (e.g., a flip-flop or latch). As used herein, a circuit 'block' may refer to two or more cells. Both cells and circuit blocks may be referred to as modules or components and may be enabled as both physical structures and simulations. Parameters, such as size, are specified for the selected cell (based on a 'standard cell') and made accessible in a database for use by the EDA product.
[0072] During analysis and extraction 626, circuit functionality is verified at the layout level, which allows for refinement of the layout design. During physical verification 628, the layout design is checked to ensure that manufacturing constraints, such as DRC constraints, electrical constraints, and lithography constraints, are correct and that the circuit functionality matches the HDL design specifications. During resolution enhancement 630, the layout geometry is transformed to improve the manufacturability of the circuit design.
[0073] During tape-out, data is created for the production of lithographic masks (after applying lithographic enhancements if appropriate).During mask data preparation 632, the 'tape-out' data is used to generate lithographic masks for the production of finished integrated circuits.
[0074] Computer systems (e.g. Figure 8 The storage subsystem of the computer system 800) may be used to store programs and data structures used by some or all of the EDA products described herein, as well as products for developing cells for the library and for physical and logical designs that use the library.
[0075] Figure 7 A diagram depicts an example simulation environment 700. Simulation environment 700 can be configured to verify the functionality of a circuit design. Simulation environment 700 can include a host system 707 (e.g., a computer as part of an EDA system) and a simulation system 702 (e.g., a set of programmable devices such as a field programmable gate array (FPGA) or a processor). The host system generates data and information by using a compiler 710 to construct the simulation system to simulate the circuit design. The circuit design to be simulated is also referred to as a design under test ('DUT'), where data and information from the simulation are used to verify the functionality of the DUT.
[0076] The host system 707 may include one or more processors. In embodiments where the host system includes multiple processors, the functions described herein as being performed by the host system may be distributed among the multiple processors. The host system 707 may include a compiler 710 to convert a specification written in a description language representing a DUT and generate data (e.g., binary data) and information used to construct the simulation system 702 to simulate the DUT. The compiler 710 may convert, change, reconfigure, add new functionality to the DUT, and / or control the timing of the DUT.
[0077] Host system 707 and emulation system 702 exchange data and information using signals carried by the emulation connection. The connection may be, but is not limited to, one or more cables, such as a cable having a pin configuration compatible with the Recommended Standard 232 (RS232) or Universal Serial Bus (USB) protocols. The connection may be a wired communication medium or a network, such as a local area network or a wide area network (e.g., the Internet). The connection may be a wireless communication medium or a network having one or more access points using a wireless protocol (e.g., Bluetooth or IEEE 802.11). Host system 707 and emulation system 702 may exchange data and information through a third device, such as a network server.
[0078] The simulation system 702 includes multiple FPGAs (or other modules), such as FPGAs 7041 and 7042 and up to 704 N Each FPGA may include one or more FPGA interfaces through which the FPGA connects to other FPGAs (and potentially other simulation components) for the FPGAs to exchange signals. FPGA interfaces may be referred to as input / output pins or FPGA pads. Although an emulator may include an FPGA, embodiments of the emulator may include other types of logic blocks for emulating the DUT in place of or in addition to the FPGA. For example, the emulation system 702 may include a custom FPGA, a dedicated ASIC for emulation or prototyping, memory, and input / output devices.
[0079] A programmable device may include an array of programmable logic blocks and a hierarchy of interconnects that enable the programmable logic blocks to interconnect according to descriptions in HDL code. Each of the programmable logic blocks can enable complex combinational functions or logic gates, such as AND and XOR logic blocks. In some embodiments, the logic blocks may also include memory elements / devices, which may be simple latches, flip-flops, or other memory blocks. Depending on the length of the interconnects between different logic blocks, signals may arrive at the input terminals of the logic blocks at different times and may therefore be temporarily stored in the memory elements / devices.
[0080] FPGA 7041 to 704 N Can be placed on one or more plates 7121 and 7122 and up to 712M Multiple boards can be placed into the simulation unit 7141. Boards within the simulation unit can be connected using the simulation unit's backplane or any other type of connection. In addition, multiple simulation units (e.g., 7141 to 714 K ) can be connected to each other by cables or any other means to form a multi-simulation unit system.
[0081] For the DUT to be simulated, the host system 707 transfers one or more bitfiles to the simulation system 702. The bitfiles may specify a description of the DUT and may further specify the partitions of the DUT with trace and injection logic created by the host system 707, the mapping of the partitions to the simulator's FPGAs, and design constraints. Using the bitfiles, the simulator configures the FPGAs to perform the functions of the DUT. In some embodiments, one or more FPGAs of the simulator may have trace and injection logic built into the FPGA's silicon. In such embodiments, the FPGAs may not be configured by the host system to simulate the trace and injection logic.
[0082] The host system 707 receives a description of the DUT to be simulated. In some embodiments, the DUT description is in a description language (e.g., register transfer language (RTL)). In some embodiments, the DUT description is in a netlist-level file or a mixture of netlist-level files and HDL files. If part or all of the DUT description is in HDL, the host system may synthesize the DUT description to create a gate-level netlist using the DUT description. The host system may use the DUT netlist to partition the DUT into a plurality of partitions, one or more of which include trace and injection logic. The trace and injection logic traces interface signals exchanged through the interface of the FPGA. In addition, the trace and injection logic may inject the traced interface signals into the logic of the FPGA. The host system maps each partition to the FPGA of the simulator. In some embodiments, the trace and injection logic is included in a set of selected partitions of the FPGA. The trace and injection logic may be built into one or more of the FPGAs of the simulator. The host system may synthesize a multiplexer to map into the FPGA. The multiplexer may be used by the trace and injection logic to inject the interface signals into the DUT logic.
[0083] The host system creates a bitfile that describes each partition of the DUT and the mapping of the partitions to the FPGAs. For partitions that contain trace and injection logic, the bitfile also describes the included logic. The bitfile can include placement and routing information as well as design constraints. The host system stores the bitfile along with information describing which FPGAs will emulate each component of the DUT (e.g., which FPGAs each component maps to).
[0084] Upon request, the host system transfers the bitfile to the emulator. The host system signals the emulator to begin emulation of the DUT. During or at the end of the DUT simulation, the host system receives simulation results from the emulator via the emulation connection. The simulation results are data and information generated by the emulator during the DUT simulation, including interface signals and their states that have been traced by each FPGA's trace and injection logic. The host system can store the simulation results and / or transmit them to another processing system.
[0085] After simulating a DUT, the circuit designer may request to debug a component of the DUT. If such a request is made, the circuit designer may specify a time period for the simulation to be debugged. The host system uses the stored information to identify which FPGAs are simulating the component. The host system retrieves the stored interface signals associated with the time period and tracked by the trace and injection logic of each identified FPGA. The host system signals the emulator to re-simulate the identified FPGA. The host system transmits the retrieved interface signals to the emulator to re-simulate the component within the specified time period. The trace and injection logic of each identified FPGA injects its corresponding interface signals received from the host system into the logic of the DUT mapped to the FPGA. In the case of multiple re-simulations of the FPGA, the results are merged to produce a complete debug view.
[0086] The host system receives, from the simulation system, signals traced by the logic of the identified FPGA during the resimulation of the component. The host system stores the signals received from the simulator. The signals traced during the resimulation may have a higher sampling rate than the sampling rate during the initial simulation. For example, in the initial simulation, the traced signals may include the saved state of the component every X milliseconds. However, during the resimulation, the traced signals may include the saved state every Y milliseconds, where Y is less than X. If the circuit designer requests to view the waveform of the signal traced during the resimulation, the host system can retrieve the stored signal and display a plot of the signal. For example, the host system can generate a waveform of the signal. The circuit designer can then request to resimulate the same component or resimulate another component during a different time period.
[0087] The host system 707 and / or the compiler 710 may include subsystems such as, but not limited to, a design synthesizer subsystem, a mapping subsystem, a runtime subsystem, a result subsystem, a debug subsystem, a waveform subsystem, and a storage subsystem. These subsystems may be configured and enabled individually or as multiple modules, or two or more subsystems may be configured as modules. Together, these subsystems form a simulator and monitor simulation results.
[0088] The design synthesizer subsystem converts the HDL representing the DUT 705 into gate-level logic. For a DUT to be simulated, the design synthesizer subsystem receives a description of the DUT. If the description of the DUT is in whole or in part in HDL (e.g., RTL or other level of representation), the design synthesizer subsystem synthesizes the HDL of the DUT to create a gate-level netlist having a description of the gate-level logic of the DUT.
[0089] The mapping subsystem partitions the DUT and maps the partitions into the simulator FPGA. The mapping subsystem uses the DUT's netlist to partition the DUT into partitions at the gate level. For each partition, the mapping subsystem retrieves a gate-level description of trace and injection logic and adds the logic to the partition. As described above, the trace and injection logic contained in the partition is used to trace signals exchanged through the interface of the FPGA to which the partition is mapped (tracing interface signals). The trace and injection logic can be added to the DUT before partitioning. For example, the trace and injection logic can be added by the design synthesizer subsystem before or after synthesizing the DUT's HDL.
[0090] In addition to including trace and injection logic, the mapping subsystem can also include additional trace logic in the partitions to track the state of certain DUT components that are not tracked by trace and injection. The mapping subsystem can include additional trace logic in the DUT before partitioning or in the partitions after partitioning. The design synthesizer subsystem can include additional trace logic in the HDL description of the DUT before synthesizing the HDL description.
[0091] The mapping subsystem maps each partition of the DUT to the FPGA of the simulator. For partitioning and mapping, the mapping subsystem uses design rules, design constraints (e.g., timing or logic constraints), and information about the simulator. For components of the DUT, the mapping subsystem stores information in the storage subsystem that describes which FPGA will simulate each component.
[0092] Using the partitions and mappings, the mapping subsystem generates one or more bitfiles that describe the created partitions and the mapping of the logic to each FPGA in the simulator. The bitfiles may contain additional information, such as constraints for the DUT and routing information for connections between and within each FPGA. The mapping subsystem may generate a bitfile for each partition of the DUT and may store the bitfiles in the storage subsystem. Upon request from the circuit designer, the mapping subsystem transfers the bitfiles to the simulator, which may use the bitfiles to configure the FPGAs to simulate the DUT.
[0093] If the emulator includes a dedicated ASIC that includes trace and injection logic, the mapping subsystem can generate a specific structure that connects the dedicated ASIC to the DUT. In some embodiments, the mapping subsystem can save information about the traced / injected signals and where the information is stored on the dedicated ASIC.
[0094] The runtime subsystem controls the simulations performed by the simulator. The runtime subsystem can start or stop the simulator from executing a simulation. In addition, the runtime subsystem can provide input signals and data to the simulator. The input signals can be provided directly to the simulator via a connection, or indirectly to the simulator via other input signal devices. For example, the host system can control the input signal device to provide the input signals to the simulator. The input signal device can be, for example, a test board (directly or via a cable), a signal generator, another simulator, or another host system.
[0095] The results subsystem processes simulation results generated by the simulator. During simulation and / or after the simulation is complete, the results subsystem receives simulation results generated during simulation from the simulator. The simulation results include signals tracked during simulation. Specifically, the simulation results include interface signals tracked by the trace and injection logic simulated by each FPGA, and may include signals tracked by additional logic included in the DUT. Each tracked signal may span multiple cycles of the simulation. The tracked signal includes multiple states, and each state is associated with the time of the simulation. The results subsystem stores the tracked signals in the storage subsystem. For each stored signal, the results subsystem may store information indicating which FPGA generated the tracked signal.
[0096] The debug subsystem allows circuit designers to debug DUT components. After the simulator has simulated the DUT and the results subsystem has received the interface signals traced by the trace and injection logic during simulation, the circuit designer can request to debug the DUT component by resimulating the component for a specific time period. In the request to debug a component, the circuit designer identifies the component and indicates the simulation time period to be debugged. The circuit designer's request can include a sampling rate, which indicates how often the state of the debugged component should be saved by the logic tracing the signals.
[0097] The debug subsystem uses the information stored in the storage subsystem by the mapping subsystem to identify one or more FPGAs of the emulator that are emulating the component. For each identified FPGA, the debug subsystem retrieves from the storage subsystem the interface signals traced by the FPGA's trace and injection logic during a time period specified by the circuit designer. For example, the debug subsystem retrieves the state traced by the trace and injection logic associated with the time period.
[0098] The debug subsystem transmits the retrieved interface signals to the emulator. The debug subsystem instructs the debug subsystem to use the identified FPGAs and, for each identified FPGA's trace and injection logic, inject its corresponding traced signal into the FPGA's logic to re-simulate the component within the requested time period. The debug subsystem may further transmit the sampling rate provided by the circuit designer to the emulator so that the trace logic traces the state at appropriate intervals.
[0099] To debug a component, the simulator can use the FPGA that the component has been mapped to. Furthermore, a re-simulation of the component can be performed at any point in time specified by the circuit designer.
[0100] For the identified FPGA, the debug subsystem can transmit instructions to the emulator to load multiple emulator FPGAs with the same configuration as the identified FPGA. The debug subsystem also signals the emulator to use the multiple FPGAs in parallel. Each FPGA from the multiple FPGAs is used with a different time window of the interface signal to generate a larger time window in a shorter amount of time. For example, the identified FPGA may take an hour or more to use a certain amount of loops. However, if multiple FPGAs have the same data and structure as the identified FPGA, and each of these FPGAs runs a subset of the loops, then the emulator may take several minutes for the FPGAs to use all the loops together.
[0101] The circuit designer can identify a hierarchy or list of DUT signals to be re-simulated. To accomplish this, the debug subsystem determines the FPGA required to simulate the hierarchy or list of signals, retrieves the necessary interface signals, and transfers the retrieved interface signals to the simulator for re-simulation. Thus, the circuit designer can identify any element (e.g., component, device, or signal) of the DUT to be debugged / re-simulated.
[0102] The waveform subsystem generates waveforms using the traced signals. If the circuit designer requests to view the waveform of a signal traced during a simulation run, the host system retrieves the signal from the storage subsystem. The waveform subsystem displays a plot of the signal. For one or more signals, the waveform subsystem can automatically generate a plot of the signal when the signal is received from the simulator.
[0103] Figure 8An example computer system 800 is illustrated within which a set of instructions may be executed, causing the machine to perform any one or more of the methodologies discussed herein. In alternative embodiments, the machine may be connected (e.g., using a network) to other machines. The machine may operate in the capacity of a server or a client user machine in server-client user network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or a client user machine in a cloud computing infrastructure or environment.
[0104] The machine may be a personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), cellular phone, network appliance, server, network router, switch or bridge, or any machine capable of executing (sequentially or otherwise) a set of instructions that specify actions to be taken by the machine. Further, while a single machine is illustrated, the term "machine" shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.
[0105] The example computer system 800 includes a processing device 802, a main memory 804 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM)), a static memory 806 (e.g., flash memory, static random access memory (SRAM), etc.), and a data storage device 818, which communicate with each other via a bus 830.
[0106] Processing device 802 represents one or more processors, such as microprocessors, central processing units, or the like. More specifically, the processing device may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor that implements other instruction sets or multiple processors that implement a combination of instruction sets. Processing device 802 may also be one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), a network processor, or the like. Processing device 802 may be configured to execute instructions 826 for performing the operations and steps described herein.
[0107] The computer system 800 may further include a network interface device 808 for communicating over a network 820. The computer system 800 may also include a video display unit 810 (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)), an alphanumeric input device 812 (e.g., a keyboard), a cursor control device 814 (e.g., a mouse), a graphics processing unit 822, a signal generating device 816 (e.g., a speaker), a video processing unit 828, and an audio processing unit 832.
[0108] The data storage device 818 may include a machine-readable storage medium 824 (also referred to as a non-transitory computer-readable medium) on which is stored one or more sets of instructions 826 or software embodying any one or more of the methodologies or functions described herein. During execution of the instructions 826 by the computer system 800, the instructions 826 may also reside, completely or at least partially, within the main memory 804 and / or within the processing device 802, with the main memory 804 and the processing device 802 also constituting machine-readable storage media.
[0109] In some embodiments, the instructions 826 include instructions that implement functionality corresponding to the present disclosure. Although the machine-readable storage medium 824 is shown as a single medium in the example embodiment, the term "machine-readable storage medium" should be considered to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more sets of instructions. The term "machine-readable storage medium" should also be considered to include any medium that can store or encode a set of instructions for execution by a machine and cause the machine and processing device 802 to perform any one or more of the methods of the present disclosure. Therefore, the term "machine-readable storage medium" should be considered to include, but not be limited to, solid-state memory, optical media, and magnetic media.
[0110] Some portions of the foregoing detailed description have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the art of data processing to most effectively convey the substance of their work to others skilled in the art. An algorithm is a sequence of operations leading to a desired result. The operations are those requiring physical manipulation of physical quantities. These quantities may take the form of electrical or magnetic signals capable of being stored, combined, compared, and otherwise manipulated. Such signals may be referred to as bits, values, elements, symbols, characters, terms, numbers, or the like.
[0111] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise apparent from this disclosure, it should be understood that throughout the description certain terms refer to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities within the computer system's registers and memories into other data similarly represented as physical quantities within the computer system's memories or registers or other such information storage devices.
[0112] The present disclosure also relates to an apparatus for performing the operations described herein. This apparatus may be specially constructed for the intended purpose, or it may comprise a computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a computer-readable storage medium, such as, but not limited to, any type of disk including floppy disks, optical disks, CD-ROMs, and magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.
[0113] The algorithms and displays presented herein are not inherently related to any particular computer or other device. Various other systems may be used in conjunction with the programs taught herein, or it may prove convenient to construct more specialized equipment to perform the methods. Furthermore, the present disclosure is not described with reference to any particular programming language. It will be appreciated that various programming languages may be used to implement the teachings of the present disclosure as described herein.
[0114] The present disclosure may be provided as a computer program product or software, which may include a machine-readable medium having instructions stored thereon, which instructions can be used to program a computer system (or other electronic device) to perform a process according to the present disclosure. A machine-readable medium includes any mechanism for storing information in a form that can be read by a machine (e.g., a computer). For example, a machine-readable (e.g., computer-readable) medium includes a machine (e.g., computer) readable storage medium, such as a read-only memory ("ROM"), a random access memory ("RAM"), a magnetic disk storage medium, an optical storage medium, a flash memory device, etc.
[0115] In the foregoing disclosure, embodiments of the present disclosure have been described with reference to specific example embodiments of the present disclosure. Obviously, various modifications may be made thereto without departing from the broader spirit and scope of the present disclosure as set forth in the claims below. Where the present disclosure refers to some elements in the singular, more than one element may be depicted in a figure, and the same elements may be designated by the same numerals. Accordingly, the present disclosure and the accompanying drawings should be regarded in an illustrative rather than a restrictive sense.
Claims
1. A method comprising: adding at least one logic level metric counting module to the circuit design; loading the circuit design into a simulation system via a processing device; applying, by the processing device, a simulation workload to the simulation system to which the circuit design is loaded; obtaining at least one logic level metric count from the at least one logic level metric count module, the at least one logic level metric count associated with at least a portion of the simulation workload; and At least one power utilization estimate of the circuit design is presented, wherein the at least one power utilization estimate of the circuit design is based on the at least one logic level metric count.
2. The method of claim 1 , wherein the at least one logic level metric count comprises: a switching count at one or more nodes of the circuit design; or A count of instances of the signal at the one or more nodes of the circuit design being at a logic level high.
3. The method of claim 1 , wherein the at least one logic level metric counting module comprises: A multiplexer with: a plurality of inputs of a plurality of sampling registers associated with a plurality of nodes of the circuit design; and a select input for selecting an input of the plurality of inputs to pass to an output of the multiplexer; and An incrementer unit is coupled to the output of the multiplexer, wherein the incrementer unit increments a count during each polling clock cycle in which the output of the multiplexer is a logic high value. 4 . The method of claim 3 , wherein each sampling register of the plurality of sampling registers stores a value of a corresponding node of the plurality of nodes within each of a plurality of simulation clock cycles of the simulation workload.
5. The method of claim 1 , wherein the at least one logic level metric counting module comprises: A multiplexer with: a plurality of inputs of a plurality of sampling registers associated with a plurality of nodes of the circuit design; and a select line for selecting an input of the plurality of inputs to pass to an output of the multiplexer; a previous value memory coupled to the output of the multiplexer to store a value of the output of the multiplexer from a current simulation clock cycle until a next simulation clock cycle of a plurality of simulation cycles of the simulation workload; a toggle detector having a first input for the output of the multiplexer and a second input for the output of the previous value memory, wherein the toggle detector outputs a logic high value when a first value at the output of the multiplexer is different from a second value at the output of the previous value memory; and An incrementer unit is coupled to the output of the toggle detector, wherein the incrementer unit increments a count during each polling clock cycle in which the output of the toggle detector is a logic high value.
6. The method of claim 5, wherein the previous value memory includes a number of storage bits at least as large as the number of the plurality of inputs to the multiplexer, wherein the previous value memory includes the same address select input as the select line of the multiplexer. The method of claim 5 , wherein the switching detector comprises an edge detector circuit.
8. The method of claim 5, wherein each sampling register of the plurality of sampling registers stores a value of a corresponding node of the plurality of nodes within each simulation clock cycle of the plurality of simulation clock cycles of the simulation workload.
9. The method of claim 1 , wherein the at least one power utilization estimate comprises: said at least one logic level metric count; Switching rate; Probability of logical high value; or Average mode power.
10. A programmable logic device, comprising: A multiplexer with: a plurality of inputs of a plurality of sampling registers associated with a plurality of nodes of the circuit design; and a select line for selecting an input of the plurality of inputs to pass to an output of the multiplexer; a previous value memory coupled to the output of the multiplexer to store a value of the output of the multiplexer from a current simulation clock cycle until a next simulation clock cycle in a plurality of simulation cycles of a simulation workload applied to the programmable logic device; a toggle detector having a first input for the output of the multiplexer and a second input for the output of the previous value memory, wherein the toggle detector is configured to output a logic high value when a first value at the output of the multiplexer is different from a second value at the output of the previous value memory; and An incrementer unit is coupled to the output of the toggle detector, wherein the incrementer unit is configured to increase a count within each polling clock cycle in which the output of the toggle detector is a logic high value. 11 . The programmable logic device of claim 10 , wherein the programmable logic device is configured to include the multiplexer, the previous value memory, the switching detector, and the incrementer unit.
12. The programmable logic device of claim 10 , wherein the circuit design is modified to include design aspects of the multiplexer, the previous value memory, the switch detector, and the incrementer unit, and wherein the programmable logic device is configured to include the multiplexer, the previous value memory, the switch detector, and the incrementer unit according to a compilation of the modified circuit design.
13. The programmable logic device of claim 10, wherein the previous value memory comprises a number of storage bits at least as large as the number of the plurality of inputs of the multiplexer.
14. The programmable logic device of claim 13, wherein the previous value memory includes an address select input that is the same as a select input of the multiplexer.
15. The programmable logic device of claim 10, wherein the switching detector comprises an edge detector circuit. 16 . The programmable logic device of claim 10 , wherein each sampling register of the plurality of sampling registers stores a value of a corresponding node of the plurality of nodes within each emulated clock cycle of the plurality of emulated clock cycles of the emulated workload.
17. The programmable logic device of claim 10 , wherein the multiplexer, the previous value memory, the toggle detector, and the incrementer unit comprise a first toggle count circuit, wherein the programmable logic device comprises a plurality of toggle count circuits including the first toggle count circuit, and wherein the programmable logic device further comprises a buffer coupled to the plurality of toggle count circuits and configured to aggregate counts from a plurality of incrementer units of the plurality of toggle count circuits.
18. A non-transitory computer-readable medium comprising a stored circuit design that, when loaded onto at least one programmable logic device, configures the at least one programmable logic device to include: A multiplexer with: a plurality of inputs of a plurality of sampling registers associated with a plurality of nodes of the stored circuit design; and a select line for selecting an input of the plurality of inputs to pass to an output of the multiplexer; and An incrementer unit is coupled to the output of the multiplexer, wherein the incrementer unit is configured to increase a count within each polling clock cycle in which the output of the multiplexer is a logic high value.
19. The non-transitory computer-readable medium of claim 18, wherein each sampling register of the plurality of sampling registers stores a value of a corresponding node of the plurality of nodes within each of a plurality of simulation clock cycles of a simulation workload.
20. The non-transitory computer-readable medium of claim 18, and wherein when loaded onto the at least one programmable logic device, the stored circuit design further configures the at least one programmable logic device to include: a plurality of logic high count circuits, wherein the multiplexer and the incrementer unit comprise a first logic high count circuit of the plurality of logic high count circuits; and A buffer, coupled to the plurality of logic high count circuits, is configured to aggregate counts from a plurality of incrementer cells of the plurality of logic high count circuits.