Design under test (DUT) processing for logic optimization
By splitting the representation of the DUT into smaller subpartitions and adopting parallel multi-threading, combined with the insertion anchor circuit instance to retain protected information, the problem of logic optimization in the prior art is solved, and faster resynthesis and compilation time and higher emulator performance are achieved.
Patent Information
- Application Number
- CN202411652530.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-22
- Filing Date
- 2024-11-19
- Publication Date
- 2025-05-23
AI Technical Summary
When performing logic optimization in the prior art, it is difficult to effectively deal with DUTs in the hierarchical structure, resulting in increased critical path delays, excessive resynthesis and compilation times, and difficulty in synchronizing timing information between different partitions.
By splitting the representation of the DUT into smaller subpartitions, partitioning is performed based on the corresponding margin of the timing endpoint, the subpartition is processed in parallel multithreading to perform logic optimization techniques, and protected information is retained by inserting the anchor circuit instance.
Reduces critical path delays, improves the clock frequency of the emulator, reduces resynthesis and compile time, improves emulator performance, while retaining important design properties.
Smart Images

Figure CN120030959A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to simulation or hardware prototyping for verifying circuit designs. Specifically, the present disclosure relates to design under test (DUT) processing for logic optimization for simulation or hardware prototyping. Background Art
[0002] Designing a circuit can be an arduous process, especially for today's complex system-on-chip (SoC) circuits. Designs are typically thoroughly tested to ensure functionality, specifications, and reliability. Designs may also be iteratively redesigned to meet target functionality, specifications, and reliability. The importance of performing these tests and redesigns prior to tapeout and manufacturing is significant due to the high cost and complexity of tapeout and manufacturing. Summary of the invention
[0003] An example is a non-transitory computer-readable storage medium comprising stored instructions. When the instructions are executed by one or more processors, the instructions cause the one or more processors to: obtain a representation of a design under test (DUT) and split the representation of the DUT into a plurality of partitions. The representation of the DUT includes optimizable leaf instances and timing paths between corresponding timing start points and timing endpoints. Splitting the representation of the DUT into a plurality of partitions is based on corresponding slacks of the timing endpoints. Each of the plurality of partitions includes one or more timing endpoints of the timing endpoint and a transfer fan-in, the transfer fan-in including one or more optimizable leaf instances of one or more timing paths along the timing path, the one or more timing paths terminating at corresponding one or more timing endpoints.
[0004] Another example is a system including a memory and a processing device. The memory stores instructions. The processing device is coupled to the memory and executes the instructions. The instructions, when executed, cause the processing device to: obtain a representation of a DUT including a plurality of partitions; for each of the plurality of partitions, determine whether the corresponding partition includes a first optimizable leaf instance driving a second optimizable leaf instance; and based on determining that the corresponding partition includes the first optimizable leaf instance driving the second optimizable leaf instance, mark the corresponding partition for logic optimization.
[0005] Another example is a method. A representation of a DUT is obtained. An anchor circuit instance is inserted into the representation of the DUT by a processing device, and the anchor circuit instance is connected to a timing path between a first port of an optimizable leaf instance and a second port of the circuit instance. The first port includes protected information. The protected information of the first port is mapped to an anchor port of the anchor circuit instance. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] The present disclosure will be more fully understood from the detailed description given below and from the accompanying drawings of embodiments of the present disclosure. These drawings are used to provide knowledge and understanding of the embodiments of the present disclosure, and do not limit the scope of the present disclosure to these specific embodiments. In addition, these drawings are not necessarily drawn to scale.
[0007] Figure 1 is a simulation and / or prototyping environment for functional verification according to some examples.
[0008] Figure 2 is a flow chart of a method of partition splitting with timing considerations according to some examples.
[0009] Figure 3 is a block schematic diagram of partitions of a representation of a design under test (DUT) according to some examples.
[0010] Figure 4 is a flow chart of a method of including and excluding partitions for analysis in a logic optimization technique according to some examples.
[0011] Figure 5 and Figure 6 is a block diagram of corresponding partitions according to some examples.
[0012] Figure 7 is a flow chart of a method for maintaining protected information in a representation of a DUT according to some examples.
[0013] Figure 8 is a portion of a representation of a DUT according to some examples.
[0014] Fig. 9 An anchor circuit instance is inserted according to some examples Figure 8 A representation of a part of the DUT.
[0015] Fig.10 is a portion of a representation of a DUT according to some examples.
[0016] Fig.11 An anchor circuit instance is inserted according to some examples Fig.10 A representation of a part of the DUT.
[0017] Fig.12 An anchor circuit instance is inserted according to some examples Fig.10 A representation of a part of the DUT.
[0018] Fig.13 is a flow chart of a method for processing a DUT for mapping to a simulation and / or prototyping system according to some examples.
[0019] Fig.14Depicted are flow diagrams of various processes used during the design and fabrication of integrated circuits, according to some examples.
[0020] Fig.15 Depicted is a diagram of an example simulation system according to some examples.
[0021] Fig.16 A diagram is depicted of an example computer system in which the example may operate. DETAILED DESCRIPTION
[0022] Aspects of the present disclosure relate to logic optimized design under test (DUT) processing for simulation or hardware prototyping. The present disclosure includes splitting partitions of a representation of the DUT, excluding trivial partitions of the representation of the DUT from logic optimization techniques, and inserting anchor circuit instances for remapping protected information from ports in the representation of the DUT.
[0023] Functional verification of the DUT may include simulating or prototyping the DUT on a simulation or hardware prototyping system. The simulation or hardware prototyping system may include one or more field programmable gate arrays (FPGAs) on which a representation of the DUT is instantiated for functional verification. The representation of the DUT may undergo synthesis, resynthesis, and compilation to obtain a representation of the DUT (e.g., as a bitstream) that can be instantiated on one or more FPGAs. For example, a hardware description language (HDL) representation of the DUT (e.g., register transfer language (RTL)) may be synthesized into an FPGA netlist, which may be further resynthesized into another (e.g., optimized) FPGA netlist. The resynthesized FPGA netlist may then be compiled into an executable file (e.g., a bitstream) that can be instantiated on one or more FPGAs.
[0024] The representation of the DUT on which synthesis is performed can be partitioned based on the hierarchical structure of the modules within the DUT (e.g., each partition includes instances of the same hierarchical level). Because a hierarchical strategy can be employed during synthesis, the resulting netlist can be suboptimal in terms of area and timing. In addition, the synthesizer rarely spends much effort to optimize path delays and often generates long logic paths. To improve critical path delays, a resynthesis of the FPGA netlist can be performed using a logic optimization tool.
[0025] General logic optimization tools can be powerful. However, there are several problems with resynthesis using general logic optimization tools. First, general logic optimization tools may not support hierarchical netlists. Second, it may not be feasible to perform logic optimization techniques on a completely flat netlist due to the large DUT and the resulting long run time. Hierarchical structures are usually viewed as partitions, and logic optimization techniques can be performed using parallel multithreading on the partitions. However, any potential solution is limited by the partitions, and the solution may be suboptimal. It will be difficult for the logic optimization tool to synchronize the timing information between different partitions. For example, if a long path passes through different partitions, the path will be considered as two short paths in the two partitions, which is misleading for the logic optimization tool. In addition, performing logic optimization techniques on large partitions containing a large number of optimizable instances may require a long run time and may limit the throughput of the netlist resynthesis. In addition, some partitions may not contribute to the resynthesized solution, and therefore, the computing resources and time of the logic optimization tool are wasted on these partitions. In addition, in many DUTs, some design properties must be retained, which may cause the solution generated by the logic optimization tool to be suboptimal.
[0026] The technical advantages of some examples of the present disclosure include, but are not limited to: improving simulation or hardware prototyping system performance and re-synthesis and compile time. These examples include splitting a partition into smaller sub-partitions based on timing path considerations. The critical path of a partition (which can be flattened into, for example, an optimizable leaf instance or an optimizable circuit unit) can be aggregated into smaller sub-partitions so that logic optimization techniques can optimize the critical path while the sub-partitions remain small enough to not grow or reduce the overhead of compile time. Sub-partitions can be processed in parallel and multi-threaded to perform logic optimization techniques, which can speed up re-synthesis. By keeping the critical path in the corresponding sub-partition, critical path delays can be reduced by logic optimization techniques so that the simulator can operate at a faster clock frequency. Therefore, re-synthesis and compile time can be reduced while improving simulator performance by a faster clock frequency.
[0027] In addition, the technical advantages of some examples of the present disclosure include, but are not limited to: speeding up resynthesis by excluding trivial partitions. Typically, the representation of a DUT may include many partitions. Some of these partitions may not significantly affect the solution of the logic optimization technique, but may cause the logic optimization technique to consume a large amount of computing resources. Therefore, some examples exclude these trivial partitions from the analysis of the logic optimization technique, which allows for a reduction in the runtime of resynthesis.
[0028] In addition, the technical advantages of some examples of the present disclosure include, but are not limited to: improving the solution of logic optimization technology while retaining protected information in the representation of the DUT. For many DUTs, the DUT will include design information to be retained and protected, such as for debugging. The protected information may include error path information and waveform observation point information. Therefore, some examples include inserting an anchor circuit instance in the representation of the DUT connected to a port with protected information, and mapping the protected information to the anchor port of the anchor circuit instance. Therefore, a logic instance with a port (e.g., an optimizable leaf instance) can be exposed to the logic optimization technology (rather than being treated as a black box that may not be optimized), which can improve the solution of the logic optimization technology while the protected information is maintained at the anchor port.
[0029] Various combinations of the above generally listed examples can be implemented according to various other examples. Therefore, various other examples can achieve any of the above technical advantages. Other advantages and benefits can be achieved through the examples.
[0030] Various algorithms or techniques described herein may be referred to as optimization algorithms or optimization techniques. The terms "optimization algorithm" and "optimization technique" should be understood to be used in the relevant art, and do not require the result of the optimization algorithm or technique to be the absolutely best or optimal result. Rather, the optimization algorithm or technique can usually determine the result based on some predefined criteria or characteristics, and the result is, for example, located at or close to a local minimum or maximum value in a mathematical space. For example, the result of the optimization algorithm may be based on reaching a gradient less than a predefined threshold in a mathematical space, and the result may not be located at a minimum or maximum value, but may be close to a local minimum or maximum value. Any tool or other device modified by "optimization" or "optimization" can refer to a tool or device that performs an optimization algorithm or optimization technique in whole or in part according to context indications.
[0031] Various modifications may be made to the examples described herein. For example, any method described herein may be performed in any logical order. Such modifications may achieve the same or similar functionality and may achieve the advantages and benefits described above. In addition, although some examples may be described in resynthesis or other contexts, the examples described herein may be implemented for processing (e.g., preprocessing) a DUT for use in a logic optimization technique, regardless of the context or purpose of the logic optimization technique.
[0032] Figure 11 is a simulation and / or hardware prototyping environment 100 ("simulation / prototyping environment 100") for functional verification according to some examples. The simulation / prototyping environment 100 includes a computer system 102 and a simulation and / or hardware prototyping system 104 ("simulation / prototyping system 104"). The example computer system 102 and the example simulation system are described in detail later. The hardware prototyping system can be the same or similar to the simulation system described later. The computer system 102 includes a synthesis tool 112, a logic optimization tool 114, and a compilation tool 116. Each of the synthesis tool 112, the logic optimization tool 114, and the compilation tool 116 operates on the computer system 102. The synthesis tool 112, the logic optimization tool 114, and the compilation tool 116 are illustrated as operating on the same computer system 102. However, in other examples, one or more of the synthesis tool 112, the logic optimization tool 114, and the compilation tool 116 can operate on different computer systems, or can be any arrangement between operating on the same computer system and operating on different computer systems. In addition, each tool in the synthesis tool 112, the logic optimization tool 114 and the compilation tool 116 can be distributed on multiple computer systems. When distributing various tools, some computer systems may be located away from other computer systems on which another tool is operating. For example, the synthesis tool 112 can operate on multiple computer systems (e.g., in a data processing field). Similarly, the logic optimization tool 114 can operate on multiple computer systems (e.g., in a data processing field). Multiple computer systems that implement one or more logic optimization tools 114 can implement multiple parallel threads. The simulation / prototype design system 104 includes one or more field programmable gate arrays (FPGAs) 122.
[0033] The synthesis tool 112, the logic optimization tool 114, and the compilation tool 116 may each be embodied as instructions stored on a non-transitory computer-readable storage medium (e.g., a memory such as a random access memory (RAM), a read-only memory (ROM), etc.). When executed by one or more processors (e.g., a processor of the computer system 102), the instructions cause the one or more processors to perform various functionalities of the respective tools described herein.
[0034] The synthesis tool 112 receives a design under test (DUT) file 132. The DUT file 132 includes a representation of a circuit to be tested before the circuit is manufactured. The DUT file 132 can be or include a representation of the circuit, such as including a hardware description language (HDL) representation, a register transfer language (RTL), a circuit schematic, etc. The synthesis tool 112 is configured and operable to receive or obtain the DUT file 132 and perform a synthesis operation on the DUT file 132 to generate a preliminary netlist corresponding to the representation of the circuit.
[0035] The logic optimization tool 114 receives the preliminary netlist. The logic optimization tool 114 is configured and operable to perform one or more methods described subsequently and perform logic optimization techniques on the preliminary netlist to obtain an optimized netlist. The compilation tool 116 receives the optimized netlist. The compilation tool 116 is configured and operable to compile the optimized netlist and generate a corresponding executable file 134 that can be executed by (multiple) FPGAs 122 of the simulation / prototyping system 104. The executable file 134 may include or be, for example, a bitstream file. The simulation / prototyping system 104 receives the executable file 134 from the computer system 102 (e.g., by direct connection, via a network, and / or via a storage system (e.g., a database)). The simulation / prototyping system 104 loads the executable file 134 onto the FPGA 122 for functional verification.
[0036] Figure 2 2 is a flow chart of a method 200 of partition splitting with timing considerations according to some examples. As described herein, the method 200 can be performed in the context of a representation of a DUT (e.g., a preliminary netlist) including a plurality of partitions. The method 200 can iterate over the plurality of partitions and further split the partitions into corresponding sub-partitions. Figure 2 The method 200 may be particularly useful for splitting very large partitions, but may be applied to any partition of a DUT. In other examples, the representation of the DUT may not include multiple partitions, and in such examples, the method 200 may be performed by treating the representation of the DUT as a single partition to be split. In the method 200, the partitions are described as being split into sub-partitions only for clarity of description; however, the sub-partitions may be simply treated as partitions.
[0037] At 202, a representation of a DUT including a plurality of partitions is obtained. The representation of the DUT and its partitions may be or include a netlist (such as an FPGA netlist) generated by a synthesis operation. In some examples, a partition may include several optimizable leaf instances, and a giant partition may be a partition including more than a certain number (e.g., a predetermined number) of optimizable leaf instances. According to some examples, an optimizable leaf instance is a smallest unit of a circuit component that can be analyzed by a logic optimization technique. In some examples, the logic optimization technique may optimize combinatorial logic, and in some examples, the logic optimization technique may optimize both combinatorial logic and sequential logic (e.g., including synchronous elements). In some examples of performing simulation or prototyping on (multiple) FPGAs based on lookup tables (LUTs), the optimizable leaf instance may be a LUT or include a LUT.
[0038] In some examples, one or more partitions of the representation of the DUT include a hierarchical structure of circuit modules, and in such examples, method 200 may include flattening each such partition to the level of an optimizable leaf instance. Any given circuit module may include submodules. A module or submodule may include one or more optimizable leaf instances. Flattening a partition includes replacing a circuit module (which may be a level of representation in a partition) with one or more optimizable leaf instances and / or non-optimizable circuit instances and corresponding connections therebetween (the corresponding connections are represented by corresponding circuit modules). Flattening the representation of the DUT may aggregate each optimizable leaf instance on the corresponding timing path to the same hierarchical level, which may result in higher quality logic optimization results. Any method described herein for operating on one or more optimizable leaf instances may include flattening the representation of the DUT (or its (multiple) partitions or (multiple) sub-partitions) to the level of an optimizable leaf instance.
[0039] At 204, a static timing analysis (STA) of the DUT is obtained. Static timing analysis can be performed during synthesis as Figure 2 STA includes the margin of each timing path in the DUT, where the timing path is from the timing start point to the timing end point. Any combinatorial logic network can have any number of timing paths through the combinatorial logic network. The timing start point can be an input port of a partition or DUT (for example, if the DUT is not partitioned) or an output port of a synchronization element (such as a trigger) in a partition. The timing end point can be an output port of a partition or DUT (for example, if the DUT is not partitioned) or an input port of a synchronization element in a partition. A synchronization element can terminate a timing path and be the beginning of another timing path. In some examples, a timing path may include a synchronization element including a corresponding timing start point or a timing end point, while in other examples, a synchronization element including a corresponding timing start point or a timing end point may not be included in the timing path. Whether a synchronization element is included in a timing path may depend on the logic optimization technique being implemented.
[0040] At 206, a partition of the representation of the DUT is selected to be split. As indicated in the following description, the method 200 iterates over each partition of the representation of the DUT to be split. The method for selecting a partition for iteration can be based on any criteria.
[0041] At 208, the timing endpoints of the selected partitions are sorted based on the corresponding margins from the STA. In some examples, the timing endpoints of the selected partitions are sorted in increasing order of margin. The timing endpoint may be an endpoint of multiple timing paths. When the timing endpoint is an endpoint of multiple timing paths, a minimum margin among the multiple timing paths is attributed to the timing endpoint for sorting purposes. The minimum margin among the multiple timing paths that terminate at the timing endpoint is used for the purpose of sorting the timing endpoint.
[0042] At 210, a subpartition is created. A subpartition is an organizational structure in which optimizable leaf instances are collected. Although the subpartition is described as being created at 210, in some examples, this may be merely for clarity of description, and in some examples, the creation of the subpartition may be performed simultaneously with the collection of optimizable leaf instances in the subpartition.
[0043] At 212, a timing endpoint of the selected partition with a desired margin (e.g., the lowest margin) is selected (for which the optimizable leaf instances have not been collected into any sub-partition). At 214, the optimizable leaf instances in the transfer fan-in of the selected timing endpoint are collected into the sub-partition. The transfer fan-in includes each timing path that terminates at the selected timing endpoint. At 216, it is determined whether the selected partition includes a timing endpoint with a transfer fan-in that has not been collected into the sub-partition. If the selected partition includes such a timing endpoint as determined in 216, it is determined at 218 whether the number of optimizable leaf instances in the sub-partition reaches or exceeds the target number. If it is determined at 218 that it is not, the method 200 iterates back to 212 to select another timing endpoint with the next lowest margin. If it is determined at 218 that the number of optimizable leaf instances in the sub-partition reaches or exceeds the target number, the method 200 iterates back to 210 to create another sub-partition in which the optimizable leaf instances are collected.
[0044] If it is determined at 216 that the selected partition does not include a timing endpoint with a propagated fan-in that has an optimizable leaf instance that has not been collected in a child partition, then at 220 it is determined whether another partition remains to be split. If so, the method 200 iterates back to 206 to select another partition, and if not, a logic optimization technique is performed on the child partition at 222. The determinations at 216 and 218 may be performed in other orders. For example, the determination of 218 may be made before 216, and each branch from 218 may subsequently include the determination of 216. The iterations from 220 to 206 in the illustrated example are used to analyze multiple partitions. In some examples, instead of and / or in addition to iterative analysis, partitions may be analyzed in parallel.
[0045] The logic optimization techniques for 222 can be or include logic rewriting, logic rebalancing, logic restructuring, etc. The logic optimization techniques can generate corresponding optimized partitions or sub - partitions, or can generate a representation of an optimized DUT including the optimized partitions or sub - partitions. The optimized partitions, sub - partitions, and / or the representation of the DUT can be or include a netlist, such as an FPGA netlist.
[0046] The logic optimization techniques can be executed on the corresponding sub - partitions with multiple parallel threads. By parallelizing the logic optimization techniques, the logic optimization techniques can be executed in a shorter time, which can increase the resynthesis throughput. Additionally, by splitting the partition, the resulting sub - partitions can have fewer optimizable leaf instances (e.g., the corresponding size of the sub - partition is smaller than the size of the partition). Fewer optimizable leaf instances in the sub - partition can enable the logic optimization techniques to execute faster.
[0047] Furthermore, by sorting the timing endpoints based on the slack and collecting the optimizable leaf instances in the (multiple) timing paths that terminate at the corresponding timing endpoints in the sub - partition, the timing paths with lower slack can be better analyzed during the logic optimization techniques. The timing paths can be maintained in the sub - partition. The logic optimization techniques can reduce the arrival time of the lower - slack timing paths, such that the simulator can operate at a higher clock frequency in some cases, thereby reducing the simulation job time and increasing the simulation throughput. By implementing the timing - driven technique to split the partition into sub - partitions, any negative performance impacts in the simulation due to executing the logic optimization techniques with a reduced solution space (e.g., due to smaller sub - partitions) can be mitigated.
[0048] Figure 3 is a block diagram of partition 300 according to some examples. Partition 300 is described to illustrate Figure 2 some operations of method 200. Partition 300 has partition input ports 302, 304, 306, 308, 310, 312 and partition output ports 322, 324, 326, 328. Partition 300 includes optimizable leaf instances 332, 334, 336, 338, 340, 342, 344, 346, 348, 350, 354, 356, 358 and a synchronization element 352. The synchronization element 352 has an input port 362 and an output port 364.
[0049] As an example, Table 1 below outlines the timing paths in partition 300, which have corresponding slack obtained from STA.
[0050]
[0051]
[0052] As shown, partition 300 includes thirteen optimizable leaf instances. In this example, partition 300 will be split into sub-partitions, where the target number of optimizable leaf instances to be included by each sub-partition is at least seven (e.g., the target number at 218 is seven). Therefore, as described below, due to method 200, partition 300 is split into two sub-partitions.
[0053] Table 2 below shows the minimum margins for the corresponding timing endpoints of the partition 300, which are calculated based on Figure 2 208 of method 200 performs sorting in an incremental manner.
[0054]
[0055]
[0056] Thus, in a first iteration at 212, a partition output port 324 (e.g., a timing endpoint) is selected, and at 214, the optimizable leaf instances 336, 338, 340, 342, 344 (which are in the pass-through fan-in of the partition output port 324) (generally indicated by the dashed cone 380) are collected into the first sub-partition. Then, at 216, it is determined that the partition includes an additional timing endpoint, and at 218, it is determined that the number of optimizable leaf instances collected in the first sub-partition is five, which is less than the target number of seven. The method 200 iterates back to 212. In a second iteration at 212, the timing endpoint with the lowest margin (for which the optimizable leaf instances have not yet been collected into any sub-partition) is the partition output port 326. At 214, the optimizable leaf instances 346, 354 in the pass-through fan-in of the partition output port 326 are collected into the first sub-partition. In the depicted example, the logic optimization technique does not optimize synchronization elements, and thus synchronization elements 352 are not collected into the first subpartition. In other examples, the logic optimization technique may optimize synchronization elements, and in such examples, synchronization elements 352 may be collected into the first subpartition at 214. Then, at 216, it is determined that the partition includes additional timing endpoints, and at 218, it is determined that the number of optimizable leaf instances collected in the first subpartition is seven, which is equal to the target number of seven.
[0057] The method 200 iterates to 210 to create an additional subpartition. The method 200 then iterates through the remaining timing endpoints at 212 and collects the corresponding optimizable leaf instances in the second subpartition at 214 until it is determined at 216 that there are no timing endpoints in the partition that have a pass-through fan-in for optimizable leaf instances that have not been collected in any subpartition. The second subpartition includes optimizable leaf instances 332, 334, 348, 350, 356, 358. The method 200 can then further iterate for additional partitions.
[0058] Figure 4 4 is a flow chart of a method 400 of including and excluding partitions for analysis in a logic optimization technique according to some examples. Partitions that can be excluded can be considered trivial partitions. At 402, a representation of a DUT including a plurality of partitions is obtained. The representation of the DUT can be or include a netlist (such as an FPGA netlist) generated by a synthesis operation. In some examples, the partitions can be Figure 2 The method 200 creates a child partition.
[0059] At 404, a partition of the representation of the DUT is selected. As indicated in the following description, method 400 iterates over each partition (or sub-partition) of the representation of the DUT. The method of selecting a partition for iteration can be based on any criteria. At 406, it is determined whether the selected partition includes a first optimizable leaf instance configured to drive a second optimizable leaf instance. The determination of 406 can be performed by iterating over the timing paths of the selected partition until the first optimizable leaf instance configured to drive the second optimizable leaf instance is identified, or all timing paths of the selected partition are exhausted without identifying the first optimizable leaf instance configured to drive the second optimizable leaf instance. In some examples, the transfer fan-in of the timing endpoint can be iterated until the first optimizable leaf instance configured to drive the second optimizable leaf instance is identified, or all timing endpoints of the selected partition are exhausted without identifying the first optimizable leaf instance configured to drive the second optimizable leaf instance.
[0060] If the selected partition includes a first optimizable leaf instance configured to drive a second optimizable leaf instance (as determined at 406), then at 408, the selected partition is marked as included for analysis in the logic optimization technique. If the selected partition does not include a first optimizable leaf instance configured to drive a second optimizable leaf instance (as determined at 406), then at 410, the selected partition is marked as excluded from analysis in the logic optimization technique. After 408 and 410, it is determined at 412 whether the representation of the DUT includes another partition that has not been analyzed. If so, the method 400 iterates back to 404 to select another partition. If not, at 414, the logic optimization technique is performed on the partition marked as included for analysis by the logic optimization technique. The logic optimization technique can be or include logic rewrite, logic rebalancing, logic reconstruction, etc. The logic optimization technique can generate a corresponding optimized partition corresponding to the partition being analyzed (e.g., the partition marked as included), or can generate a representation of the optimized DUT including the optimized partition. The representation of the optimized partitions and / or DUT may be or include a netlist (such as an FPGA netlist).
[0061] Figure 4The method 400 may allow for a reduction in the time used for logic optimization techniques. A partition that does not include a first optimizable leaf instance configured to drive a second optimizable leaf instance may have little impact on the solution space available for the logic optimization technique, and therefore, such a partition may be considered a trivial partition. Performing a logic optimization technique on a trivial partition will cost computing resources and time, with little or no benefit to the final solution and subsequent simulator performance. By excluding trivial partitions from the logic optimization technique, computing resources may be saved, the time for the logic optimization technique may be reduced, and no or little adverse impact on the results of the logic optimization technique and subsequent simulator performance may be achieved.
[0062] Figure 5 and Figure 6 is a block schematic diagram of a corresponding partition 500, 600 according to some examples. Figure 5 The partition 500 includes a first optimizable leaf instance configured to drive a second optimizable leaf instance, and Figure 6 The partition 600 does not include a first optimizable leaf instance configured to drive a second optimizable leaf instance.
[0063] Figure 5 The partition 500 of FIG. 500 includes partition input ports 502, 504, 506, partition output ports 512, 514, optimizable leaf instances 522, 524, 528, and a synchronization element 526. The synchronization element 526 has an input port 532 and an output port 534. To determine whether the partition 500 includes a first optimizable leaf instance that is configured to drive a second optimizable leaf instance, the transfer fan-in (e.g., timing endpoint) of the partition output port 512 can be traversed, where the optimizable leaf instance 522 that is configured to drive the optimizable leaf instance 528 is found. Therefore, the partition 500 can be marked as including a first optimizable leaf instance that is configured to drive a second optimizable leaf instance. Figure 4 The method 400 is analyzed in a logic optimization technique.
[0064] Figure 6The partition 600 includes partition input ports 602, 604, 606, partition output ports 612, 614, optimizable leaf instances 624, 628, and a synchronization element 626. The synchronization element 626 has an input port 632 and an output port 634. To determine whether the partition 600 includes a first optimizable leaf instance configured to drive a second optimizable leaf instance, the transitive fan-in (e.g., timing endpoint) of the partition output port 612 can be traversed (e.g., including a timing path from the partition input port 602 through the optimizable leaf instance 628 to the partition output port 612, and a timing path from the output port 634 of the synchronization element 626 through the optimizable leaf instance 628 to the partition output port 612), wherein the optimizable leaf instance configured to drive another optimizable leaf instance is not identified. Then, the transitive fan-in of the partition output port 614 (e.g., including the timing path from the output port 634 of the synchronization element 626 to the partition output port 614) can be traversed, wherein no optimizable leaf instance configured to drive another optimizable leaf instance is identified. Then, the transitive fan-in (e.g., timing endpoint) of the input port 632 of the synchronization element 626 can be traversed (e.g., including the timing path from the partition input port 604 through the optimizable leaf instance 624 to the input port 632 of the synchronization element 626, and the timing path from the partition input port 606 through the optimizable leaf instance 624 to the input port 632 of the synchronization element 626) can be traversed, wherein no optimizable leaf instance configured to drive another optimizable leaf instance is identified. Thus, the partition 600 can be marked as being from according to Figure 4 The method 400 performs analysis elimination in logic optimization techniques.
[0065] Figure 7 7 is a flow diagram of a method 700 for maintaining protected information in a representation of a DUT according to some examples. At 702, a representation of the DUT is obtained. In some examples, the representation of the DUT can be or include a netlist (such as an FPGA netlist) generated by a synthesis operation. In general, the method 700 can traverse the representation of the DUT to identify ports with protected information, although such traversal is not specifically illustrated.
[0066] At 704, a port with protected information of an optimizable leaf instance in a representation of the DUT is identified. The protected information may be or include any information about the optimizable leaf instance (such as false path information, waveform observations, or other information) that will survive the logic optimization technique. At 706, an anchor circuit instance is inserted into the representation of the DUT and connected to a timing path between the identified port of the optimizable leaf instance and another port of another circuit instance. The anchor circuit instance may be any circuit instance that maintains the logic functionality in the (multiple) timing paths into which the anchor circuit instance is inserted. For example, the anchor circuit instance may be a buffer circuit. At 708, the protected information of the identified port is mapped to the anchor port of the anchor circuit instance. The mapping may include: changing the information identifying the identified port of the optimizable leaf instance in the representation of the DUT to the port of the anchor instance to which the protected information is mapped, or the mapping may include: changing the parameters of the netlist from the identified port of the optimizable leaf instance to specify (e.g., in the HDL code) the port of the anchor instance.
[0067] The method for connecting an anchor circuit instance to a timing path and determining which port of the anchor circuit instance is the anchor port may depend on whether the identified port is an input port or an output port of the optimizable leaf instance and / or may depend on the type of information that is protected information. Figure 8 , Fig. 9 , Fig.10 , Fig.11 and Fig.12 Examples of how an anchor circuit instance may be inserted and which port of the anchor circuit instance may be the anchor port are illustrated according to some examples.
[0068] Figure 8 An optimizable leaf instance 802 and a circuit instance 804 are shown in a representation of a DUT. The circuit instance 804 can be another optimizable leaf instance, a non-optimizable leaf instance such as a synchronization element (e.g., a flip-flop), or another circuit component. The optimizable leaf instance 802 includes an input port 806 having protected information. A first timing path 812 is connected between an output port 808 of the circuit instance 804 and an input port 806 of the optimizable leaf instance 802. Figure 8In the illustrated example, circuit instance 804 is a driver circuit and optimizable leaf instance 802 is a load circuit. Output port 808 of circuit instance 804 is configured to output and drive a signal received at input port 806 of optimizable leaf instance 802 in the representation of the DUT. In the representation of the DUT, output port 808 of circuit instance 804 is electrically connected to input port 806 of optimizable leaf instance 802. Second timing path 814 and third timing path 816 are shown as connected to output port 808 of circuit instance 804, but subsequent circuit instances along these timing paths 814, 816 are not shown to avoid obscuring the aspects described herein. Although timing paths 812, 814, 816 are described separately, each of timing paths 812, 814, 816 may include or represent one or more timing paths.
[0069] In the case where the input port 806 is identified as a port having protected information at 704 of the method 700, an anchor circuit instance 902 (e.g., a buffer circuit) is inserted at 706 of the method 700, which is connected to the first timing path 812, such as Fig. 9 At 708 of method 700 , the protected information of input port 806 is mapped to output port 904 of anchor circuit instance 902 .
[0070] like Fig. 9 , the anchor circuit instance 902 is inserted in the first timing path 812 between the output port 808 of the circuit instance 804 and the input port 906 (e.g., without protected information) of the optimizable leaf instance 802. The input port of the anchor circuit instance 902 is electrically connected to the output port 808 of the circuit instance 804, and the output port 904 of the anchor circuit instance 902 is electrically connected to the input port 906 of the optimizable leaf instance 802.
[0071] Since the port with protected information (e.g., input port 806) is an input port, if the output port 808 of the circuit instance 804 is electrically connected to other circuit instances, such as illustrated by timing paths 814, 816, then the anchor circuit instance 902 is inserted downstream from (multiple branches) to other load circuit instances. Fig. 9 As illustrated in , anchor circuit instance 902 is not connected in timing paths 814 , 816 , but an input port of anchor circuit instance 902 is electrically connected to timing paths 814 , 816 by being electrically connected to output port 808 of circuit instance 804 .
[0072] Fig.10An optimizable leaf instance 1002 and a circuit instance 1004 are shown in a representation of a DUT. The circuit instance 1004 can be another optimizable leaf instance, a non-optimizable leaf instance such as a synchronization element (e.g., a flip-flop), or another circuit component. The optimizable leaf instance 1002 includes an output port 1006 having protected information. A first timing path 1012 is connected between the output port 1006 of the optimizable leaf instance 1002 and an input port 1008 of the circuit instance 1004. Fig.10 In the illustrated example, the optimizable leaf instance 1002 is a driver circuit and the circuit instance 1004 is a load circuit. The output port 1006 of the optimizable leaf instance 1002 is configured to output and drive a signal received at an input port 1008 of the circuit instance 1004 in the representation of the DUT. In the representation of the DUT, the output port 1006 of the optimizable leaf instance 1002 is electrically connected to the input port 1008 of the circuit instance 1004. The second timing path 1014 and the third timing path 1016 are shown as being connected to the output port 1006 of the optimizable leaf instance 1002, but subsequent circuit instances along these timing paths 1014, 1016 are not shown to avoid obscuring the aspects described herein. Although the timing paths 1012, 1014, 1016 are described separately, each of the timing paths 1012, 1014, 1016 can include or represent one or more timing paths.
[0073] In the case where the output port 1006 is identified as a port having protected information at 704 of the method 700, at 706 of the method 700, an anchor circuit instance 1102, 1202 (e.g., a buffer circuit) is inserted, which is connected to the first timing path 1012, as shown in FIG. Fig.11 and Fig.12 At 708 of method 700, the protected information of output port 1006 is mapped to input port 1104 and / or output port 1106 of anchor circuit instance 1102, as shown in FIG. Fig.11 , or mapped to the input port 1204 and / or output port 1206 of the anchor circuit instance 1202, as shown in Fig.12 In some examples, the protected information may be mapped to one of the input ports 1104, 1204 or the output ports 1106, 1206. In some examples, some of the protected information may be mapped to the input ports 1104, 1204, while other protected information may be mapped to the output ports 1106, 1206.
[0074] like Fig.11As shown in FIG. 1 , the anchor circuit instance 1102 is inserted in the first timing path 1012 between the output port 1108 (e.g., not having protected information) of the optimizable leaf instance 1002 and the input port 1008 of the circuit instance 1004. The input port 1104 of the anchor circuit instance 1102 is electrically connected to the output port 1108 of the optimizable leaf instance 1002, and the output port 1106 of the anchor circuit instance 1102 is electrically connected to the input port 1008 of the circuit instance 1004 and the corresponding input ports of the circuit instances in the other timing paths 1014, 1016.
[0075] Since the port with protected information (e.g., output port 1006) is an output port, if the output port 1006 of the optimizable leaf instance 1002 is electrically connected to other circuit instances, such as illustrated by timing paths 1014, 1016, the anchor circuit instance 1102 is inserted upstream from (multiple branches) to other load circuit instances. Fig.11 As shown in , an anchor circuit instance 1102 is connected in each timing path 1012 , 1014 , 1016 .
[0076] like Fig.12 , the anchor circuit instance 1202 is inserted into the first timing path 1012 connected between the output port 1208 (e.g., without protected information) of the optimizable leaf instance 1002 and the input port 1008 of the circuit instance 1004. The input port 1204 of the anchor circuit instance 1202 is electrically connected to the output port 1208 of the optimizable leaf instance 1002, and the output port 1206 of the anchor circuit instance 1202 can remain electrically floating.
[0077] Since the port with protected information (e.g., output port 1006) is an output port, if the output port 1006 of the optimizable leaf instance 1002 is electrically connected to other circuit instances, such as illustrated by timing paths 1014, 1016, the anchor circuit instance 1202 can be connected to the output port 1208 of the optimizable leaf instance 1002 via any of the timing paths 1012, 1014, 1016 because the potentials (e.g., voltages) on the timing paths 1012, 1014, 1016 are the same. Fig.12 As shown in , anchor circuit instance 1202 is connected to timing paths 1012 , 1014 , 1016 .
[0078] If the protected information includes false path information, then the anchor circuit instance is serially inserted in the timing path(s) connected to the port of the optimizable leaf instance having the protected information, as determined by Fig. 9 and Fig.11As shown. A false path can be a path that is topologically present in the DUT but does not have functionality and / or does not require timing. As an example, false path information can be marked on a timing arc that includes a pair of driver-reader ports on the same network, which specifies that the timing arc is part of a false path. In order to retain the false path information, the connection from the driver port to the reader port needs to be retained. By inserting an anchor instance between the driver (output) port and the reader (input) port, the connection between the ports can be restored after logic optimization. If the port with protected information is an input port, such as Figure 8 As in , the anchor circuit instance is serially inserted in the timing path(s) connected to the input port, the anchor circuit may be downstream of any branch of another timing path, the other timing path may be connected to the output port of the circuit instance, and is connected to the input port of the optimizable leaf instance. If the port with protected information is an output port, such as Fig.10 As in, the anchor circuit instance is inserted serially in the timing path(s) connected to the output port, and the anchor circuit may be upstream from any branch between the timing paths that may be connected to the output port.
[0079] If the protected information includes waveform observation point information, then the anchor circuit instance can be inserted in parallel with the timing paths connected to the ports of the optimizable leaf instance with the protected information, such as Fig.12 . The waveform observation point includes a location where the signal can be retained and can be used for debugging in subsequent steps. If the signal is an intermediate signal that does not have a one-to-one mapping before and after logic optimization, the signal may be lost after logic optimization. In some examples, the port with waveform observation point information can be regarded as a primary output to be retained after logic optimization. Since the waveform observed at the port can be the same along any timing path connected to the port, more freedom can be obtained in inserting anchor circuit instances, and anchor circuit instances can be inserted in parallel.
[0080] review Figure 7 , at 710, a logic optimization technique is performed on a representation of the DUT (or its partition(s)) that includes the anchor circuit instance. During the logic optimization technique, the anchor circuit instance is treated as a black box that is not affected by the optimization, while the optimizable leaf instances may be exposed to the logic optimization technique and may be optimized. The logic optimization technique may be or include logic rewrite, logic rebalancing, logic reconstruction, and the like. The logic optimization technique may produce an optimized representation of the DUT (or its partition(s)) that includes the anchor circuit instance. The optimized representation of the DUT may be or include a netlist (such as an FPGA netlist).
[0081] At 712, the protected information is mapped from the anchor port to the port of the optimizable leaf instance in the representation of the optimized DUT. The mapping may include changing the information identifying the port of the anchor instance in the representation of the DUT back to the identified port of the optimizable leaf instance, or may include changing a parameter of the netlist from the port of the anchor instance to specify (e.g., in the HDL code) the identified port of the optimizable leaf instance. At 714, the anchor circuit instance is removed from the representation of the optimized DUT. Removing the anchor circuit instance may be handled by, for example, making the output port of the anchor circuit instance the same net as the input port of the anchor circuit instance (which may effectively short the anchor circuit instance, thereby restoring any timing paths without the anchor circuit instance).
[0082] By inserting anchor circuit instances as described above, protected information can be anchored (e.g., maintained) in the representation of the DUT during logic optimization techniques while exposing optimizable leaf instances to the logic optimization techniques. Without inserting anchor circuit instances, optimizable leaf instances having ports with protected information may be treated as black boxes that are not affected by optimizations by the logic optimization techniques. Since the results of the logic optimization techniques do not necessarily have a one-to-one mapping with the representation of the DUT before the logic optimization techniques, optimizable leaf instances may be processed in this manner, which may result in the loss of protected information. Anchor circuit instances allow ports and / or timing paths to be maintained in a one-to-one mapping in the representation of the optimized DUT while exposing optimizable leaf instances to the logic optimization techniques. Exposing optimizable leaf instances to logic optimization techniques can produce a better or more optimal representation of the optimized DUT.
[0083] Fig.13 1 is a flow chart of a method 1300 for processing a DUT to map to a simulation and / or prototyping system according to some examples. As shown, the method 1300 includes Figure 2 , Figure 4 and Figure 7 Other examples may include any combination or arrangement of the operations of methods 200, 400, 700, including omitting any operation of methods 200, 400, 700.
[0084] At 1302, a representation of the DUT is obtained, such as Figure 2 , Figure 4 and Figure 7 At 1304, partition splitting is performed, which includes Figure 2 At 1306, trivial partition elimination is performed, which includes Figure 4 When partition splitting is performed in 1304, the sub-partitions may be considered as partitions in 1306. At 1308, anchor circuit instance insertion is performed, which includes Figure 7At 1310, a logic optimization technique is performed on the representation of the DUT, such as Figure 2 , Figure 4 and Figure 7 222, 414, 710 in. The logic optimization technique may be or include logic rewriting, logic rebalancing, logic reconstruction, etc. At 1312, anchor circuit instance removal is performed, which includes Figure 7 712-714 of .
[0085] In some examples, implementation Fig.13 The method 1300 produces improved results. In an example, by splitting the partition as in 1304, the runtime of resynthesis is improved by more than 20%, and simulator performance is improved by 1%, compared to arbitrarily splitting the partition in two. By splitting the partition as in 1304, the runtime of resynthesis is improved by more than 50%, compared to resynthesis without splitting the partition. The average compile time overhead during resynthesis is reduced from more than 70 minutes to 25 minutes. The longest compile time overhead during resynthesis is reduced from more than 30 minutes to less than 15 minutes.
[0086] Fig.14 A set of example processes 1400 used during the design, verification, and manufacture of an article such as an integrated circuit are illustrated for converting and verifying design data and instructions representing the integrated circuit. Each of these processes can be constructed and enabled as multiple modules or operations. The term "EDA" represents the term "electronic design automation". These processes begin with creating a product idea 1410 using information provided by a designer, which is converted to create an article using a set of EDA processes 1412. When the design is complete, the design is taped out 1434, at which time the artwork (e.g., geometric pattern) of the integrated circuit is sent to a preparation facility to manufacture a mask set, which is then used to manufacture the integrated circuit. After tape-out, semiconductor chips are manufactured 1436, and packaging and assembly processes 1438 are performed to produce 1440 finished integrated circuits.
[0087] Specifications for circuits or electronic structures can range from low-level transistor material layouts to high-level description languages. High-level representations can be used to design circuits and systems using hardware description languages (HDLs) such as VHDL, Verilog, SystemVerilog, SystemC, MyHDL, or OpenVera. HDL descriptions can be converted to logic-level register transfer level (RTL) descriptions, gate-level descriptions, layout-level descriptions, or mask-level descriptions. Each lower level of representation as a more detailed description adds more useful details to the design description, such as more details of the modules that contain the description. Each lower level of representation as a more detailed description can be computer-generated, derived from a design library, or created by another design automation process. An example of a specification language at a lower level of representation language that specifies a more detailed description is SPICE, which is used for detailed descriptions of circuits with many analog components. The description at each level of representation can be used by a corresponding system for that layer, such as a formal verification system. The design process can use Fig.14 The sequence depicted in . The described process is implemented by an EDA product (or EDA system).
[0088] During system design 1414, the functionality of the integrated circuit to be manufactured is specified. The design may be optimized for desired characteristics such as power consumption, performance, area (physical and / or lines of code), and cost reduction. The design may be divided into different types of modules or components at this stage.
[0089] During logic design and functional verification 1416, modules or components in a circuit are specified in one or more description languages and the functional accuracy of the specifications is checked. For example, components of a circuit can be verified to generate outputs that match the requirements of the specification of the circuit or system being designed. Functional verification can use simulators and other programs such as test bench generators, static HDL checkers, and formal verifiers. In some embodiments, special systems of components called "emulators" or "prototyping systems" are used to accelerate functional verification.
[0090] During synthesis and design for test 1418, the HDL code is converted to a netlist. In some embodiments, the netlist can be a graph structure, wherein the nodes of the graph structure represent components of the circuit, and wherein the edges of the graph structure represent how the components are interconnected. Both the HDL code and the netlist are layered artifacts that an EDA product can use to verify that an integrated circuit performs according to a specified design when manufactured. The netlist can be optimized for a target semiconductor manufacturing technology. In addition, the finished integrated circuit can be tested to verify that the integrated circuit meets the requirements of the specification.
[0091] During netlist verification 1420, the netlist is checked for compliance with timing constraints and for correspondence with the HDL code. During design planning 1422, an overall floor plan for the integrated circuit is constructed and analyzed for timing and top-level routing.
[0092] During layout or physical implementation 1424, physical placement (positioning of circuit components such as transistors or capacitors) and routing (connection of circuit components through multiple conductors) occur, and selection of cells from a library to enable specific logic functions can be performed. As used herein, the term "cell" can specify a group of transistors, other components, and interconnections that provide Boolean logic functions (such as AND, OR, NOT, XOR), storage functions (such as flip-flops or latches), etc. As used herein, a circuit "block" can refer to two or more cells. Both cells and circuit blocks can be referred to as modules or components, and can be enabled as physical structures and in simulations. Parameters such as size are specified for the selected cell (based on a "standard cell"), and are made accessible in a database for use by EDA products.
[0093] During analysis and extraction 1426, circuit functionality is verified at the layout level, allowing the layout design to be refined. During physical verification 1428, the layout design is checked to ensure that manufacturing constraints such as DRC constraints, electrical constraints, lithography constraints are correct and that the circuit device functionality matches the HDL design specifications. During resolution enhancement 1430, the geometry of the layout is transformed to improve the way the circuit design is manufactured.
[0094] During tape-out, data is created for producing lithography masks (after applying lithography enhancements, if applicable).During mask data preparation 1432, tape-out data is used to produce lithography masks, which are used to produce finished integrated circuits.
[0095] Computer systems (such as Fig.16 Computer system 1600 or Fig.15 The storage subsystem of the host system 1507) can be used to store programs and data structures used by some or all of the EDA products described herein, as well as products used to develop libraries and physical and logical design units that use the libraries.
[0096] Fig.15A diagram of an example simulation environment 1500 is depicted. The simulation environment 1500 can be configured to verify the functionality of a circuit design. The simulation environment 1500 can include a host system 1507 (e.g., a computer as part of an EDA system) and a simulation system 1502 (e.g., a set of programmable devices, such as a field programmable gate array (FPGA) or a processor). The host system generates data and information by using a compiler 1510 to construct the simulation system to simulate the circuit design. The circuit design to be simulated is also referred to as a design under test (DUT), where data and information from the simulation are used to verify the functionality of the DUT.
[0097] The host system 1507 may include one or more processors. In embodiments where the host system includes multiple processors, the functions described herein as being performed by the host system may be distributed among the multiple processors. The host system 1507 may include a compiler 1510 to convert specifications written in a description language representing the DUT and generate data (e.g., binary data) and information that are used to construct the simulation system 1502 to simulate the DUT. The compiler 1510 may convert, change, restructure, add new functionality to the DUT, and / or control the timing of the DUT.
[0098] The host system 1507 and the simulation system 1502 use signals carried by the simulation connection to exchange data and information. The connection can be (but is not limited to) one or more cables, such as a cable with a pin structure compatible with the Recommended Standard 232 (RS232) or Universal Serial Bus (USB) protocol. The connection can be a wired communication medium or network, such as a local area network or a wide area network (such as the Internet). The connection can be a wireless communication medium or network, which has one or more access points using a wireless protocol (such as Bluetooth or IEEE802.11). The host system 1507 and the simulation system 1502 can exchange data and information through a third device (such as a network server).
[0099] The simulation system 1502 includes a plurality of FPGAs (or other modules), such as FPGA 1504 1 and 1504 2 And additional FPGA until 1504 N. Each FPGA may include one or more FPGA interfaces through which the FPGA is connected to other FPGAs (and potentially other simulation components) so that the FPGAs exchange signals. FPGA interfaces may be referred to as input / output pins or FPGA pads. Although an emulator may include an FPGA, embodiments of an emulator may include other types of logic blocks instead of an FPGA, or used in conjunction with an FPGA to simulate a DUT. For example, the simulation system 1502 may include a custom FPGA, a dedicated ASIC for simulation or prototyping, memory, and input / output devices.
[0100] A programmable device may include an array of programmable logic blocks and a hierarchical structure of interconnections that enable the programmable logic blocks to be interconnected according to descriptions in the HDL code. Each programmable logic block may enable complex combinational functions or enable logic gates, such as AND and XOR logic blocks. In some embodiments, the logic blocks may also include memory elements / devices, which may be simple latches, flip-flops, or other memory blocks. Depending on the length of the interconnections between different logic blocks, signals may arrive at the input terminals of the logic blocks at different times and may therefore be temporarily stored in the memory elements / devices.
[0101] FPGA 1504 1 -804 N Can be placed onto one or more boards 1512 1 and 1512 2 and additional boards until 1512 M Multiple boards can be placed into the simulation unit 1514 1 The boards within a simulation unit can be connected using the simulation unit's backplane or any other type of connection. In addition, multiple simulation units (such as 1514 1 and 1514 2 To 1514 K ) are connected to each other to form a multi-simulation unit system.
[0102] For a DUT to be simulated, the host system 1507 transmits one or more bit files to the simulation system 1502. The bit file may specify a description of the DUT, and may further specify partitions of the DUT created by the host system 1507 using the trace and injection logic, mapping of the partitions to the FPGA of the simulator, and design constraints. Using the bit file, the simulator constructs the FPGA to perform the functions of the DUT. In some embodiments, one or more FPGAs of the simulator may have the trace and injection logic built into the silicon of the FPGA. In such an embodiment, the FPGA may not be constructed by the host system to simulate the trace and injection logic.
[0103] The host system 1507 receives a description of the DUT to be simulated. In some embodiments, the DUT description uses a description language (e.g., register transfer language (RTL)). In some embodiments, the DUT description uses a netlist-level file or a mixture of a netlist-level file and an HDL file. If part of the DUT description or the entire DUT description uses HDL, the host system can synthesize the DUT description to create a gate-level netlist using the DUT description. The host system can use the netlist of the DUT to divide the DUT into multiple partitions, wherein one or more partitions include tracking and injection logic. The tracking and injection logic tracks interface signals exchanged via the interface of the FPGA. In addition, the tracking and injection logic can inject the tracked interface signals into the logic of the FPGA. The host system maps each partition to the FPGA of the simulator. In some embodiments, the tracking and injection logic is included in the selected partition of the FPGA group. The tracking and injection logic can be built into one or more FPGAs of the simulator. The host system can synthesize a multiplexer to be mapped into the FPGA. The tracking and injection logic can use the multiplexer to inject the interface signal into the DUT logic.
[0104] The host system creates a bit file that describes each partition of the DUT and the mapping of the partition to the FPGA. For partitions that include trace and injection logic, the bit file also describes the included logic. The bit file can include placement and routing information and design constraints. The host system stores the bit file and information describing which FPGAs will emulate each component of the DUT (e.g., which FPGA each component is mapped to).
[0105] Upon request, the host system transmits the bit file to the simulator. The host system sends a signal to the simulator to start the simulation of the DUT. During the simulation of the DUT or at the end of the simulation, the host system receives the simulation results from the simulator through the simulation connection. The simulation results are the data and information generated by the simulator during the simulation of the DUT, which include the interface signals tracked by the trace and injection logic of each FPGA and the status of the interface signals. The host system can store the simulation results and / or transmit the simulation results to another processing system.
[0106] After simulation of the DUT, the circuit designer may request to debug a component of the DUT. If such a request is made, the circuit designer may specify a time period of the simulation to be debugged. The host system uses the stored information to identify which FPGAs are simulating the component. The host system retrieves the stored interface signals associated with the time period, which are tracked by the trace and injection logic of each identified FPGA. The host system sends a signal to the simulator to re-simulate the identified FPGA. The host system transmits the retrieved interface signals to the simulator to re-simulate the component within the specified time period. The trace and injection logic of each identified FPGA injects its respective interface signals received from the host system into the logic of the DUT mapped to the FPGA. In the case of multiple re-simulations of the FPGA, the merged results produce a complete debug view.
[0107] The host system receives from the simulation system signals tracked by the logic of the identified FPGA during the re-simulation of the component. The host system stores the signals received from the simulator. The signals tracked during the re-simulation may have a higher sampling rate than the sampling rate during the initial simulation. For example, in the initial simulation, the tracked signals may include component states saved once every X milliseconds. However, in the re-simulation, the tracked signals may include states saved once every Y milliseconds, where Y is less than X. If the circuit designer requests to view the waveform of the signal tracked during the re-simulation, the host system may retrieve the stored signal and display a graph of the signal. For example, the host system may generate a waveform of the signal. Afterwards, the circuit designer may request to re-simulate the same component or re-simulate another component in a different time period.
[0108] The host system 1507 and / or the compiler 1510 may include subsystems, such as but not limited to a design synthesizer subsystem, a mapping subsystem, a runtime subsystem, a result subsystem, a debugging subsystem, a waveform subsystem, and a storage subsystem. The subsystem may be constructed and enabled as a single or multiple modules, or two or more subsystems may be constructed as one module. These subsystems together construct a simulator and monitor simulation results.
[0109] The design synthesizer subsystem converts the HDL representing the DUT 1505 into gate-level logic. For a DUT to be simulated, the design synthesizer subsystem receives a description of the DUT. If the description of the DUT is in full or in part in HDL (e.g., RTL or other level of representation), the design synthesizer subsystem synthesizes the HDL of the DUT to create a gate-level netlist, in which the DUT is described in the form of gate-level logic.
[0110] The mapping subsystem partitions the DUT and maps the partitions into the simulator FPGA. The mapping subsystem uses the netlist of the DUT to partition the DUT at the gate level into multiple partitions. For each partition, the mapping subsystem retrieves the gate-level description of the tracking and injection logic and adds the logic to the partition. As described above, the tracking and injection logic included in the partition is used to track the signals exchanged via the interface of the FPGA to which the partition is mapped (tracking interface signals). The tracking and injection logic can be added to the DUT before the partition is made. For example, the design synthesizer subsystem can add the tracking and injection logic before or after synthesizing the HDL of the DUT.
[0111] In addition to including trace and injection logic, the mapping subsystem may include additional trace logic in the partitions to track the states of certain DUT components that are not tracked by trace and injection. The mapping subsystem may include the additional trace logic in the DUT prior to partitioning or in the partitions after partitioning. The design synthesizer subsystem may include the additional trace logic in the HDL description of the DUT prior to synthesizing the HDL description.
[0112] The mapping subsystem maps each partition of the DUT to the FPGA of the simulator. For the partitioning and mapping, the mapping subsystem uses the design rules, design constraints (such as timing or logic constraints), and information about the simulator. For the components of the DUT, the mapping subsystem stores information in the storage subsystem that describes which FPGAs will simulate each component.
[0113] Using the partitioning and mapping, the mapping subsystem generates one or more bit files that describe the created partitions and the mapping of the logic to each FPGA of the simulator. The bit file may include additional information, such as the constraints of the DUT and the connections between the FPGAs and the routing information of the connections within each FPGA. The mapping subsystem may generate a bit file for each partition of the DUT and may store the bit file in the storage subsystem. Based on a request from a circuit designer, the mapping subsystem transmits the bit file to the simulator and the simulator may use the bit file to construct an FPGA to simulate the DUT.
[0114] If the simulator includes a dedicated ASIC with trace and injection logic, the mapping subsystem can generate a specific structure that connects the dedicated ASIC to the DUT. In some embodiments, the mapping subsystem can save information about the traced / injected signals and the location where the information is stored on the dedicated ASIC.
[0115] The runtime subsystem controls the simulation performed by the simulator. The runtime subsystem can cause the simulator to start or stop performing a simulation. In addition, the runtime subsystem can provide input signals and data to the simulator. The input signal can be provided directly to the simulator through a connection, or it can be provided indirectly to the simulator through other input signal devices. For example, the host system can control the input signal device to provide the input signal to the simulator. The input signal device can be, for example, a test board (directly or through a cable), a signal generator, another simulator, or another host system.
[0116] The result subsystem processes the simulation results generated by the simulator. During simulation and / or after the simulation is completed, the result subsystem receives the simulation results generated from the simulator during simulation. The simulation results include the signals tracked during simulation. In particular, the simulation results include interface signals tracked by the tracking and injection logic simulated by each FPGA, and may include signals tracked by additional logic included in the DUT. Each tracked signal may span multiple cycles of the simulation. The tracked signal includes multiple states, and each state is associated with the time of the simulation. The result subsystem stores the tracked signals in the storage subsystem. For each stored signal, the result subsystem may store information indicating which FPGA generated the tracked signal.
[0117] The debug subsystem allows the circuit designer to debug a DUT component. After the simulator has simulated the DUT and the results subsystem has received the interface signals traced by the trace and injection logic during simulation, the circuit designer can request to debug a component of the DUT by re-simulating the component for a specific time period. In the request to debug a component, the circuit designer identifies the component and indicates the time period of the simulation to be debugged. The circuit designer's request can include a sampling rate that indicates how often the logic that traces the signals should save the state of the debugged component.
[0118] The debug subsystem uses the information stored in the storage subsystem by the mapping subsystem to identify one or more FPGAs of the emulator that is emulating the component. For each identified FPGA, the debug subsystem retrieves from the storage subsystem the interface signals tracked by the trace and injection logic of the FPGA during the time period indicated by the circuit designer. For example, the debug subsystem retrieves the state tracked by the trace and injection logic associated with the time period.
[0119] The debug subsystem transmits the retrieved interface signals to the emulator. The debug subsystem instructs the debug subsystem to use the identified FPGAs and instructs the trace and injection logic of each identified FPGA to inject its respective traced signals into the logic of the FPGA to re-simulate the component within the requested time period. The debug subsystem may also transmit the sampling rate provided by the circuit designer to the emulator so that the trace logic traces the state at appropriate intervals.
[0120] To debug a component, the simulator can use the FPGA to which the component has been mapped. In addition, re-simulation of the component can be performed at any point specified by the circuit designer.
[0121] For the identified FPGA, the debug subsystem can send instructions to the emulator to load multiple emulator FPGAs with the same configuration as the identified FPGA. The debug subsystem also signals the emulator to use multiple FPGAs in parallel. Each FPGA from the multiple FPGAs is used with a different time window of the interface signal to generate a larger time window in a shorter amount of time. For example, the identified FPGA may take an hour or more to use a certain number of cycles. However, if multiple FPGAs have the same data and structure as the identified FPGA, and each of these FPGAs runs a subset of cycles, the emulator may take several minutes to make the FPGA use all cycles together.
[0122] The circuit designer can identify a hierarchy or list of DUT signals to be re-simulated. To accomplish this, the debug subsystem determines the FPGAs required to simulate the hierarchy or list of signals, retrieves the necessary interface signals, and transmits the retrieved interface signals to the simulator for re-simulation. Thus, the circuit designer can identify any element (e.g., component, device, or signal) of the DUT to be debugged / re-simulated.
[0123] The waveform subsystem generates waveforms using the traced signals. If the circuit designer requests to view the waveform of a signal traced during the simulation run, the host system retrieves the signal from the storage subsystem. The waveform subsystem displays a graph of the signal. For one or more signals, the waveform subsystem can automatically generate a graph of the signal when the signal is received from the simulator.
[0124] Fig.16An example machine of a computer system 1600 is illustrated in which a set of instructions may be executed to cause the machine to perform any one or more of the methodologies discussed herein. In alternative implementations, the machine may be connected to (e.g., connected to a network) other machines in a LAN, an intranet, an extranet, and / or the Internet. The machine may operate in the capacity of a server or a client machine in a client-server network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or a client machine in a cloud computing infrastructure or environment.
[0125] The machine may be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a web appliance, a server, a network router, a switch or a bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by the machine. In addition, while a single machine is illustrated, the term "machine" should also be understood to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.
[0126] The example computer system 1600 includes a processing device 1602, a main memory 1604 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM)), a static memory 1606 (e.g., flash memory, static random access memory (SRAM), etc.), and a data storage device 1618, which communicate with each other via a bus 1630.
[0127] Processing device 1602 represents one or more processors, such as a microprocessor, a central processing unit, etc. More specifically, the processing device can be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor that implements other instruction sets, or a processor that implements a combination of instruction sets. Processing device 1602 can also be one or more special processing devices, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, etc. Processing device 1602 can be configured to execute instructions 1626 to perform the operations and steps described herein.
[0128] The computer system 1600 may also include a network interface device 1608 to communicate over a network 1620. The computer system 1600 may also include a video display unit 1610 (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)), an alphanumeric input device 1612 (e.g., a keyboard), a cursor control device 1614 (e.g., a mouse), a graphics processing unit 1622, a signal generating device 1616 (e.g., a speaker), a graphics processing unit 1622, a video processing unit 1628, and an audio processing unit 1632.
[0129] The data storage device 1618 may include a machine-readable storage medium 1624 (also referred to as a non-transitory computer-readable storage medium) on which is stored one or more sets of instructions 1626 or software embodying any one or more of the methods or functions described herein. The instructions 1626 may also reside, completely or at least partially, within the main memory 1604 and / or within the processing device 1602 during execution thereof by the computer system 1600, the main memory 1604 and the processing device 1602 also constituting machine-readable storage media.
[0130] In some implementations, the instructions 1626 include instructions for implementing functionality corresponding to the present disclosure. Although the machine-readable storage medium 1624 is shown as a single medium in the example implementation, the term "machine-readable storage medium" should be understood to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more instruction sets. The term "machine-readable storage medium" should also be understood to include any medium that can store or encode instruction sets for machine execution and cause the machine and processing device 1602 to perform any one or more methods of the present disclosure. The term "machine-readable storage medium" should therefore be understood to include, but is not limited to, solid-state memory, optical media, and magnetic media.
[0131] Some parts of the above detailed description have been presented in the form of algorithms and symbolic representations of operations on data bits in computer memory. These algorithmic descriptions and representations are the means by which those skilled in the art of data processing use to most effectively communicate their work content to other persons skilled in the art. An algorithm can be a sequence of operations that lead to a desired result. These operations are those operations that require physical manipulation of physical quantities. These quantities can take the form of electrical or magnetic signals that can be stored, combined, compared, and otherwise manipulated. These signals can be referred to as bits, values, elements, symbols, characters, terms, numbers, etc.
[0132] It should be remembered, however, that all of these terms and similar terms are associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise expressly stated, as will be apparent from this disclosure, it should be understood that throughout the description certain terms refer to actions and processes of computer systems or similar electronic computing devices that manipulate data represented as physical (electronic) quantities within the computer system's registers and memories and convert them into other data similarly represented as physical quantities within the computer system's memories or registers or other such information storage devices.
[0133] The present disclosure also relates to an apparatus for performing the operations herein. The apparatus may be specially constructed for the intended purpose, or it may include a computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a non-transitory computer-readable storage medium, such as but not limited to any type of disk (including floppy disks, optical disks, CD-ROMs, and magneto-optical disks), read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, or any type of medium suitable for storing electronic instructions, each of which is coupled to a computer system bus.
[0134] The algorithms and displays presented herein are not inherently related to any particular computer or other device. Various other systems may be used together with the programs taught herein, or it may prove convenient to construct a more specialized device to perform the method. In addition, the present disclosure is not described with reference to any particular programming language. It should be appreciated that a variety of programming languages may be used to implement the teachings of the present disclosure described herein.
[0135] The present disclosure may be provided as a computer program product or software, which may include a machine-readable storage medium having instructions stored thereon, which may be used to program a computer system (or other electronic device) to perform a process according to the present disclosure. A machine-readable storage medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). For example, a machine-readable (e.g., computer-readable) storage medium includes a machine-readable (e.g., computer-readable) storage medium (such as a read-only memory (ROM), a random access memory (RAM)), a disk storage medium, an optical storage medium, a flash memory device, etc.
[0136] In the foregoing disclosure, the implementation of the present disclosure has been described with reference to the specific example implementations of the present disclosure. Obviously, various modifications may be made thereto without departing from the broader spirit and scope of implementation of the present disclosure as set forth in the following claims. When the present disclosure refers to some elements in the singular, more elements may be depicted in the accompanying drawings, and similar elements are marked with similar numbers. Therefore, the present disclosure and the accompanying drawings should be regarded as illustrative, not restrictive.
Claims
1. A non-transitory computer-readable storage medium comprising stored instructions that, when executed by one or more processors, cause the one or more processors to: obtaining a representation of a design under test (DUT), the representation of the DUT comprising optimizable leaf instances and timing paths between corresponding timing start points and timing end points; and The representation of the DUT is split into a plurality of partitions based on the respective margins of the timing endpoints, each of the plurality of partitions comprising one or more timing endpoints of the timing endpoints and a propagation fan-in comprising one or more optimizable leaf instances of one or more timing paths along the timing paths, the one or more timing paths terminating at the respective one or more timing endpoints.
2. The non-transitory computer-readable storage medium of claim 1, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to: perform a logic optimization on at least one of the plurality of partitions.
3. The non-transitory computer-readable storage medium of claim 1 , wherein the instructions, when executed by the one or more processors, cause the one or more processors to split the representation of the DUT into the plurality of partitions based on the respective margins of the timing endpoints, and further cause the one or more processors to: Iteratively, until each of the timing endpoints that has a transitive fan-in that includes an optimizable leaf instance has been collected in a partition: Create partitions; and Iteratively collecting in the corresponding partition the optimizable leaf instances in the delivery fan-in of the timing endpoint with the lowest margin that have not been collected in any partition until a minimum target number of optimizable leaf instances has been collected in the corresponding partition.
4. The non-transitory computer-readable storage medium of claim 1 , wherein the instructions, when executed by the one or more processors, further cause the one or more processors to: flatten the representation of the DUT to the level of the optimizable leaf instances.
5. The non-transitory computer-readable storage medium of claim 1 , wherein when executed by the one or more processors, the instructions further cause the one or more processors to: For each partition of the plurality of partitions, determining whether the corresponding partition includes a first optimizable leaf instance driving a second optimizable leaf instance; and Based on determining that the corresponding partition includes the first optimizable leaf instance that drives the second optimizable leaf instance, marking the corresponding partition for logic optimization.
6. The non-transitory computer-readable storage medium of claim 5, wherein when executed by the one or more processors, the instructions further cause the one or more processors to: inserting an anchor circuit instance into a partition of the plurality of partitions, and the anchor circuit instance is connected to a timing path of the timing path, the timing path being between a first port of an optimizable leaf instance of the optimizable leaf instances and a second port of a circuit instance, the first port comprising protected information; and The protected information of the first port is mapped to an anchor port of the anchor circuit instance.
7. The non-transitory computer-readable storage medium of claim 1 , wherein when executed by the one or more processors, the instructions further cause the one or more processors to: inserting an anchor circuit instance into a partition of the plurality of partitions, and the anchor circuit instance is connected to a timing path of the timing path, the timing path being between a first port of an optimizable leaf instance of the optimizable leaf instances and a second port of a circuit instance, the first port comprising protected information; and The protected information of the first port is mapped to an anchor port of the anchor circuit instance.
8. A system comprising: a memory storing instructions; as well as a processing device coupled to the memory and executing the instructions, which when executed cause the processing device to: obtaining a representation of a design under test (DUT), the representation of the DUT comprising a plurality of partitions; for each partition of the plurality of partitions, determining whether the corresponding partition includes a first optimizable leaf instance driving a second optimizable leaf instance; as well as Based on determining that the corresponding partition includes the first optimizable leaf instance that drives the second optimizable leaf instance, marking the corresponding partition for logic optimization.
9. The system of claim 8, wherein the instructions, when executed, further cause the processing device to perform the logic optimization on the partition marked for the logic optimization, wherein the logic optimization excludes the other partition based on another determination that the other partitions do not include the first optimizable leaf instance driving the second optimizable leaf instance.
10. The system of claim 8, wherein the instructions, when executed, further cause the processing device to: inserting an anchor circuit instance into a partition marked for inclusion for analysis in the logic optimization technique, and the anchor circuit instance is connected to a timing path between a first port of an optimizable leaf instance and a second port of the circuit instance, the first port including protected information; and The protected information of the first port is mapped to an anchor port of the anchor circuit instance.
11. A method comprising: Obtain a representation of the design under test (DUT); inserting, by a processing device, an anchor circuit instance into the representation of the DUT, and the anchor circuit instance being connected to a timing path between a first port of an optimizable leaf instance and a second port of a circuit instance, the first port comprising protected information; and The protected information of the first port is mapped to an anchor port of the anchor circuit instance.
12. The method according to claim 11, further comprising: performing logic optimization on a representation of the DUT including the anchor circuit instance; as well as The protected information is mapped from the anchor port of the anchor circuit instance to the first port in the optimized representation of the DUT.
13. The method of claim 11, wherein inserting the anchor circuit instance comprises serially inserting and connecting the anchor circuit instance in the timing path. The method of claim 13 , wherein the protected information comprises false path information.
15. The method of claim 11, wherein inserting the anchor circuit instance comprises inserting and connecting the anchor circuit instance in parallel with the timing path. The method of claim 15 , wherein the protected information includes waveform observation point information.
17. The method of claim 11, wherein: Inserting the anchor circuit instance includes inserting the anchor circuit instance in the timing path from an output port of the circuit instance to an input port of the optimizable leaf instance; The input port of the optimizable leaf instance is the first port; The output port of the circuit instance is the second port; and The anchor port is an output port of the anchor circuit instance.
18. The method of claim 11, wherein: Inserting the anchor circuit instance includes inserting the anchor circuit instance in each timing path from an output port of the optimizable leaf instance, the timing path from the output port of the optimizable leaf instance to an input port of the circuit instance; The output port of the optimizable leaf instance is the first port; and The input port of the circuit instance is the second port.
19. The method of claim 11, wherein: Inserting the anchor circuit instance includes inserting the anchor circuit instance connected to the timing path, the timing path from the output port of the optimizable leaf instance to the input port of the circuit instance, the input port of the anchor circuit instance being connected to the timing path; The output port of the optimizable leaf instance is the first port; The input port of the circuit instance is the second port; and The output port of the anchor circuit instance is floating.
20. The method of claim 11, wherein the anchor circuit instance comprises a buffer circuit.