Route-based timing-driven placement using iterative modification of a pseudo netlist

The path-based timing-driven placement method using iterative pseudo netlist modification addresses the limitations of HPWL by enhancing timing compliance and resource efficiency in IC design.

JP7785080B2Active Publication Date: 2025-12-12INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023534941
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-16
Filing Date
2021-12-15
Publication Date
2025-12-12
Estimated Expiration
2041-12-15

AI Technical Summary

Technical Problem

Current Half-Perimeter Wire Length (HPWL) techniques in electronic design automation (EDA) are not timing-aware, leading to IC designs that may not meet timing constraints, and existing solutions are computationally intensive, limited by global placement quality and timing analysis accuracy, and prone to saturation and tortuous paths.

Method used

A path-based timing-driven placement method using iterative modification of a pseudo netlist, where timing-critical paths are identified and pseudo two-pin nets are created to enhance wire-length-driven placement, iteratively adjusting the placement to improve timing compliance.

Benefits of technology

This approach significantly improves timing performance of IC designs, reducing resource usage and achieving better compliance with timing constraints compared to prior art methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007785080000001
    Figure 0007785080000001
  • Figure 0007785080000002
    Figure 0007785080000002
  • Figure 0007785080000003
    Figure 0007785080000003
Patent Text Reader

Abstract

Performing an initial wirelength-driven placement for the integrated circuit design embodied in the unplaced netlist using a computerized placer to obtain a data structure representing an initial placement of logic gates; identifying at least one timing-critical source / sink path between the at least one pair of source / sink endpoints in the data structure representing the initial placement; creating a new pseudo two-pin net for each of the at least one pair of source / sink endpoints to produce an updated netlist; performing a revised wirelength-driven placement on the updated netlist to obtain a data structure representing a revised placement.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to electrical, electronic, and computer technologies, and more particularly to semiconductor electronic design automation (EDA). [Background technology]

[0002] EDA involves the use of software tools to design electronic systems such as integrated circuits (ICs) and printed circuit boards. Generally, an IC has a data signal and a clock, and the data signal must arrive at a particular node at a precise time relative to the time of the corresponding clock cycle of the device at that node. If the data signal does not arrive in time, either the clock is too early or the data signal takes too long to propagate (the path is too slow).

[0003] Currently, Half-Perimeter Wire Length (HPWL) techniques are used for placement within EDA processes. However, HPWL is not timing-aware, such that IC designs developed using HPWL techniques may not meet timing constraints. Previous attempts to address the issue of HPWL not being timing-aware have been excessively computationally intensive, limited by the quality of global placement and / or the accuracy of timing analysis, and prone to saturation and / or tortuous paths between sources and sinks. Summary of the Invention

[0004] The present principles provide a technique for path-based timing-driven placement using iterative modification of a pseudo netlist. In one aspect, an exemplary method for improving timing performance of an electronic circuit designed using electronic design automation includes performing an initial wire-length-driven placement for an integrated circuit design embodied in an unplaced netlist using a computerized placer to obtain a data structure representing an initial placement of logic gates, identifying at least one timing-critical source / sink path between at least one pair of source / sink endpoints in the data structure representing the initial placement, creating a new pseudo two-pin net for each pair of the at least one pair of source / sink endpoints to create an updated netlist, and performing a modified wire-length-driven placement on the updated netlist to obtain a data structure representing the modified placement.

[0005] In one or more embodiments, at least one timing-critical source-sink path includes multiple timing-critical source / sink paths, and at least one pair of source / sink endpoints includes multiple pairs of source / sink endpoints, and the steps of identifying the multiple timing-critical source / sink paths, creating a new pseudo two-pin net for each of the multiple pairs of endpoints, and performing a modified length-driven placement are repeated for multiple full iterations.

[0006] The steps of performing an initial wire length driven placement and performing a modified wire length driven placement each include, for example, applying a half-circle approximate wire length driven placement, and the computerized placer includes a half-circle approximate wire length driven computerized placer.

[0007] In another aspect, an exemplary computer includes a memory and at least one processor coupled to the memory, wherein the at least one processor operates to improve timing performance of an electronic circuit designed using electronic design automation by performing an initial wirelength-driven placement on an integrated circuit design embodied in an unplaced netlist using a computerized placer to obtain a data structure representing an initial placement of logic gates; identifying at least one timing-critical source / sink path between at least one pair of source / sink endpoints in the data structure representing the initial placement; creating a new pseudo two-pin net for each of the at least one pair of source / sink endpoints to create an updated netlist; and performing a revised wirelength-driven placement on the updated netlist to obtain a data structure representing a revised placement.

[0008] In yet another aspect, an exemplary method for improving timing performance of an electronic circuit designed using electronic design automation includes obtaining, from a computerized placer, initial wirelength-driven placement results for an integrated circuit design embodied in an unplaced netlist, the results including a data structure representing an initial placement of logic gates; obtaining, from a computerized timer, at least one timing-critical source / sink path between at least one pair of source / sink endpoints in the data structure representing the initial placement; creating a new pseudo two-pin net for each of the at least one pair of source / sink endpoints to create an updated netlist; and providing the updated netlist to the computerized placer to facilitate the computerized placer performing a modified wirelength-driven placement on the updated netlist to produce a data structure representing a modified placement.

[0009] "Facilitating" an action, as used herein, includes performing an action, simplifying an action, assisting in the performance of an action, or causing an action to be performed. Thus, by way of example and not limitation, instructions executing on one processor may facilitate the performance of an action by instructions executing on a remote processor by sending appropriate data or commands that cause or assist in the performance of the action. For the avoidance of doubt, where an action is facilitated by an actor other than the actor performing the action, the action is nevertheless performed by some entity or combination of entities.

[0010] One or more embodiments of the present invention or elements thereof may be implemented in the form of a computer program product including a computer-readable storage medium having computer-usable program code for performing the illustrated method steps. Furthermore, one or more embodiments of the present invention or elements thereof may be implemented in the form of a system (or apparatus) including a memory and at least one processor coupled to the memory and operative to perform the illustrated method steps. In yet another aspect, one or more embodiments of the present invention or elements thereof may be implemented in the form of a means for performing one or more of the method steps described herein, which may include (i) hardware modules, (ii) software modules stored on a computer-readable storage medium (or multiple such media) and implemented on a hardware processor, or (iii) a combination of (i) and (ii), any of which implements specific techniques described herein.

[0011] The techniques of the present invention provide substantial beneficial technical effects. For example, the embodiments described in this Summary section provide the following advantages:

[0012] Chips designed using aspects of the present invention are superior (eg, have better timing compliance) when compared to chips designed using prior art techniques.

[0013] Computers performing EDA using embodiments of the present invention achieve better results than prior art and, in at least some cases, use fewer resources (such as CPU or memory) than prior art.

[0014] These and other features and advantages of the present invention will become apparent from the following detailed description of illustrative embodiments of the invention, when read in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0015] [Figure 1] FIG. 1 is a block diagram of an EDA process in which embodiments of the present invention may be used for placement. [Figure 2] FIG. 1 illustrates a global placement according to an aspect of the present invention. [Figure 3] 10A and 10B are diagrams illustrating aspects of half-circle approximate wiring length placement that can be improved in accordance with aspects of the present invention. [Figure 4] FIG. 1 illustrates a Timing Driven Attraction (TDA) according to an embodiment of the present invention. [Figure 5] FIG. 1 illustrates timing-driven attraction (TDA) according to an embodiment of the present invention. [Figure 6] FIG. 1 illustrates timing-driven attraction (TDA) according to an embodiment of the present invention. [Figure 7] 1 is a flowchart according to an aspect of the present invention. [Figure 8] 1 is a flowchart according to an aspect of the present invention. [Figure 9] FIG. 1 illustrates exemplary experimental results. [Figure 10] FIG. 1 illustrates a computer system that may be useful in implementing one or more aspects and / or elements of the present invention. [Figure 11]FIG. 1 is a flow diagram of a design process used in semiconductor design, manufacturing, and / or testing. [Figure 12] 10A-10C illustrate further aspects of IC fabrication from physical design data. [Figure 13] FIG. 1 illustrates a flow diagram of an exemplary high-level electronic design automation (EDA) tool in which aspects of the present invention may be used. [Figure 14] FIG. 2 illustrates elements of a TDA in a data structure according to an embodiment of the present invention. [Figure 15] FIG. 1 is a block diagram according to an aspect of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0016] Figure 1 shows the physical synthesis flow process used in semiconductor design. In step 101, a potential design is rendered in a hardware description language such as VHDL. In step 103, logic synthesis is performed to convert from the hardware description language to actual interconnected logic gates. In step 105, initial (global) placement is performed, i.e., translating the abstract description of the hardware (interconnected logic gates) into something that can be laid out on a chip. In step 107, virtual timing optimization is performed. In step 109, clock optimization is performed. In step 111, wiring synthesis, coarse optimization, and fine optimization are performed. In step 113, routing is performed.

[0017] Figure 2 shows the progress of the global placement process starting from netlist 201. First, the total wire length is minimized, which results in a "tight" placement at 203. To meet a given density target, the placement is gradually "widened" as seen at 205, 207, 209, and 211, finally resulting in a legal placement 213 that obeys the constraints and does not have any overlaps.

[0018] One or more embodiments are useful during global placement step 105 and / or clock optimization 109, by way of example and not limitation. Generally, attention is focused on producing a good global placement in step 105, along with a good final result from the physical synthesis flow process. Indeed, physical synthesis's ability to meet timing closure is constrained by the global placement. Traditional global placement algorithms are connectivity-driven linear optimizations (using half-perimeter approximated wirelength (HPWL)). Referring to FIG. 3, in the HPWL approach, wirelength is approximated by half the perimeter of the smallest bounding rectangle 301 that encloses all terminals 303-1, 303-2, 303-3, 303-4, and 303-5 in a net. State-of-the-art HPWL minimization uses force-directed placement. A better global placement HPWL solution (meeting the same density / crowding constraints) often translates to better physical synthesis results. HPWL placement also has the advantage of being relatively fast. Table 311 in Figure 3 shows an example runtime breakdown for the physical synthesis flow process. The placement portion uses only 1 minute 40 seconds of runtime, while the optimization portion takes 10.5 hours and the routing portion takes 17.3 hours.

[0019] As those skilled in the art will appreciate, some paths in a circuit layout are typically more timing critical than others, meaning that low signal propagation times are required along those paths so that the circuit can operate successfully at the desired clock frequency. However, HPWL is not timing-aware. For complex, high-performance logic macros, this can result in a placement solution where physical synthesis is insufficient to meet cycle-time constraints.

[0020] Many attempts and techniques have been used to address the limitations of HPWL, each with its own advantages and disadvantages. Net-weight-driven placers prioritize timing-critical nets based on high net weights, which has the advantage of providing timing information for HPWL optimization. However, this approach saturates quickly, is overly aggressive for high-fanout nets, relies too heavily on timing accuracy, and demonstrates a dynamic where solving one problem can lead to other problems. Hypergraph placement based on storage elements (e.g., flip-flops) helps balance the distance between timing endpoints. However, this approach often introduces congestion issues and can result in meandering paths in combinational logic, where endpoints are closed but the paths are still long. Critical path correction algorithms, used after global placement, can identify critical paths and are effective at reducing the distance signals travel between timing endpoints, but are limited by the quality of the global placement solution. Timing models implemented within the placement procedure (path-based) are mathematically interesting, but are extremely CPU-intensive and overly constrained by the unoptimized netlist.

[0021] One or more embodiments advantageously provide a global placement technique using path-based timing-driven placement (also referred to herein as timing-driven attraction (TDA)) that uses iterative modification of a pseudo netlist to provide timing information to an HPWL placer to bring timing endpoints (i.e., endpoints of timing-critical paths) closer together, correct critical paths, or a combination thereof. The pseudo netlist modification, also referred to herein as "attraction," is a "fake" ("pseudo") two-pin net created only within the HPWL model, from sink to source. The exemplary approach is timing-driven in the sense that the worst "paths" (source timing endpoint to sink timing endpoint) are selected based on timing estimates, and attraction is propagated along those worst-case paths (source / sink pairs undergo attraction). The exemplary approach is implemented in the placer, for example, by re-running global placement with new attraction (i.e., a new netlist with attraction added). An exemplary approach is iterative, for example, in the sense that multiple global placements are performed, incrementally introducing attractions and / or increasing the weight of previous attractions.

[0022] View 401 in Figure 4 is a logic diagram illustrating an example of timing-driven attraction (TDA) according to an embodiment of the present invention for a subset of the exemplary netlist. Boxes 403-1, 403-2, 403-3, 403-4, and 403-5 are timing endpoints, and view 405 on the right represents the physical layout (with actual placement, i.e., represented by the physical HPWL model) corresponding to the elements within dashed box 407 in view 401. There are multiple logic gates, and to avoid confusion, only AND gates 421 and 411, and OR gates 413, 417, and 419 are numbered.

[0023] Figure 5 shows a logic diagram with three timing endpoints 503-1, 503-2, and 503-3, with 503-2 and 503-3 determined to be the worst endpoints. The bold lines 505-1, 505-2, 505-3, 505-4, 505-5, and 505-6 (not all of which are numbered separately to avoid confusion) interconnecting endpoints 503-1 and 503-3 through various logic gates represent attraction. Note that OR gate 599 has an output to AND gate 597 and another output to AND gate 595, but only the output from 599 to 595 is subject to attraction because only that output is on the critical path. Thus, the most timing-critical start / endpoint pair is selected, resulting in a new attraction in the netlist and iterating over the global placement.

[0024] Figure 6 reproduces view 405 from Figure 4 (the default model for NHP-driven placement) and adds a new view 605 (a model with a net critical source-sink pair) that shows the addition of attractions 687 and 689. In view 405 without attraction 687, there is no incentive for gate 419 to move closer to source 421 because the dashed bounding box (and therefore the HPWL) remains unchanged. However, the addition of attractions 687 and 689 in view 605 does provide incentives for elements 421, 419, and 403-4 to move closer together.

[0025] Referring now to the flowchart of FIG. 7, a placement flow using TDA in accordance with an aspect of the present invention is shown. The incoming netlist 801 has no placement. In step 805, HPWL-driven placement is performed using any now-known or later-developed HPWL-based placement algorithm. In step 807, timing estimation is performed. Advantageously, in one or more embodiments, the timing estimation does not need to be accurate in an absolute sense (which is advantageous because accurate timing in an absolute sense is difficult to obtain at the beginning of the synthesis flow), but only needs to be relatively accurate (i.e., one can determine which paths are most critical to facilitate attraction insertion, without needing to know how many picoseconds it will degrade). Furthermore, in one or more embodiments, any now-known or later-developed timer (program that implements timing) may be used (which advantageously allows the technology to easily adapt to improved timing estimation methods). In one or more embodiments, a combination of buffer estimation and reasonable metal layer assumptions is used.

[0026] In step 809, attraction generation is performed as described. As indicated by dashed box 803 and arrow 811, steps 805, 807, and 809 are repeated in an iterative manner. In each iteration, all attractions are treated equally; i.e., everything becomes a new net. It is only desired to correct critical paths and "nudge" the endpoints closer together. With each iteration, in one or more embodiments, all previous attractions are retained. Thus, improved paths maintain the HPWL forces that improved them, and repeatedly impaired paths can be further improved (i.e., are given higher priority). Once the iterations are complete, proceed to timing-based optimization 813.

[0027] FIG. 8 provides an example of TDA use in a physical design closure flow. One or more embodiments are reusable in the sense that they proceed from a hardware description language (e.g., VHDL 901) to gates (from logic synthesis 903 to unplaced netlist 905), which can be used for (first) global placement 907 and proceed through TDA iteration 909. Further, in a subsequent flow, storage elements can be clustered and fixed (so they are no longer movable) at 911, and the remaining circuit elements (other than the fixed storage elements) can then be optimized by iterating placement 913 and TDA 915 to reduce and correct other critical paths. In subsequent stages of the physical synthesis flow, after certain optimizations have been performed, gates in difficult-to-route areas can be widened at 917 using TDA attraction. Advantageously, many of the difficult and complex algorithms during the steps in the dashed box can be completely replaced with TDA placement. Following step 917, additional optimizations can be performed at 919, and routing can be performed at 921.

[0028] Advantageously, compared to net-weight-driven placers, one or more embodiments using TDA add new nets as opposed to simply adjusting the weights of existing nets. Compared to prior art latch-gate techniques, one or more embodiments using TDA offer the ability to pull endpoints closer together without stripping any combinational logic outside of the placement model, since the actual netlist is adjusted. Whereas critical path correction algorithms are performed as a "repair" optimization after global placement, one or more embodiments using TDA are global in nature as opposed to repair optimization. Also, it is only natural that one or more embodiments using TDA do not focus on specifically correcting critical paths, i.e., do not look at these existing corrections, but rather correct paths further, resulting in one or more embodiments using TDA. One or more embodiments using TDA use standard wirelength and placer density / congestion objective functions and use timing only to adjust the placement netlist model as opposed to requiring a timing model in the placer.

[0029] 9 , one or more embodiments provide significant improvements in the worst-case slack and total negative slack (Figure of Merit = FOM) output of physical synthesis and / or the ability to replace many complex placement solution algorithms. Note that even with overall improvements, one or more embodiments typically produce some worst-case HPWL. Display 931 shows a path 939 before TDA (for movable endpoints) and a path 941 after TDA. Display 933 shows a path 943 before TDA (for fixed endpoints) and a path 945 after TDA. Displays 935 and 937 show bar graphs of worst-case slack (ps) and total negative slack for experimental results of seven different exemplary timing-challenged designs (comparing baselines and results using TDA).

[0030] Those skilled in the art will recognize that, based on the discussion herein, one or more embodiments can produce timing-driven placement using a TDA technique, which in one or more embodiments includes selecting critical start / endpoint pairs (using any type of timing estimation), adding "fake" two-pin nets to the netlist between these identified endpoints of all sink-source pairs, and performing placement (using any length-driven placer). One or more embodiments further perform these steps iteratively.

[0031] In one or more embodiments, to select critical start / endpoint pairs (using any timing estimate), a timing simulation is performed in which arrival times (At) and required arrival times (RAT) propagate back and forth through the network (net) from the storage elements to obtain slack values ​​(AT minus RAT) at all sink pins of the net. Knowing the slack at those sink pins helps determine the critical two-pin connections (i.e., between the endpoints involved).

[0032] Advantageously, one or more embodiments can be implemented with any wirelength-driven placer and with any timing estimation technique. Furthermore, one or more embodiments can be easily extended to N+1 synthesis, where the worst result from a previous physical synthesis flow run is automatically selected as the worst endpoint. That is, in some cases, timing information at the time of placement is not necessarily required, but can be from a previous run. Furthermore, one or more embodiments can be extended beyond just physical synthesis flows. For example, referring to FIG. 8 , TDA is shown being used at various points in the design / synthesis process, e.g., 909, 915, and 917. One or more embodiments can provide data flow between routing 921 and the unplaced netlist 905, i.e., the entire flow in FIG. 8 can be scrutinized, with the results fed back into the next round. That is, the intelligence used to implement attraction can utilize the output of routing step 921. In one or more embodiments, an initial placement is performed without attraction to obtain estimated timing, and then placement is reiterated using attraction (using current or previous timing). TDA can thus be carried out outside the scope of the physical synthesis flow, if desired.

[0033] Additionally, one or more embodiments may support user input, allowing a user to specify known critical endpoints (i.e., critical endpoints may, but need not, be provided by a timing program). Additionally, one or more embodiments may be reusable during any placement optimization, not just global placement. For example, one or more embodiments may be used during congestion-driven spreading to reduce timing impacts.

[0034] It has been found that focusing solely on the total wirelength created by the placer (e.g., using HPWL) is not sufficient. One or more embodiments advantageously address the difficult problem of incorporating timing information into placement with a flexible approach that can be extended to any HPWL-driven placer and any timing estimation technique (though, of course, those skilled in the art will recognize that timing estimation requires considerable care). In one or more embodiments, a TDA placement flow has been shown (based on experimental evaluation) to significantly outperform other methods of timing-driven placement.

[0035] FIG. 15 shows an exemplary block diagram including a placer 1501, a timer 1503, an attraction generator 1505, and a shared data structure 1507 that enables data communication and sharing between blocks 1501, 1503, and 1505.

[0036] From the foregoing discussion, it will be appreciated that an exemplary method for improving timing performance of an electronic circuit designed using electronic design automation in accordance with one aspect of the present invention includes (for example) performing initial wirelength-driven placement for an integrated circuit design embodied in an unplaced netlist using a computerized placer 1501 in step 805 to obtain a data structure representing an initial placement of logic gates. Placer 1501 may include a commercially available or academic research placer such as one that implements HPWL techniques, and suitably the selected placer may be capable of providing the attractions described herein by allowing the addition of connections that do not necessarily correspond to the physical pins of the logic gates, as discussed further below.

[0037] A further step 807 includes identifying at least one timing-critical source / sink path between at least one pair of source / sink endpoints in the data structure representing the initial configuration. This step can be performed, for example, using a timer 1503. Non-limiting examples of suitable timers include Cadence Tempus™ (a trademark of Cadence Design Systems, Inc., San Jose, Calif., USA) and Synopsys Primetime® (a registered trademark of Synopsys, Inc., Mountain View, Calif., USA).

[0038] A still further step 809 (see also the discussion of Figures 4 and 5) involves creating a new pseudo 2-pin net (attraction) for each of the at least one pair of source / sink endpoints to create an updated netlist, which can be performed by attraction generator 1505, which implements the logical steps set forth herein to create the pseudo nets.

[0039] Yet a further step (repeating step 805 on the updated netlist) involves performing a modified wirelength-driven placement on the updated netlist to obtain a data structure representing the modified placement, which may be done using placer 1501.

[0040] Here, we provide further notes regarding the attraction and updated netlist used in one or more embodiments. Traditionally, when people think of a netlist, they think of it as a gate in the actual technology, where, for example, an AND gate has two inputs. In one or more embodiments, when a pseudo two-pin net is added, for example, terminating in that AND gate, this net does not actually represent a new signal; rather, it is simply a new net for optimizing HPWL-driven placement. In one or more cases, the updated netlist is a modified traditional netlist with increased attraction, which is not associated with a physical input / output (I / O) for a logic gate, but rather is provided to enhance HPWL calculations. Traditional netlists are representations of actual logic signals. For example, the output of an inverter drives the input of a NAND gate, and this relationship is a two-pin net. An "attraction" in one or more embodiments of the present invention is a two-pin net that does not represent an actual logic signal present on the chip, but is used by the HPWL-driven placement engine in an equivalent manner to such a signal.

[0041] Referring now also to FIG. 14, which shows an exemplary portion of data structure 1507 of FIG. 15, one or more embodiments are shown to create an attraction that results in a new connection between two logic gates. The endpoints of the pseudo two-pin net are, in one or more instances, logic gates. One or more embodiments assign weights to the connection. The weight is an attribute (like a label) for the new connection. Existing nets may also have such labels. The concept of weights for nets is itself well known in placement algorithms. In one exemplary approach, all nets, both traditional nets corresponding to physical connections and nets corresponding to attractions according to aspects of the present invention, are assigned a uniform weight. In one or more exemplary embodiments, iteration 811 is performed, and if step 809 is reached a second time and an attraction already exists, the weight is increased (rather than creating a new attraction, which may not be done in an alternative approach). In some instances, the amount of increase is a user-configurable parameter. One or more embodiments simply increase the weight linearly.

[0042] Thus, by way of review, in one or more embodiments, a two-pin pseudo-net is created by first programming a connection between two logic gates into a data structure using known techniques and assigning initial weights to the connection. For example, in FIG. 14, in the first iteration, an attraction with weight "Weight_1_2" is created between Gate 1 and Gate 2, an attraction with weight "Weight_6_4" is created between Gate 6 and Gate 4, and an attraction with weight "Weight_12_99" is created between Gate 12 and Gate 99. In one exemplary embodiment, all connections are given uniform weights (i.e., Weight_1_2 = Weight_6_4 = Weight_12_99). Those skilled in the art will be familiar with such uniform weighting for conventional representations of actual logic signals and will be able, with the teachings herein, to assign weights to the "attractions" of the present invention. When step 809 reaches the second iteration (the second iteration), the weights are adjusted. When step 807 is repeated a second time, the critical source / sink paths will be determined again. Some pairs of logic gates (e.g., Gate 49 and Gate 51) that did not have an attraction the first time are now found to be critical, and they get an initial weight Weight_49_51 (which can be 1 or another value). Some pairs of logic gates that were critical the first time and have an attraction, e.g., Gate 6 and Gate 4, are again critical, and their weights are increased (e.g., by a factor of two or more, typically 2*Weight_6_4) (instead of introducing a second attraction between Gate 6 and Gate 4, which is not an alternative approach). Net weighting (for conventional representations of actual logic signals) is itself known to those familiar with HPWL. It can be increased by a constant amount per iteration, or proportional to the timing criticality of the net, or in some cases, even not increased at all. The simplest approach is for weight = w in the first iteration, 2w in the second iteration, 3w in the third iteration, etc.It is also possible to add only a small weight, or, as discussed, give more weight to the connection with the worst negative slack. Furthermore, one skilled in the art knows how to create a connection between two gates and assign it a weight (as a traditional representation of an actual logic signal), and can assign weights to the "attractions" of the present invention according to the teachings herein. Non-limiting examples of known weighting techniques include Chentouf, Mohamed, and Zine El Abidine Alaoui Ismaili, "A Novel Net Weighting Algorithm for Power and Timing-Driven Placement," VLSI Design, October 18, 2018; Goplen, Brent, Prashant Saxena, and Sachin Sapatnekar, "Net weighting to reduce repeater counts during placement," in Proceedings 42nd Design Automation Conference, 2005, June 13, 2005 (pp. 503-508), IEEE; and Ren, Haoxing, David Zhigang Pan, and David S. Kung, "Sensitivity guided net weighting for placement-driven synthesis," IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, April 25, 2005, 24(5):711-21. The weights in Figure 14 can thus be included in data structure 1507 and shared with placer 1501. Because the attraction is a pseudo net to enhance placement, physical pins do not need to be specified.

[0043] In one or more embodiments, at least one timing-critical source-sink path includes multiple timing-critical source / sink paths, and at least one pair of source / sink endpoints includes multiple pairs of source / sink endpoints.

[0044] One or more embodiments further include repeating the steps of identifying multiple timing-critical source / sink paths, creating a new pseudo 2-pin net for each of the multiple pairs of endpoints, and performing modified wirelength-driven placement for multiple full iterations. See arrow 811. As used herein, multiple full iterations includes one run plus at least one additional run. Stopping an iteration can be, for example, based on a certain number of iterations, a "convergence" criterion (e.g., timing metrics stop improving and no critical paths remain, or timing gets worse in subsequent iterations due to saturation), or after a certain CPU utilization. In a non-limiting example, referring to FIG. 8, there can be eight iterations 909 and four iterations 915.

[0045] In one or more embodiments, performing the initial wirelength-driven placement and performing the modified wirelength-driven placement each include applying a half-circle approximate wirelength-driven placement, and the computerized placer 1501 includes a half-circle approximate wirelength-driven computerized placer.

[0046] In some cases, identifying multiple timing-critical source / sink paths between pairs of source / sink endpoints in the data structure representing the initial configuration involves obtaining input from an expert in the art using a computerized user interface, but more typically, identifying multiple timing-critical source / sink paths between pairs of source / sink endpoints in the data structure representing the initial configuration involves using a computerized timing estimation routine (timer 1503).

[0047] In one or more embodiments, using the computerized timing estimation routine includes using the computerized timing estimation routine to obtain results that are relatively accurate between source / sink endpoint pairs. As discussed elsewhere herein, "relatively accurate" has the express meaning that source / sink endpoint pairs are accurately ranked from worst to best, but the actual arrival time / slack values ​​are not necessarily accurate.

[0048] Referring to 931 in FIG. 9 , in one or more embodiments, performing a modified wirelength-driven placement on the updated netlist to obtain a data structure representing the modified placement includes moving pairs of endpoints closer to each other.

[0049] In some cases, performing a modified wirelength-driven placement on the updated netlist to obtain a data structure representing the modified placement includes reducing curvature of paths between pairs of endpoints (as seen in both 931 and 933). In some such cases 931, performing a modified wirelength-driven placement on the updated netlist to obtain a data structure representing the modified placement further includes moving pairs of endpoints closer together.

[0050] As shown at 933, pairs of endpoints are optionally fixed relative to each other when performing a modified wirelength-driven placement on the updated netlist to obtain a data structure representing the modified placement.

[0051] One or more embodiments further include retaining the pseudo 2-pin net during multiple full iterations, such that Weight_1_2 and Weight_12_99 are retained in the second iteration even though they are not critical in the second iteration, as in FIG. 14 .

[0052] One or more embodiments further include performing logic synthesis 103 to obtain the integrated circuit design embodied in the unplaced netlist, and, after multiple full iterations, performing virtual timing optimization 107, clock optimization 109, wiring synthesis and optimization 111, and routing 113 on the final version of the data structure representing the modified placement. The "final version" of the data structure may be, for example, from the penultimate iteration or the last iteration, depending on which iteration typically provides the best performance.

[0053] Referring to FIG. 8 (dotted arrow from 921 back to 905), one or more embodiments further include repeating the steps of identifying multiple timing-critical source / sink paths based on the routing results, creating new pseudo two-pin nets for each pair of endpoints, and performing modified length-driven placement for a new plurality of full iterations.

[0054] One or more embodiments further include fabricating a physical integrated circuit based on the routing 921, e.g., concluding step 921, preparing a layout and instantiating the layout as a design structure, according to which the physical integrated circuit can be fabricated. See Figures 11-13 discussed elsewhere herein.

[0055] One or more embodiments further include, at 911, clustering and fixing storage elements in a final version of the data structure representing the modified placement, and repeating the steps of identifying multiple timing-critical source / sink paths, creating new pseudo 2-pin nets for each pair of endpoints, and performing the modified wirelength-driven placement for a new plurality of full iterations after clustering and fixing storage elements (see 915; also note step 913 of removing attractions from 907 and 909). Note also that 911 includes physical clustering, also known as clock optimization or clock synthesis.

[0056] Furthermore, consider "congestion-driven" spread. Placement algorithms, often referred to as physical synthesis flows with objectives other than simply improving HPWL, use placement engines to ensure minimal degradation of HPWL. Advantageously, using aspects of the present invention in such cases further enhances these algorithms to minimize timing degradation.

[0057] With buffering and reasonable metal layer assumptions, timing information early in the physical synthesis flow is usually quite inaccurate. Those skilled in the art will recognize that appropriate steps should be taken to ensure critical paths are obtained. Long wires may need to include buffers that repeat and strengthen the signal. The best buffering / repeater solution for long wires cannot be determined until the terminal locations are known, but the terminals are not known where they are until some knowledge of the criticality of that particular net exists. Timing information is desired to perform placement, but accurate timing requires knowing where they will be placed and how they will be optimized. Because this problem is difficult to solve, time-of-flight calculations can be performed based on specific techniques, providing an approximation of the delay given the distance between the terminals (source and sink of a net) when the connections are properly optimized. If the path is short, it may be left alone. If the path is long, several buffers are assumed. The cost of time-of-flight calculations can be estimated for a particular metal layer. The time-of-flight includes the propagation delay as well as the delay associated with intermediate buffer / repetition stages. Those skilled in the art are familiar with the problem of running timing early in the design process when the locations of all circuit elements are not yet known, and recognize that assumptions must be made about buffer insertion and metal layers. Higher metal layers allow for wider wiring. Paths that are critical when routed on lower layers may not be critical at all when routed on higher layers. It can be unrealistic to assume that all connections are on the lowest metal layer. Techniques can be used to estimate which routers will be used given a two-pin net (source and sink), how the router will route the wires between these terminals, and what optimizations will determine whether several repeaters need to be inserted along the wires. This is a known problem whenever timing is run early in the design process (hypothetical timing, timing estimation).Given the teachings herein, one skilled in the art will be able to implement suitable buffering and reasonable metal layer assumptions.

[0058] One or more embodiments provide a method for improving timing performance of an electronic circuit designed using electronic design automation, the method including: obtaining, from a computerized placer 1501, initial wirelength-driven placement results for an integrated circuit design embodied in an unplaced netlist, the results including a data structure representing an initial placement of logic gates; obtaining, from a computerized timer 1503, at least one timing-critical source / sink path between at least one pair of source / sink endpoints in the data structure representing the initial placement; creating (e.g., by generator 1505) a new pseudo two-pin net for each of the at least one pair of source / sink endpoints to create an updated netlist; and providing the updated netlist to a computerized placer to generate a data structure representing the modified placement, the data structure representing the modified placement, to facilitate the computerized placer performing a modified wirelength-driven placement on the updated netlist.

[0059] In one or more embodiments, the resulting layout is instantiated as a design structure. See the discussion of FIG. 11. A physical integrated circuit is then fabricated according to the design structure. See again the discussion of FIG. 11. See also FIG. 12. Once physical design data is obtained based in part on the analysis process described herein, integrated circuits designed accordingly can be manufactured according to known processes generally described with reference to FIG. 12. Typically, wafers with multiple copies of the final design are fabricated and cut (i.e., diced) so that each die is one copy of the integrated circuit. In block 1110, the process includes fabricating a mask for lithography based on the final physical layout. In block 1120, fabricating the wafer includes using the mask to perform photolithography and etching. Once the wafer is diced, testing and sorting each die is performed to filter out any defective die at 1130.

[0060] One or more embodiments include a computer including a memory 28 and at least one processor 16 coupled to the memory and operative to perform or otherwise facilitate any one, some, or all of the method steps described herein (as shown in FIG. 10). In one or more embodiments, the amount of computer resources / CPU time and ergonomic design time required during a design cycle can be reduced using aspects of the present invention. Alternatively, a different, better chip (e.g., in terms of timing performance) can be designed with the same resources and engineering design time.

[0061] 11 , in one or more embodiments, at least one processor operates to generate a design structure for the circuit design, and in at least some embodiments, the at least one processor further operates to control integrated circuit manufacturing equipment to fabricate a physical integrated circuit in accordance with the design structure. Thus, the layout can be instantiated as a design structure, and the design structure can be provided to the fabrication equipment to facilitate fabrication of the physical integrated circuit in accordance with the design structure. The physical integrated circuit will be improved (e.g., by improved placement / timing performance compared to a design that does not use aspects of the present invention for EDA).

[0062] FIG. 13 shows the flow of an exemplary high-level electronic design automation (EDA) tool responsible for creating an optimized microprocessor (or other IC) design to be manufactured. A designer can start with a high-level logic description 1201 of the circuit (e.g., VHDL or Verilog). A logic synthesis tool 1203 compiles the logic and optimizes it with estimated timing information, without any sense of this physical representation. A placement tool 1205 takes the logic description and places each component, focusing on minimizing congestion in each area of ​​the design. A clock synthesis tool 1207 optimizes the clock tree network by cloning / balancing / buffering latches or registers. A timing closure step 1209 performs several optimizations on the design, including buffering, wiring tuning, and circuit repowering; the goal is to produce a design that can be routed without timing violations and excessive power consumption. The routing stage 1211 takes the placed / optimized design and determines how to create wiring that connects all of the components without causing manufacturing violations. Post-routing timing closure 1213 performs another set of optimizations to resolve any violations remaining after routing. Design refinement 1215 then adds additional metal shapes to the netlist to comply with manufacturing requirements. A check step 1217 analyzes whether the design violates any requirements, such as manufacturing, timing, power, electromigration (e.g., using techniques disclosed herein), or noise. If the design is complete, the final step 1219 is to generate a layout for the design that represents all of the shapes to be fabricated in the to-be-fabricated design 1221.

[0063] One or more embodiments of the present invention or elements thereof may be implemented in the form of an apparatus including a memory and at least one processor coupled to the memory and operative to perform exemplary method steps. Figure 10 illustrates a computer system that may be useful in implementing one or more aspects and / or elements of the present invention, referred to herein as a cloud computing node, but which may also be representative of a server, general-purpose computer, etc., that may be provided in a cloud or locally.

[0064] In cloud computing node 10, there are computer systems / servers 12 that operate in numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, or configurations, or combinations thereof, that may be suitable for use with computer system / server 12 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices.

[0065] The computer system / server 12 may be described in the general context of computer system-executable instructions, such as program modules, executed by a computer system. Generally, program modules may include routines, programs, objects, components, logic, and data structures that perform particular tasks or implement particular abstract data types. The computer system / server 12 may be practiced in a distributed cloud computing environment where tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media, including memory storage devices.

[0066] 10, computer system / server 12 in cloud computing node 10 is shown in the form of a general-purpose computing device. Components of computer system / server 12 may include, but are not limited to, one or more processors or processing units 16, system memory 28, and a bus 18 coupling various system components, including system memory 28, to processor 16.

[0067] Bus 18 represents any one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor bus or local bus using any of a variety of bus architectures. By way of example and not limitation, such architectures include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0068] Computer system / server 12 typically includes a variety of computer system-readable media, which may be any available media that can be accessed by computer system / server 12 and includes both volatile and nonvolatile media, removable and non-removable media.

[0069] System memory 28 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Computer system / server 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be provided for reading from and writing to non-removable, non-volatile magnetic media (not shown, but commonly referred to as a "hard drive"). Although not shown, a magnetic disk drive may be provided for reading from and writing to removable, non-volatile magnetic disks (e.g., "floppy disks"), and an optical disk drive may be provided for reading from or writing to removable, non-volatile optical disks, such as CD-ROMs, DVD-ROMs, or other optical media. In such cases, each disk drive may be connected to bus 18 by one or more data media interfaces. As further shown and described below, memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of embodiments of the present invention.

[0070] A program / utility 40 having a set (at least one) of program modules 42 may be stored in memory 28, by way of example and not limitation, as well as an operating system, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data, or any combination thereof, may include an implementation of a networking environment. The program modules 42 generally perform the functions and / or methodologies of embodiments of the present invention described herein.

[0071] The computer system / server 12 may also communicate with one or more external devices 14, such as a keyboard, pointing device, or display 24, one or more devices that allow a user to interact with the computer system / server 12, or any device (e.g., a network card, modem, etc.) that allows the computer system / server 12 to communicate with one or more other computing devices, or a combination thereof. Such communication may occur via an input / output (I / O) interface 22. Furthermore, the computer system / server 12 may communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof, via a network adapter 20. As shown, the network adapter 20 communicates with other components of the computer system / server 12 via a bus 18. Although not shown, it should be understood that other hardware and / or software components may be used in conjunction with the computer system / server 12. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems.

[0072] Thus, one or more embodiments may utilize software running on a general-purpose computer or workstation. Referring to FIG. 10 , such an implementation may employ, for example, a processor 16, memory 28, and a display 24 and an input / output interface 22 to external devices 14, such as a keyboard or pointing device. As used herein, the term “processor” is intended to include any processing device, such as one that includes a CPU (central processing unit) or other forms of processing circuitry, or both. Furthermore, the term “processor” may refer to multiple individual processors. The term “memory” is intended to include memory associated with a processor or CPU, such as, for example, RAM (random access memory) 30, ROM (read-only memory), fixed memory devices (e.g., hard drive 34), removable memory devices (e.g., diskettes), and flash memory. Furthermore, as used herein, the phrase “input / output interface” is intended to contemplate, for example, an interface to one or more mechanisms for inputting data into a processing unit (e.g., a mouse) and one or more mechanisms for providing results associated with the processing unit (e.g., a printer). The processor 16, memory 28, and input / output interface 22 may be interconnected, for example, via a bus 18 as part of the data processing unit 12. Suitable interconnections, for example via the bus 18, may also be provided to a network interface 20, such as a network card, which may be provided to interface with a computer network, and to a media interface, such as a diskette or CD-ROM drive, which may be provided to interface with suitable media.

[0073] Thus, computer software containing instructions or code for carrying out the techniques of the present invention may be stored in one or more of the associated memory devices (e.g., ROM, fixed, or removable memory) as described herein, and may be loaded partially or wholly (e.g., into RAM) and implemented by a CPU when ready for use. Such software may include, but is not limited to, firmware, resident software, microcode, and the like.

[0074] A data processing system suitable for storing and / or executing program code will include at least one processor 16 coupled directly or indirectly via a system bus 18 to a memory element 28. The memory element may include local memory used during the actual implementation of the program code, bulk storage, and cache memory 32 that provides temporary storage of at least some program code to reduce the number of times the code must be retrieved from bulk storage during implementation.

[0075] Input / output devices, or I / O devices (including but not limited to keyboards, displays, and pointing devices) can be coupled to the system either directly or through intervening I / O controllers.

[0076] Network adapters 20 may also be coupled to the system to enable the data processing system to become coupled to other data processing systems or remote printers or storage devices through intervening private or public networks. Modems, cable modems, and Ethernet cards are just a few of the commercially available types of network adapters.

[0077] As used herein, including in the claims, "server" includes a physical data processing system (e.g., system 12 as shown in FIG. 10) running a server program. It will be understood that such a physical server may or may not include a display and keyboard.

[0078] It should be noted that any of the methods described herein may include the additional step of providing a system including separate software modules embodied on a computer-readable storage medium, where the modules may include, for example, any or all of the appropriate elements shown in block diagrams and / or described herein, including, by way of example and not limitation, any one, some, or all of the modules / blocks and / or sub-modules / sub-blocks illustrated (e.g., in FIG. 15). Method steps may also be performed using separate software modules and / or sub-modules of the above-described systems running on one or more hardware processors, such as 16. Furthermore, a computer program product may include a computer-readable storage medium having code adapted to be implemented to perform one or more method steps described herein, including providing a system having separate software modules.

[0079] One example of a user interface that may be used is Hypertext Markup Language (HTML) code served by a server or the like to a browser on a user's computing device, which is parsed by the browser on the user's computing device to create a graphical user interface (GUI).

[0080] Exemplary System and Product Details

[0081] The present invention may be a system, method, or computer program product, or a combination thereof, at any possible level of technical detail integration. The computer program product may include a computer-readable storage medium or media having computer-readable program instructions for causing a processor to perform aspects of the present invention.

[0082] A computer-readable storage medium may be a tangible device capable of holding and storing instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or ridge structures in grooves with instructions recorded on them, and any suitable combination of the foregoing. Computer-readable storage medium, as used herein, should not be construed as being a transitory signal per se, such as an electric wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse passing through a fiber optic cable), or an electrical signal transmitted through a wire.

[0083] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to a computer-readable storage medium within each computing / processing device for storage.

[0084] Computer-readable program instructions for carrying out the operations of the present invention may be either source or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or object-oriented programming languages ​​such as Smalltalk® or C++, and procedural programming languages ​​such as the “C” programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or a connection to an external computer may be made (e.g., through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer readable program instructions to personalize the electronic circuitry by utilizing state information of the computer readable program instructions to perform aspects of the present invention.

[0085] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0086] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, the instructions executing on the processor of which cause the computer or other programmable data processing apparatus to cause the machine to implement the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored on a computer-readable storage medium capable of directing a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner, such that the computer-readable storage medium on which the instructions are stored comprises an article of manufacture containing instructions implementing aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.

[0087] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device, causing the computer, other programmable apparatus, or other device to perform a series of operational steps, resulting in a computer-implemented process, such that the instructions, which execute on a computer, other programmable data processing apparatus, or other device, perform the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams.

[0088] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions that implement a specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a special-purpose hardware-based system that performs the specified functions or operations or executes a combination of special-purpose hardware and computer instructions.

[0089] Exemplary design processes used in semiconductor design, manufacturing, and / or testing

[0090] In one or more embodiments, the TDA techniques herein are integrated with semiconductor integrated circuit design simulation, testing, layout, and / or manufacturing. In this regard, FIG. 11 illustrates a block diagram of an exemplary design flow 700 used, for example, in semiconductor IC logic design, simulation, testing, layout, and manufacturing. The design flow 700 includes a process, machine, or mechanism, or a combination thereof, for processing a design structure or device to generate a logically or otherwise functionally equivalent representation of the design structure or device or both, such as one that can be analyzed using the techniques disclosed herein. The design structure processed and / or generated by the design flow 700 may be encoded on a machine-readable storage medium to include data and / or instructions that, when executed or otherwise processed on a data processing system, generate a logically, structurally, mechanically, or otherwise functionally equivalent representation of a hardware component, circuit, device, or system. The machine includes, but is not limited to, any machine used in an IC design process, such as the design, manufacturing, or simulation of a circuit, component, device, or system. For example, a machine may include a lithography machine, a machine and / or equipment for generating a mask (e.g., an electron beam writer), a computer or equipment for simulating a design structure, any apparatus used in a manufacturing or testing process, or any machine for programming a functionally equivalent representation of a design structure into any medium (e.g., a machine for programming a programmable gate array).

[0091] The design flow 700 may differ depending on the type of representation being designed. For example, a design flow 700 for building an application-specific integrated circuit (ASIC) may differ from a design flow 700 for designing a standard component, or from a design flow 700 for instantiating a design into a programmable array, such as a programmable gate array (PGA) or field programmable gate array (FPGA) offered by Altera® or Xilinx®.

[0092] FIG. 11 illustrates multiple such design structures, preferably including an input design structure 720, processed by the design process 710. The design structure 720 may be a logic simulation design structure generated and processed by the design process 710 to produce a logically equivalent functional representation of a hardware device. The design structure 720 may also or alternatively include data and / or program instructions that, when processed by the design process 710, generate a functional representation of the physical structure of the hardware device. Whether representing functional and / or structural design features, the design structure 720 may be generated using electronic computer-aided design (ECAD), such as implemented by a core developer / designer. Once encoded, such as on a gate array or storage medium, the design structure 720 may be accessed and processed by one or more hardware and / or software modules in the design process 710 to simulate or otherwise functionally represent an electronic component, circuit, electronic or logic module, apparatus, device, or system. As such, design structure 720 may include files or other data structures, including human- and / or machine-readable source code, compiled structures, and computer-executable code structures, that, when processed by a design or simulation data processing system, functionally simulate or otherwise represent a circuit or other level hardware logic design. Such data structures may include hardware description language (HDL) design entities or other data structures that conform to and / or are compatible with low-level HDL design languages, such as Verilog and VHDL, and / or high-level design languages, such as C or C++.

[0093] Design process 710 preferably uses and incorporates hardware and / or software modules for synthesizing, translating, or otherwise manipulating design / simulation functional equivalents of components, circuits, devices, or logic structures to generate netlist 780, which may contain design structures such as design structure 720. Netlist 780 may include, for example, compiled or otherwise manipulated data structures representing lists of wires, discrete components, logic gates, control circuits, I / O devices, models, etc., that describe connections to other elements and circuits in an integrated circuit design. Netlist 780 may be synthesized using an iterative process in which netlist 780 is resynthesized one or more times depending on device design specifications and parameters. As with the other design structure types described herein, netlist 780 may be recorded on a machine-readable data storage medium or programmed into a programmable gate array. The medium may be a non-volatile storage medium such as a magnetic or optical disk drive, a programmable gate array, compact flash, or other flash memory. Additionally or alternatively, the medium may be system or cache memory, a buffer area, or other suitable memory.

[0094] Design process 710 may include hardware and software modules for processing various input data structure types, including netlist 780. Such data structure types may reside, for example, in library elements 730 and include sets of commonly used elements, circuits, and devices, including models, layouts, and symbolic representations for a given manufacturing technology (e.g., different technology nodes, such as 32 nm, 45 nm, 90 nm, etc.). Data structure types may further include design specifications 740, characterization data 750, verification data 760, design rules 770, and test data files 785, which may include input test patterns, output test results, and other test information. Design process 710 may also include standard mechanical design processes, such as stress analysis, thermal analysis, mechanical event simulation, and process simulation of operations such as cast, mold, and die press formation. Those skilled in the art of mechanical design will recognize the range of possible mechanical design tools and applications that may be used in design process 710 without departing from the scope and spirit of the present invention. The design process 710 may also include modules for performing standard circuit design processes such as timing analysis, verification, design rule checking, place and route operations, etc. Path-based timing-driven placement using iterative modification of a pseudo netlist can be performed as described herein.

[0095] Design process 710 uses and incorporates logical and physical design tools, such as HDL compilers and simulation model building tools, to process design structure 720 along with some or all of the illustrated supporting data structures, in addition to any additional mechanical design or data (if applicable), to generate second design structure 790. Design structure 790 resides on a storage medium or programmable gate array in a data format used for the exchange of mechanical device and structure data (e.g., information stored in IGES, DXF, Parasolid XT, JT, DRG, or any other suitable format for storing or rendering such mechanical design structures). Like design structure 720, design structure 790 preferably resides on a data storage medium and includes one or more files, data structures, or other computer-encoded data or instructions that, when processed by an ECAD system, generate a logically or otherwise functionally equivalent form, such as one or more IC designs. In one embodiment, design structure 790 may include a compiled, executable HDL simulation model that functionally simulates the device being analyzed.

[0096] Design structure 790 may also use a data format used for the exchange of integrated circuit layout data and / or a symbolic data format (e.g., information stored in GDSII (GDS2), GL1, OASIS, map files, or any other suitable format for storing such design data structures). Design structure 790 may include information such as, for example, symbolic data, map files, test data files, design content files, manufacturing data, layout parameters, wires, metal levels, vias, shapes, data for routing manufacturing lines, and any other data (e.g., lib files) required by a manufacturer or other designer / developer to produce a device or structure as described herein. Design structure 790 may then proceed to stage 795, where, for example, design structure 790 may proceed to tape-out, be released for manufacturing, be released to a mask house, be sent to another design house, be returned to a customer, etc.

[0097] The description of various embodiments of the present invention is presented for illustrative purposes and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations that do not depart from the scope and spirit of the described embodiments will be apparent to those skilled in the art. The terminology used herein has been chosen to best explain the principles of the embodiments, practical applications or technical improvements over technology found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.

Claims

1. 1. A method for improving timing performance of an electronic circuit designed using electronic design automation, comprising: performing an initial wirelength-driven placement for the integrated circuit design embodied in the unplaced netlist using a computerized placer to obtain a data structure representing an initial placement of logic gates; identifying at least one timing-critical source / sink path between at least one pair of source / sink endpoints in the data structure representing the initial configuration; creating a new pseudo two-pin net for each pair of the at least one pair of source / sink endpoints to create an updated netlist, the new pseudo two-pin net lacking a corresponding electrical signal of an electronic circuit; performing a modified wirelength-driven placement on the updated netlist to obtain a data structure representing a modified placement; A method comprising:

2. 2. The method of claim 1 , wherein the at least one timing-critical source / sink path comprises a plurality of timing-critical source / sink paths, and the at least one pair of source / sink endpoints comprises a plurality of pairs of source / sink endpoints.

3. 3. The method of claim 2, further comprising: repeating the steps of identifying the plurality of timing-critical source / sink paths, creating the new pseudo two-pin net for each of the plurality of pairs of endpoints, and performing the modified wire-length-driven placement for a plurality of full iterations.

4. 4. The method of claim 3, wherein the steps of performing the initial wirelength-driven placement and performing the modified wirelength-driven placement each include applying a half-circle approximate wirelength-driven placement, and the computerized placer includes a half-circle approximate wirelength-driven computerized placer.

5. 4. The method of claim 3, wherein identifying the plurality of timing-critical source / sink paths between the pairs of source / sink endpoints in the data structure representing the initial placement comprises obtaining input from an expert in the art using a computerized user interface.

6. 4. The method of claim 3, wherein identifying the plurality of timing-critical source / sink paths between pairs of source / sink endpoints in the data structure representing the initial placement comprises using a computerized timing estimation routine.

7. 7. The method of claim 6, wherein using the computerized timing estimation routine includes using the computerized timing estimation routine to obtain relatively accurate results between the source / sink endpoint pair.

8. 4. The method of claim 3, wherein performing the modified wirelength-driven placement on the updated netlist to obtain the data structure representing a modified placement comprises moving pairs of the endpoints closer to each other.

9. 4. The method of claim 3, wherein performing the modified length-driven placement on the updated netlist to obtain the data structure representing a modified placement includes reducing curved portions of paths between the pair of endpoints.

10. 10. The method of claim 9, wherein performing the modified wirelength-driven placement on the updated netlist to obtain the data structure representing a modified placement further comprises moving pairs of the endpoints closer to each other.

11. 10. The method of claim 9, wherein the pairs of endpoints are fixed relative to each other when performing the modified wirelength-driven placement on the updated netlist to obtain the data structure representing a modified placement.

12. 4. The method of claim 3, further comprising the step of maintaining said pseudo 2-pin net during said plurality of full iterations.

13. performing logic synthesis to obtain the integrated circuit design embodied in the unplaced netlist; after said plurality of full iterations, performing virtual timing optimization, clock optimization, wiring synthesis and optimization, and routing on a final version of said data structure representing said modified placement; The method of claim 3 further comprising:

14. 14. The method of claim 13, further comprising: repeating the steps of identifying the plurality of timing-critical source / sink paths based on the routing results, creating the new pseudo two-pin net for each pair of the endpoints, and performing the modified wire-length-driven placement for a new plurality of full iterations.

15. The method of claim 13 further comprising fabricating a physical integrated circuit based on the routing.

16. clustering and fixing storage elements in a final version of the data structure representing the modified arrangement; performing a further modified wirelength-driven placement based on the clustering and fixing step; repeating the steps of identifying the plurality of timing-critical source / sink paths, creating the new pseudo 2-pin net for each pair of the endpoints, and performing the modified wirelength-driven placement for a new plurality of full iterations after the further modified wirelength-driven placement based on the clustering and fixing step; The method of claim 3 further comprising:

17. Memory and at least one processor coupled to the memory; wherein the at least one processor: performing an initial wirelength-driven placement of the integrated circuit design embodied in the unplaced netlist using a computerized placer to obtain a data structure representing an initial placement of logic gates; identifying at least one timing-critical source / sink path between at least one pair of source / sink endpoints in the data structure representing the initial configuration; creating a new pseudo two-pin net for each pair of the at least one pair of source / sink endpoints to create an updated netlist, the new pseudo two-pin net lacking a corresponding electrical signal of an electronic circuit; performing a modified wirelength-driven placement on the updated netlist to obtain a data structure representing a modified placement; to improve timing performance of an electronic circuit designed using electronic design automation.

18. 18. The computer of claim 17, wherein the at least one timing-critical source / sink path includes a plurality of timing-critical source / sink paths, and the at least one pair of source / sink endpoints includes a plurality of pairs of source / sink endpoints, and the at least one processor is further operative to repeat the steps of identifying the plurality of timing-critical source / sink paths, creating the new pseudo two-pin net for each pair of the plurality of pairs of endpoints, and performing the modified wire-length-driven placement for a plurality of full iterations.

19. The at least one processor after said plurality of full iterations, performing virtual timing optimization, clock optimization, wiring synthesis and optimization, and routing on a final version of said data structure representing said modified placement; preparing a layout based on the routing; instantiating the layout as a design structure; providing the design structure to fabrication equipment to facilitate fabrication of a physical integrated circuit in accordance with the design structure; 20. The computer of claim 18, further operative to:

20. One or more computer-readable storage media, the one or more computer-readable storage media comprising: first program instructions executable by a computer system to cause the computer system to perform an initial wirelength-driven placement for an integrated circuit design embodied in an unplaced netlist using a computerized placer to obtain a data structure representing an initial placement of logic gates; second program instructions executable by the computer system to cause the computer system to identify at least one timing-critical source / sink path between at least one pair of source / sink endpoints in the data structure representing the initial configuration; and third program instructions executable by the computer system to cause the computer system to create a new pseudo two-pin net for each pair of the at least one pair of source / sink endpoints to create an updated netlist, the new pseudo two-pin net lacking a corresponding electrical signal of an electronic circuit; and fourth program instructions executable by the computer system to cause the computer system to perform a modified wirelength-driven placement on the updated netlist to obtain a data structure representing a modified placement; and A computer-readable storage medium having stored thereon.

21. 21. The computer-readable storage medium of claim 20, further comprising fifth program instructions executable by the computer system to cause the at least one timing-critical source / sink path to include a plurality of timing-critical source / sink paths, the at least one pair of source / sink endpoints including a plurality of pairs of source / sink endpoints, and to cause the computer system to repeat the steps of identifying the plurality of timing-critical source / sink paths, creating the new pseudo two-pin net for each of the plurality of pairs of endpoints, and performing the modified wire-length-driven placement for a plurality of full iterations.

22. sixth program instructions executable by the computer system to cause the computer system to perform virtual timing optimization, clock optimization, wiring synthesis and optimization, and routing on a final version of the data structure representing the modified placement after the plurality of full iterations; seventh program instructions executable by the computer system to cause the computer system to prepare a layout based on the routing; and eighth program instructions executable by the computer system to cause the computer system to instantiate the layout as a design structure; and ninth program instructions executable by the computer system to cause the computer system to provide the design structure to fabrication equipment to facilitate fabrication of a physical integrated circuit in accordance with the design structure; 22. The computer-readable storage medium of claim 21, further comprising:

23. 1. A method for improving timing performance of an electronic circuit designed using electronic design automation, comprising: obtaining initial wirelength-driven placement results from a computerized placer for the integrated circuit design embodied in an unplaced netlist, said results including a data structure representing an initial placement of logic gates; obtaining, from a computerized timer, at least one timing-critical source / sink path between at least one pair of source / sink endpoints in the data structure representing the initial configuration; creating a new pseudo 2-pin net for each pair of the at least one pair of source / sink endpoints to create an updated netlist, the new pseudo 2-pin net lacking a corresponding electrical signal of an electronic circuit; providing the updated netlist to the computerized placer to facilitate the computerized placer performing a modified wirelength-driven placement on the updated netlist to produce a data structure representing the modified placement; A method comprising:

24. 24. The method of claim 23, wherein the at least one timing-critical source / sink path comprises a plurality of timing-critical source / sink paths, and the at least one pair of source / sink endpoints comprises a plurality of pairs of source / sink endpoints.

25. 25. The method of claim 24, further comprising: repeating the steps of obtaining the plurality of timing-critical source / sink paths, creating the new pseudo 2-pin net for each pair of the plurality of pairs of endpoints, and providing the updated netlist for a plurality of full iterations.

Citation Information

Patent Citations

  • Method for designing semiconductor integrated circuit device

    JP1996077219A

  • Incremental timing optimization and placement

    US20100257498A1

  • Method and apparatus for implementing engineering change orders in integrated circuit designs

    US5953236A