Method for constructing a hierarchical clock tree for an integrated circuit

By building a hierarchical clock tree, existing tools solve the problems of long design time, waste of resources and high power consumption in adjacent designs, and achieve faster timing convergence and smaller design area, supporting multi-layer parallel processing and flexible clock tree planning.

CN112100971BActive Publication Date: 2025-07-04SAMSUNG ELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010534261.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-28
Filing Date
2020-06-11
Publication Date
2025-07-04
Estimated Expiration
2040-06-11

AI Technical Summary

Technical Problem

Existing industry-standard EDA tools lack efficient hierarchical clock implementation methods, especially for adjacency designs, resulting in increased design time, waste of resources, high power consumption and large design area.

Method used

By building a layered clock tree in an integrated circuit, including building a clock distribution network at the first layer, pushing it to the second layer to implement a partitioned clock tree, and calculating the combined timing, adjusting the timing using engineering change commands and target constraints, supporting multi-layer parallel balance, and decoupling the dependence between the clock distribution network and the partitioned clock tree.

Benefits of technology

Faster timing convergence, lower power consumption and smaller design area, providing full-chip view and flexible clock tree planning, supporting adjacency design and channel-based parallel processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112100971B_ABST
    Figure CN112100971B_ABST
Patent Text Reader

Abstract

A method for constructing a hierarchical clock tree for an integrated circuit is disclosed. The method for constructing a hierarchical clock tree for an integrated circuit may include: constructing a clock distribution network on a first layer; pushing the clock distribution network to a second layer; implementing a partitioned clock tree in a partition on the second layer; and calculating a combined timing of the clock distribution network and the partitioned clock tree on the second layer. The step of implementing the partitioned clock tree may include: constructing a partitioned clock tree in a partition on the second layer; calculating a test timing of the partitioned clock tree; calculating a target timing constraint of the partitioned clock tree based on the timing of the clock distribution network and the test timing of the partitioned clock tree; and adjusting the timing of one or more of the partitioned clock trees based on the target constraint.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority and benefit of U.S. Provisional Patent Application No. 62 / 863,259, titled "Methods and Apparatus for Hierarchical Clock Implementation for Abutted Design," filed on Jun. 18, 2019, and U.S. Patent Application No. 16 / 666,389, filed on Oct. 28, 2019, which are hereby incorporated by reference. Technical Field

[0002] The present disclosure generally relates to clock trees for integrated circuits, and more particularly, to hierarchical clock tree implementation. Background Art

[0003] Hierarchical design is widely used in very large scale integration (VLSI) to design highly complex integrated circuits (ICs). Hierarchical design generally involves decomposing a complex design into smaller physical blocks that can be more easily designed individually, and then combining the blocks into a larger overall design. The blocks in a hierarchical design are typically arranged in a channel-based design or an abutted design. In a channel-based design, the blocks are separated by channels through which clocks and other signals are distributed to the blocks. In an abutted or channel-less design, the blocks are placed adjacent to each other with no space between the blocks.

[0004] A clock tree is used to distribute clock signals throughout an integrated circuit. The clock tree is designed through a process that seeks to minimize delay, which is the delay from the root clock to the point of use, and skew, which is the difference in the arrival times of clock transitions at different points on the integrated circuit. During the design process, many parameters of the clock tree are typically adjusted through multiple iterations to meet the timing requirements and constraints of the clock tree. When the targets for the timing requirements and constraints have been met, the design is said to have achieved timing closure. VLSI designs are performed on industry-standard electronic design automation (EDA) tools that typically have automated workflows for many of the routine tasks performed by a designer. However, industry-standard EDA tools do not have methods or workflows for efficient hierarchical clock implementation, particularly for abutted designs. Summary of the Invention

[0005] The object of the present disclosure is to provide a hierarchical clock implementation that saves design time and / or resources and reduces power consumption and / or design area.

[0006] A method for constructing a hierarchical clock tree for an integrated circuit may include: constructing a clock distribution network on a first layer; pushing the clock distribution network to a second layer; implementing a partitioned clock tree in a partition on the second layer; and calculating a combined timing of the clock distribution network and the partitioned clock tree on the second layer. The step of implementing the partitioned clock tree may include: constructing a partitioned clock tree in a partition on the second layer; calculating a test timing of the partitioned clock tree; calculating a target timing constraint of the partitioned clock tree based on the timing of the clock distribution network and the test timing of the partitioned clock tree; and adjusting the timing of one or more in the partitioned clock tree based on the target constraint. The step of calculating the combined timing of the clock distribution network and the partitioned clock tree on the second layer may include merging the partitioned clock trees. The method may further include: adjusting the timing of one or more in the partitioned clock tree on the second layer. The timing of one or more in the partitioned clock tree on the second layer may be adjusted by an engineering change order (ECO). The timing of one or more in the partitioned clock tree may be adjusted by adjusting one or more target constraints of one or more in the partitioned clock tree on the second layer. The method may further include: determining that a timing target is not met by adjusting the timing of one or more in the partitioned clock tree and / or adjusting the clock tree distribution network. The method may further include: balancing the clock distribution network and balancing the one or more partitioned clock trees in parallel. The method may further include: pushing the clock distribution network to a third layer; and implementing a partitioned clock tree in a partition on the third layer. The clock distribution network may be pushed to one or more in a partition on the second layer. The second layer may include a block layer. The second layer may be lower than the first layer.

[0007] A method for constructing a hierarchical clock tree for an integrated circuit may include: constructing a clock distribution network on a first layer; pushing the clock distribution network to a partition at a second layer; calculating a test timing of the partition at the second layer; calculating a combined timing of the clock distribution network and the test timing of the partition at the second layer; calculating a partition layer target constraint based on the combined timing of the clock distribution network and the test timing of the partition at the second layer; and calculating a modified timing at the partition layer based on the target constraint. The method may further include: merging partitions at the partition layer; calculating a modified combined timing of the clock distribution network and the modified timing of the partition at the second layer; and checking whether the modified combined timing meets a design target. The method may further include: balancing the hierarchical clock tree by adjusting the modified timing of the partition at the second layer. The timing at the second layer may be adjusted by an engineering change order (ECO). The timing of the second layer may be adjusted by adjusting the target constraint at the second layer. The timing may include delay. The timing may include skew. The partition may include adjacent blocks. The second layer may include channel-based blocks. The partition may include multi-instantiated modules (MIMs). The dependency between the clock distribution network and the partition on the second layer may be decoupled. The method may further include: balancing the clock distribution network and the partition at the second layer in parallel.

[0008] A method of constructing a clock tree for an integrated circuit may include: constructing a top-level clock distribution network; calculating the distribution delay to the endpoints of the clock distribution network; pushing the top-level clock distribution network down to the block level; constructing a clock tree in the blocks at the endpoints; calculating the block-level insertion delay of the clock tree in the blocks at the endpoints; combining the distribution delay and the block-level insertion delay to calculate the clock tree insertion delay from the root of the top-level clock distribution network; calculating the delay target constraint for the blocks based on the clock tree insertion delay from the root of the top-level clock distribution network; recalculating the block-level insertion delay based on the delay target constraint; merging the clock trees at the block level; and recalculating the clock tree insertion delay from the root of the top-level clock distribution network. The method may further include: checking whether the recalculated clock tree insertion delay from the root of the top-level clock distribution network meets the design goal.

[0009] The method may further include: determining that the insertion delay from the root of the top-level clock distribution network does not meet the design goal; and using an engineering change order (ECO) to change the clock cells on the clock tree. The method may further include: determining that the insertion delay from the root of the top-level clock distribution network does not meet the design goal; and recalculating the delay target constraint for the blocks. The step of pushing the top-level clock distribution network down to the block level may include creating an ECO file for the sub-blocks. The step of pushing the top-level clock distribution network down to the block level may include creating a configuration file for the sub-blocks. The method may further include: determining that the sub-block layout is changed to push the top-level clock distribution network down to the block level; and modifying the top-level clock distribution network structure at the block level to save the distribution delay to the endpoints of the clock distribution network.

[0010] According to the disclosure, a method for implementing a hierarchical clock distribution network is provided. The network supports faster timing convergence, parallel balancing at different layers, more stringent clock timing control, etc. Thus, an implementation that saves design time and / or resources and reduces power consumption and / or design area is provided. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Throughout the drawings, for illustrative purposes, the drawings are not necessarily drawn to scale, and elements of similar structure or function are generally represented by the same reference numerals. The drawings are only intended to facilitate the description of the various embodiments described herein. The drawings do not depict every aspect of the teachings disclosed herein and do not limit the scope of the claims. The drawings, together with the specification, illustrate example embodiments of the present disclosure and are used, together with the description, to explain the principles of the present disclosure.

[0012] Figure 1 is a flowchart showing an embodiment of a method for constructing a hierarchical clock tree for an integrated circuit according to the present disclosure.

[0013] Figure 2is a plan view showing blocks arranged on an integrated circuit for use with an embodiment of a method for constructing a hierarchical clock tree for an integrated circuit.

[0014] Figure 3 is a plan view showing an embodiment of a clock distribution network constructed according to an embodiment of a method for constructing a hierarchical clock tree for an integrated circuit in accordance with the present disclosure.

[0015] Figure 4 is a plan view showing an embodiment of a clock distribution network and a partitioned clock tree pushed to another layer constructed according to an embodiment of a method for constructing a hierarchical clock tree for an integrated circuit in accordance with the present disclosure.

[0016] Figure 5A and Figure 5B together form a flowchart showing an example embodiment of a method for constructing a hierarchical clock tree for an integrated circuit in accordance with the present disclosure.

[0017] Figure 6 is a plan view showing an embodiment of a clock distribution network constructed according to a method for constructing a hierarchical clock tree for an integrated circuit in accordance with the present disclosure.

[0018] Figure 7 is a plan view showing a virtual load added to a clock distribution network for calculating clock timing according to a method for constructing a hierarchical clock tree for an integrated circuit in accordance with the present disclosure Figure 6 to.

[0019] Figure 8 is a plan view showing adjacent clock terminals of a clock distribution network added according to a method for constructing a hierarchical clock tree for an integrated circuit in accordance with the present disclosure Figure 6 to.

[0020] Figure 9 is a plan view showing a push - down clock cell and a clock spine segment in an enlarged portion of Figure 6 according to a method for constructing a hierarchical clock tree for an integrated circuit in accordance with the present disclosure.

[0021] Figure 10 shows an embodiment of a computing system in accordance with the present disclosure. Detailed Description

[0022] In a hierarchical design, a clock tree can be partitioned into a top-level or global part and a block-level or local part. The top-level or global part sends clock signals to different blocks or parts of an integrated circuit, and the block-level or local part propagates the clock signals to individual sequential elements that use the clock signals. To design a clock hierarchy based on a channel design, clock tree synthesis (CTS) is typically used to design the block-level clock tree for each block. Once the block-level clock trees are ready, the designer can create the top-level clock distribution network in the channels between the blocks, typically based on the clock tree insertion delays for each block and / or sub-block. Since the layout, routing, and balancing of the top-level clock tree may have to wait until all the clock trees for the blocks and / or sub-blocks are complete, this process can result in slower timing convergence. Additionally, compared to an abutment design, a channel-based floorplan is typically less efficient in terms of space (chip area) and power consumption.

[0023] Although an abutment hierarchical design is generally more efficient than a channel-based design, designing the clock distribution network for an abutment design can be very challenging. For example, it may be necessary to plan and / or design the top-level clock distribution network at the block level, which may require stitching together individual blocks that may not provide a good view of the overall clock distribution network. Additionally, due to differences in the sizes of different blocks and / or sub-blocks and / or the number of clock leaf cells, the insertion delays of the clock trees in a block or sub-block can vary. Furthermore, even for the same block, the designer may see different clock tree insertion delays for different runs due to different clock leaf cell layouts. This can lead to a need for many design iterations to balance the clock distribution network and the block-level trees, increasing the time required to achieve timing convergence. Therefore, CTS may not work for an abutment design for the top-level clock tree. Additionally, industrial standard tools do not have the ability to hierarchically push down clock trees to facilitate faster timing convergence. Therefore, only industrial standard tools can be used to create a channel-based design, which may result in increased chip area and greater power consumption.

[0024] Most high-frequency abutment designs typically use a clock mesh structure that can be implemented as a grid of metal traces driven by many clock drivers. Clock mesh structures typically achieve better skew and latency, but they also typically consume more power and / or chip area. The power consumption in a clock mesh design can be driven by the additional capacitance of the mesh structure and the clock gates that are pushed towards the leaf cells, which can result in less efficient clock gating. Additionally, clock mesh design and workflows can be more complex and time-consuming than clock tree synthesis. For example, clock mesh design can involve a large number of SPICE simulations, timing backannotation, and transition times.

[0025] Figure 1An embodiment of a method for constructing a hierarchical clock tree for an integrated circuit according to the present disclosure is shown. The method begins with step 100 having a partitioned floorplan as shown in Figure 2 which shows various physical partitions that can be arranged on an integrated circuit (IC) 113 or a portion of the IC 113. In an embodiment of Figure 2 , partitions 112A, 112B, 112C, and 112D are arranged in an adjacent design, while partitions 112E, 112F, 112G, and 112H are arranged in a channel-based design. The partitions can be implemented as, for example, blocks, sub-blocks, modules, regions, etc. Figure 2 In step 102 of

[0026] , as shown in Figure 1 , a clock distribution network 116 can be constructed on, for example, another layer above the partition layer. The clock distribution network 116 can distribute clock signals from a clock source or root 118 to endpoints 120A, 120B, 120C, 120D, 120E, 120F, 120G, and 120H (which can be collectively referred to as 120). For simplicity of illustration, only the routing of the clock distribution network 116 is shown, but it can also include buffers, gates, and any other clock units located at appropriate points in the network. Figure 3 In step 104 of

[0027] , in this case, the clock distribution network 116 can then be pushed down to the partition layer. In the portion of the floorplan having adjacent partitions 112A, 112B, 112C, and 112D, the clock distribution network 116 can be pushed into the adjacent partitions. In the portion of the floorplan having channel-based partitions 112E, 112F, 112G, and 112H, the clock distribution network 116 is mainly pushed into the channels between the partitions, where the clock distribution network 116 has endpoints located at the edges of the partitions. However, in some cases, for example, for partitions 112F and 112H, some endpoints and segments of the clock distribution network 116 can be pushed down into the partitions. Figure 1 In step 106 of

[0028] , as shown in Figure 1 , in step 106 of Figure 4As shown, the partitioned clock trees 122A-122H can be implemented within the partitions 112A-112H. However, the present invention is not limited thereto, and one or more partitioned clock trees can be implemented within each of one or more partitions. The endpoints 120 of the clock distribution network 116 can be used as the roots of the partitioned clock trees 122A-122H (which can be collectively referred to as 122). In some embodiments, the partitioned clock trees 122 can be implemented by constructing the partitioned clock trees and calculating the clock timing of the partitioned clock trees. In other embodiments, the partitioned clock trees 122 can be implemented by calculating the test timing and constraints at the partition level and by various additional techniques in accordance with the principles of the present disclosure described in more detail below. For simplicity of illustration, only the wiring of the partitioned clock trees 122 is shown, but they can also include buffers, gates, and any other clock units located at appropriate points in the tree.

[0029] Thus, a hierarchical clock tree can be constructed starting from the clock root 118 and passing through the clock distribution network 116, the endpoints 120, and the partitioned clock trees 122.

[0030] At Figure 1 step 108, after the hierarchical clock tree has been constructed as described above, the timing of the clock distribution network 116 can be combined with the timing of the partitioned clock trees 122 to calculate the overall hierarchical clock tree timing starting at the clock root 118. Herein, the timing can refer to any clock-related timing parameter (such as, by way of example, delay (or insertion delay) and skew). Additionally, any timing parameter can be evaluated at one or more corners including any or all relevant process, voltage, and temperature (PVT) corners. At Figure 1 step 110, the overall hierarchical clock tree timing can then be checked to see if the design goals are met. If not, the clock distribution layer and / or the partition layer can be adjusted by returning to any of steps 102, 104, and 106. If the overall hierarchical clock tree timing meets the design goals, the method can end at step 111. All or any part of the overall hierarchical clock tree including the wiring, buffers, and / or clock units can then be saved, for example, by locking them to prevent the EDA tools from modifying all or any part of the overall hierarchical clock tree during the implementation of the partition.

[0031] In accordance with the principles of the present disclosure, the above for Figures 1 to 4The described methods and structures can achieve a wide range of benefits, applications, as well as additional features and techniques. For example, the clock tree designer can be provided with a full-chip and / or multi-layer view of the clock tree planning that allows for more stringent control and / or increased flexibility in clock distribution and implementation. As another example, a non-grid hierarchical clock tree can be implemented in an adjacent design. The above methods and structures can also enable the designer to manually implement the clock distribution network in a manner that can lead to faster timing convergence independent of the EDA tools used. They can also work with adjacent designs and channel-based designs and accommodate multi-instance partitions (such as multi-instantiated modules (MIM) and multi-instantiated blocks (MIB)). The above methods and structures can also decouple the dependencies between the clock distribution network and the partitioned clock tree, thus enabling additional new technologies. For example, depending on the implementation details and specific circumstances, in the case of decoupled dependencies between layers, the clock distribution network may not require balancing, and balancing can be achieved through adjustments and / or budgets of constraints at the partitioned layers. This can in turn lead to minimizing or reducing the delay at the clock distribution network. As another example, for instance, once the delay budget is completed, decoupling the dependencies between the clock distribution network and the partitioned layer clock tree can enable the balancing operation to be performed in parallel on the clock distribution network and the partitioned layer clock tree. Any or all of the benefits, applications, features, and techniques described above can lead to faster timing convergence, lower power consumption, and smaller design area for individual components and / or the overall hierarchical clock tree.

[0032] The inventive principles disclosed in this patent are not limited to Figures 1 to 4 the details shown therein. For example, the clock distribution network 116 and the partitioned clock tree 122 are shown to have a simple tree structure as can be achieved using clock tree synthesis (CTS) and a single clock root. However, multiple clock roots and trees, as well as other topologies and arrangements (such as multi-source CTS (MSCTS), grid structures, or any hybrid combination thereof), can be used. For Figures 1 to 4 the embodiments shown can be implemented as an entire integrated circuit (IC), or only a portion of the IC. It can also be implemented as part of a larger hierarchical clock tree. Additionally, other embodiments can be implemented in which all or part of the clock distribution network 116 can be pushed across more than one layer.

[0033] Figure 5A 、 Figures 5B to 9 FIG. shows a more detailed exemplary embodiment of a method for constructing a hierarchical clock tree for an integrated circuit according to the present disclosure. For illustrative purposes, this embodiment is described in the context of a full-chip design, where the clock distribution network can be designed at the top layer and the partitioned layers can have adjacent blocks. However, the principles of the present disclosure are not limited to these or any other exemplary implementation details.

[0034] As an introductory overview, and as described in more detail below, a designer may begin with a floorplan having adjacent blocks that have been laid out during physical design. A clock designer may plan and construct a top-level clock distribution network to distribute clock signals from a top-level root clock to individual blocks. The insertion delay from the root clock to the endpoints of the clock distribution network at each block may be computed or measured. Once the top-level clock distribution network meets top-level clock timing goals such as latency and / or skew, the top-level clock distribution network may be pushed down into the blocks and / or sub-blocks at the block level. Since the block-level clock timing for different blocks and / or sub-blocks may be different, test clock timing such as latency and / or skew may be computed for each block and / or sub-block using, for example, clock tree synthesis (CTS). Using insertion delay as an example, the test insertion delay for each block or sub-block may be added to the insertion delay of the top-level clock distribution network to the endpoint at that block, thereby determining the total test latency for each block or sub-block starting at the top-level clock root. This may be repeated for any or all of the blocks to determine the block or sub-block having the longest total test latency. The total test latency of each other block or sub-block may be subtracted from the longest total test latency to compute the result that may be used as a target latency constraint for each other block or sub-block. CTS may be run again for each block or sub-block using the target latency constraint to compute the new or modified latency for each other block or sub-block. The blocks and / or sub-blocks may then be merged to create the overall hierarchical clock tree. The recomputed latency for each block or sub-block may then be added to the insertion delay of the top-level clock distribution network to the endpoint at that block, thereby determining the new or modified total latency for each block or sub-block. If the new or modified total latency meets the design goals, the hierarchical clock tree is considered balanced. If not, the block-level clock timing may be adjusted by engineering change order (ECO) and re-merging the blocks, or by adjusting the constraints for one or more of the blocks or sub-blocks and re-running CTS. If the hierarchical clock tree is not balanced by ECO or adjusting the constraints, the top-level clock distribution network may be re-planned and / or re-constructed, and the test timing process may be repeated.

[0035] Referring to Figure 5A and 5B ,the method may begin at step 123, where the designer may start with a floorplan having blocks that have been laid out during the layout phase of the physical design process. At step 124, the designer may plan a top-level clock distribution network that may include all of the structures required to distribute the source clock signal to all of the blocks. Figure 6Shows an exemplary block layout. In this example, blocks Z0 and Y0 are single-instance blocks, while blocks A0 and A1, C0 and C1, and B0 - B3 are sets of multi-instantiated blocks (MIBs). As an example, all MIBs may have a block-level clock tree implemented using clock tree synthesis (CTS) starting at a single-ended buffer, which may be referred to as a CTS root buffer, in each block. Also as an example, blocks Z0 and Y0 may use multi-source CTS (MSCTS), where each of the four-ended buffers in each block can be one of the multiple leaf points for the MSCTS in that block. The fine structure 164 located around the periphery of the block and / or sub-block can be a macro cell that can be placed before any standard cells in the remaining area of each block or sub-block. The macro cells can be placed first because they may have higher specifications than the standard cells.

[0036] Although this step can be automated, in this embodiment, the designer can manually plan and construct the top-level clock distribution network, which can be done independently of the EDA tool used for the design and which can provide the designer with a full-chip view of the hierarchical clock tree. This can reduce the number of design iterations and lead to faster timing convergence.

[0037] In step 126, the designer can construct the top-level clock distribution network, for example, based on a configuration file that can specify the topology of the clock distribution network. At this point in the design process, since the routing can be only topological and the routing to the buffer terminals may not be complete, the clock buffers may be placed in illegal locations.

[0038] Word terminals can be used to represent physical connections including the physical location and / or shape of the physical connection. Word ports can be used to represent logical connections. Depending on the context, word pins can be used interchangeably to represent terminals or ports.

[0039] Figure 6 Shows the layout of the clock distribution network 160, which can start at the clock root 162 and can have a backbone that branches to endpoints indicated by circles 168A to 168R that are not part of the design. For simplicity of illustration, only the routing of the clock distribution network 160 is shown, but it can also include buffers, gates, and any other clock units located at appropriate points in the network.

[0040] In step 128, as Figure 7As shown, a single dummy flip-flop or other sequential logic load 170A, 170B, etc. (collectively referred to as 170) can be added at each endpoint so that one or more EDA tools can calculate the insertion delay from the clock root 162 to each endpoint. The insertion delay can be calculated for all relevant process, voltage, and temperature (PVT) corners to find the maximum insertion delay to each endpoint at each PVT corner. The delay values calculated for each sub-block can be saved for use in step 140, and the delay values calculated for each sub-block can be used in step 140 to calculate the total delay of each sub-block. For simplicity, the fine structure 164 has been omitted from Figure 7 the fine structure 164 is omitted.

[0041] In step 130, the top-level clock distribution network including wiring and clock buffers can be pushed down to blocks and sub-blocks including multi-instantiated blocks that can be properly processed. The top-level and block-level connections can be modified as needed, and as Figure 8 shown, pairs of adjacent block terminals 172A, 172B, etc. (collectively referred to as 172) including feedthrough can be created at each location where the wiring of the clock distribution network passes between adjacent blocks. For simplicity, the fine structure 164 has been omitted from Figure 8 the fine structure 164. During the push-down, a sub-block profile can be created for each sub-block. An engineering change order (ECO) file can also be created for each sub-block. For multi-instantiated blocks, profiles and / or ECO files can be created only for the master instance that can be determined, for example, by the top-level floorplan designer. This can capture the location where the buffers are placed and the topology of the wiring. Layout and routing blocks can be generated to block the legal locations of the layout buffers. The ECO file and the profile can provide two separate or complementary ways to implement the push-down. In addition, the top-level and / or block-level profiles can provide a unique way for the designer to construct the clock distribution network and / or the hierarchical clock tree.

[0042] In step 132, if the push-down has resulted in a floorplan change for any block or sub-block that may require a change to the pushed-down clock distribution network, the method can proceed to step 134. In addition, if any sub-block may require a change to the pushed-down clock distribution network to meet the new top-level and / or block-level delay and / or skew targets, the method can proceed to step 134. Otherwise, the method can proceed to step 136.

[0043] In step 134, the pushed-down clock distribution network including wiring and / or cells can be modified to meet the new delay and / or skew targets.

[0044] In step 136, the construction of the entire hierarchical clock tree can be started based on the combined configuration file generated by the pushdown. For each block, the portion of the top-level clock distribution network in the block can be constructed at the block layer with the wiring topology and buffer layout of the portion of the top-level clock distribution network in the block recreated. Depending on the implementation details, the recreation can be substantially accurate. At this step, the wiring to the terminals of the buffers in the top-level clock distribution network can also be completed. The clock gates and flip-flops can be reordered using each distribution endpoint with techniques such as clock gate merging and splitting. Figure 9 An example embodiment showing what the pushdown clock cells and clock spine segments 160A through 160G for block Y0 might look like. As an example, some of the clock buffer cells are indicated as 161A and 161B.

[0045] In step 138, the clock tree for each block or sub-block can be constructed, and the insertion delay for each block or sub-block can be calculated starting from the endpoints of the clock distribution network. This can be described as calculating the block layer test timing. In this example embodiment, clock tree synthesis (CTS) can be used, but any other suitable technique can also be used to calculate or measure the insertion delay for each block or sub-block. As an example, a clock grid structure and the accompanying timing analysis can be used for some or all of the block layer clock trees.

[0046] In step 140, the total delay or combinational delay from the top-level clock root can be calculated for each block or sub-block. This can be achieved by summing the insertion delay from the clock root 162 to each endpoint calculated in step 128 and the insertion delay of the corresponding block or sub-block calculated in step 138. Then, the block or sub-block with the longest total insertion delay among all blocks or sub-blocks can be identified. For all other blocks, the insertion delay of each block or sub-block can be subtracted from the longest total insertion delay, and the result of the subtraction can be used as the clock delay target for that block. Thus, the block layer CTS target constraints can be obtained from the top-level clock distribution network and the block layer CTS test results. The clock delay target for each block can be used as the clock insertion delay constraint in step 142.

[0047] In step 142, the insertion delay constraint calculated in step 140 can be used to recalculate the insertion delay for each block or sub-block.

[0048] In some embodiments, for example, as Figure 1 disclosed in step 106 of, steps 138 through 142 can be collectively referred to as an example of implementing a partitioned clock tree.

[0049] In Figure 5BIn step 144, blocks and / or sub-blocks can be merged to form a complete top-down hierarchical clock tree structure and a merged database with the insertion delay of each block or sub-block, where the insertion delay of each block or sub-block has been recalculated in step 142 using the insertion delay constraints from step 140. Then, the total delay starting from the top-level clock root can be recalculated for each block or sub-block using the delay data from the merged database. The delay can be evaluated at one or more corners including any or all relevant process, voltage, and temperature (PVT) corners.

[0050] In step 146, the recalculated delay of each block or sub-block can be checked against the design goals for skew and / or delay. If the goals are met, the method can terminate in step 148. All or any part of the overall hierarchical clock tree including routing, buffers, and / or clock cells can then be saved, for example, by locking them to prevent the EDA tool from modifying all or any part of the overall hierarchical clock tree during the implementation of the partition.

[0051] If the design goals are not met in step 146, the method can proceed to step 150, where in step 150, one or more ECOs can be used to make minor changes to one or more blocks or sub-blocks. Then, the blocks and / or sub-blocks can be re-merged in step 144, and the blocks and / or sub-blocks can be re-checked in step 146. The method can make one or more attempts through the loop of steps 150, 144, and 146 to achieve the clock timing goal.

[0052] In step 150, if it is determined that the skew and / or delay goals cannot be met using the ECO, the method can proceed to step 152, where in step 152, one or more of the insertion delay constraints of one or more of the blocks and / or sub-blocks can be adjusted. Then, the method can return to step 142, where in step 142, the insertion delay of one or more blocks or sub-blocks can be recalculated using one or more of the adjusted insertion delay constraints from step 152. The method can make one or more attempts through the loop of steps 152, 142, 144, 146, and 150 to achieve the clock timing goal. The method can also go back and forth between the inner block layer loop of step 150 and the outer block layer loop of step 152.

[0053] In step 152, if it is determined that the skew and / or delay goals cannot be met by adjusting one or more of the insertion delay constraints of one or more of the blocks and / or sub-blocks, the method can proceed to step 126, where in step 126, the designer, who may include automated processing in some embodiments throughout this disclosure, can re-plan and / or re-build the top-level clock distribution tree, but this time benefiting from the knowledge obtained through Figure 5A and Figure 5B the main flow of the method.

[0054] When the method proceeds through Figure 5A and Figure 5B steps, the designer can obtain an increasingly better chip-level or multi-level view and understanding of the entire hierarchical clock tree, which cannot be obtained using conventional EDA tools. When the method reaches step 146, the designer can not only have a profound understanding of the timing behavior of the entire hierarchical clock tree, but the designer can also have the benefits of ECO and constraint-based tuning processes, which enable the designer to quickly adjust clock timing at the block level and can lead to faster timing convergence as well as lower total power consumption and smaller design area. Even if the designer returns to step 126 to re-plan and / or re-construct the top-level clock distribution tree, the designer can generally be able to quickly adjust the clock distribution tree, which results in meeting the timing design goals with little or no adjustment required at the block level. In addition, the decoupling of the dependencies between the top-level clock distribution tree and the block-level clock tree can enable the top-level and block-level balancing operations to be performed in parallel. Moreover, because the designer can be able to use ECO and / or budgeting / tuning at the block level to balance the entire hierarchical clock tree, in some embodiments and / or scenarios, the top-level clock distribution tree may not require any balancing, which can reduce or minimize top-level latency.

[0055] Regardless of the path taken by the method through Figure 5A and Figure 5B once the clock timing goals are met at step 146, the method can terminate at step 148. All or any part of the entire hierarchical clock tree including wiring, buffers, and / or clock cells can then be saved, for example, by locking them to prevent the EDA tools from modifying all or any part of the entire hierarchical clock tree during the implementation of the partition.

[0056] Referring to Figure 6 , the fine structure 164 located around the periphery of the block and / or sub-block can be a macro cell that can be placed before any standard cells that can be laid out in the remaining area of each block or sub-block. The macro cells can be laid out first because they can have higher specifications than standard cells. In some embodiments, it may be beneficial or necessary to avoid placing clock cells (such as buffers, gates, etc.) for the clock distribution network in positions occupied by memory cells. However, even though the components and wiring within the clock cells can be laid out and wired in the lower layers of the integrated circuit, most of the wiring between the cells for the clock distribution network can be wired through the higher layers of the integrated circuit. Since the wiring at the higher layers can have thicker wires and larger spacing between the wires, this can reduce the resistance, capacitance, and / or other potential adverse characteristics of the wiring.

[0057] In some additional embodiments, a hierarchical clock tree may be built from the bottom up. Using this method, the top-level clock designer may convert the entire hierarchical clock tree structure, including all clock cell locations and routing, into each block layer. The clock ports may be aligned, and the clock delay may be calculated by summing the insertion delays from each block or sub-block. This may be achieved, for example, by manual calculation and / or scripts. In other additional embodiments, buffers may be automatically placed and / or sized based on information such as timing per unit length, metal information, etc., which may be provided, for example, by a look-up table or other sources.

[0058] In addition to those mentioned above, and depending on implementation details and circumstances, the principles of the present disclosure may provide any or all of the following benefits and / or features: faster timing convergence for the entire design, especially at the interfaces between MIM and non-MIM blocks; easy delay and skew control for the entire hierarchical clock tree design; overall better clock trees for lower power, low delay, and skew; independence from integrated circuit technology; support for any number of blocks and / or sub-blocks; support for any layout, including straight layout shapes and adjacent sub-blocks; support for any user-specified non-default routing rules for the clock network; support for any conventional standard cell clock driver or custom clock cell; support for multiple clocks; support for a multi-layer hierarchy including pushing down one or more layers at a time; support for adjacent designs, which may save project execution time and design area compared to non-adjacent designs where top-level balancing may have to wait until all sub-blocks are complete; support for CTS for sub-blocks, which may save design time and / or resources and reduce power consumption and / or design area compared to clock grids; reduction of the offline end time; and / or more stringent control of clock timing and flexibility in clock distribution and implementation. In some embodiments, some principles of the present disclosure may provide top-down hierarchical clock balancing and bottom-up clock network tuning.

[0059] Figure 10 An embodiment of a computing system according to the present disclosure is shown. Figure 10The system 300 can be used to implement any or all of the methods and / or apparatuses described in the present disclosure. The system 300 can include a central processing unit (CPU) 302, a memory 304, a storage device 306, a user interface 308, a network interface 310, and a power supply 312. The clock tree building logic 307 can include logic for implementing any of the features described in the present disclosure, the features including building a clock distribution network, pushing the network to different layers, implementing a partitioned clock tree, making an ECO, calculating timing, etc. In different embodiments, the system can omit any of these components, or can include any component and duplicates of any other type of component for implementing any method and / or apparatus described in the present disclosure or any additional number of any component and any other type of component for implementing any method and / or apparatus described in the present disclosure.

[0060] The CPU 302 can include any number of cores, caches, buses, and / or interconnect interfaces and / or controllers. The memory 304 can include any arrangement of dynamic RAM and / or static RAM, non-volatile memory (such as flash memory), etc. The storage device 306 can include a hard disk drive (HDD), a solid state drive (SSD), and / or any other type of data storage device or any combination thereof. The user interface 308 can include any type of human-machine interface device (such as a keyboard, a mouse, a monitor, a video capture or transmission device, a microphone, a speaker, a touch screen, etc.) and any virtual or remote version of such devices. The network interface 310 can include one or more adapters or other devices to communicate via Ethernet, Wi-Fi, Bluetooth, or any other computer network arrangement to enable the components to communicate via a physical and / or logical network (such as an intranet, the Internet, a local area network, a wide area network, etc.). The power supply 312 can include a battery and / or any form of power supply capable of receiving power from an AC or DC power source and converting it into a form suitable for use by the components of the system 300.

[0061] Any or all components of the system 300 can be interconnected via a system bus 301, which can collectively include various interfaces including the following: a power bus, an address and data bus, a high-speed interconnect such as Serial ATA (SATA), a Peripheral Component Interconnect (PCI), a Peripheral Component Express (PCI-e), a System Management Bus (SMB), and any other type of interface that can enable the components to work together and is located locally at one location and / or distributed between different locations.

[0062] System 300 may also include various chip sets, interfaces, adapters, glue logic, embedded controllers (such as programmable or non-programmable logic devices or arrays), application specific integrated circuits (ASICs), embedded computers, smart cards, etc. arranged such that the various components of system 300 can work together to implement any of the methods and / or apparatuses described in this disclosure. Any component of system 300 can be implemented in hardware, software, firmware, or any combination thereof. In some embodiments, any or all components can be implemented in virtualized form and / or in a cloud-based implementation with, for example, flexible resource configurations within a data center or distributed across multiple data centers.

[0063] The blocks or steps of the methods or algorithms and functions described in connection with the embodiments disclosed herein can be implemented directly in hardware included in system 300, in software modules executed by a processor, or in a combination of both. If implemented in software, the functions can be stored on or transmitted over a tangible, non-transitory computer-readable medium as one or more instructions or code. The software modules can reside in random access memory (RAM), flash memory, read only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disk, removable disk, CD ROM, or any other form of storage medium.

[0064] Unless otherwise clear from the context, the terms (such as "first" and "second") used in this disclosure and the claims may be used only for the purpose of distinguishing the things they modify and may not indicate any spatial or temporal order.

[0065] The various details and embodiments described above can be combined to produce additional embodiments in accordance with the inventive principles of this patent disclosure. Since the inventive principles of this patent disclosure can be modified in arrangement and detail without departing from the inventive concept, such changes and modifications are considered to fall within the scope of the claims.

Claims

1. A method for constructing a hierarchical clock tree for an integrated circuit, the method comprising: Constructing a clock distribution network on a first layer; Pushing the clock distribution network to a second layer; Implementing a partitioned clock tree in a partition on the second layer; And Calculating a combined timing of the clock distribution network and the partitioned clock tree on the second layer, wherein the step of implementing the partitioned clock tree includes: constructing a partitioned clock tree in a partition on the second layer; calculating a test timing of the partitioned clock tree; calculating a target timing constraint of the partitioned clock tree based on the timing of the clock distribution network and the test timing of the partitioned clock tree; and adjusting the timing of one or more in the partitioned clock tree based on the target timing constraint.

2. The method according to claim 1, wherein, The step of calculating the combined timing of the clock distribution network and the partitioned clock tree on the second layer includes merging the partitioned clock tree.

3. The method according to claim 1 or claim 2 further comprises: Adjusting the timing of one or more in the partitioned clock tree on the second layer.

4. The method according to claim 3, wherein The timing of one or more in the partitioned clock tree on the second layer is adjusted by an engineering change order.

5. The method according to claim 3, wherein, The timing of one or more in the partitioned clock tree is adjusted by adjusting one or more target constraints of one or more in the partitioned clock tree on the second layer.

6. The method according to claim 3, further comprising: Determining that a timing target is not met by adjusting the timing of one or more in the partitioned clock tree; And Adjusting the clock distribution network.

7. The method according to claim 1 or claim 2 further comprises: Balancing the clock distribution network in parallel with one or more in the balanced partitioned clock tree.

8. The method according to claim 1 or claim 2, further comprising: Pushing the clock distribution network to a third layer; And Implementing a partitioned clock tree in a partition on the third layer.

9. The method according to claim 1, wherein The clock distribution network is pushed to one or more in a partition on the second layer.

10. The method according to claim 1, wherein The second layer includes a block layer.

11. The method according to claim 1, wherein, The second layer is lower than the first layer.

12. A method for constructing a hierarchical clock tree for an integrated circuit, the method comprising: Constructing a clock distribution network on a first layer; Pushing the clock distribution network to a partition at a second layer; Constructing a partitioned clock tree in the partition at the second layer; Calculating a test timing of the partitioned clock tree; Calculating a combined timing of the timing of the clock distribution network and the test timing of the partitioned clock tree; Calculating a target constraint of the partitioned clock tree based on the combined timing of the timing of the clock distribution network and the test timing of the partitioned clock tree; And Calculating a modified timing of the partitioned clock tree based on the target constraint.

13. The method according to claim 12, further comprising: Merging the partitioned clock tree; Calculating a modified combined timing of the timing of the clock distribution network and the modified timing of the partitioned clock tree; And Checking whether the modified combined timing meets a design target.

14. The method according to claim 13 further comprises: Balancing the hierarchical clock tree by adjusting the modified timing of the partitioned clock tree.

15. The method according to claim 14, wherein The modified timing of the partitioned clock tree is adjusted by an engineering change order.

16. The method according to claim 14, wherein, The modified timing of the partitioned clock tree is adjusted by adjusting the target constraint of the partitioned clock tree.

17. The method according to claim 12, wherein, Timing includes delay.

18. The method according to claim 12, wherein Timing includes skew.

19. The method according to claim 12, wherein, A partition includes adjacent blocks.

20. The method according to claim 12, wherein The second layer includes channel-based blocks.

21. The method according to claim 12, wherein, A partition includes multi-instantiated modules.

22. The method according to claim 12, wherein, The dependency between the clock distribution network and the partitioned clock tree is decoupled.

23. The method according to claim 22 further comprises: Balancing the clock distribution network and the partitioned clock tree in parallel.

24. A method for constructing a hierarchical clock tree for an integrated circuit, the method comprising: Constructing a top-level clock distribution network; Calculate the distribution delay to the endpoints of the top-level clock distribution network; Push the top-level clock distribution network down to the block level; Build a clock tree in the block at the endpoint; Calculate the block-level insertion delay of the clock tree in the block at the endpoint; Combine the distribution delay with the block-level insertion delay to calculate the clock tree insertion delay from the root of the top-level clock distribution network; Calculate the delay target constraint for the block based on the clock tree insertion delay from the root of the top-level clock distribution network; Recalculate the block-level insertion delay based on the delay target constraint; Merge the clock trees at the block level; And Recalculate the clock tree insertion delay from the root of the top-level clock distribution network.

25. The method according to claim 24, further comprising: Check whether the recalculated clock tree insertion delay from the root of the top-level clock distribution network meets the design goal.

26. The method according to claim 24 or claim 25, further comprising: Determine that the insertion delay from the root of the top-level clock distribution network does not meet the design goal; And Use an engineering change order to change the clock cells on the clock tree.

27. The method according to claim 24 or claim 25, further comprising: Determine that the insertion delay from the root of the top-level clock distribution network does not meet the design goal; And Recalculate the delay target constraint for the block.

28. The method according to claim 24, wherein The step of pushing the top-level clock distribution network down to the block level includes creating an engineering change order file for the sub-block.

29. The method according to claim 24, wherein, The step of pushing the top-level clock distribution network down to the block level includes creating a configuration file for the sub-block.

30. The method according to claim 24 or claim 25, further comprising: Determine that the sub-block layout is changed to push the top-level clock distribution network down to the block level; And Modify the top-level clock distribution network structure at the block level to preserve the distribution delay to the endpoints of the top-level clock distribution network.

31. The method according to claim 24 or claim 25, further comprising: Save at least a portion of the hierarchical clock tree to prevent modification during block-level implementation.

Citation Information

Patent Citations

  • Methodology to optimize hierarchical clock skew by clock delay compensation

    US20050102643A1

  • Incremental clock tree synthesis

    US20140189627A1

  • Method and system for providing hybrid clock distribution

    US7392495B1