Efficient clock tree synthesis method based on two-partition four-branch tree

By adopting a clock tree synthesis method based on two-partition four-branch H-tree in large-scale integrated circuit design, the problems of clock path delay, poor clock offset control effect and excessive buffers in large-scale integrated circuit design are solved, and more efficient clock tree comprehensive performance and flexibility are achieved.

CN120145987AActive Publication Date: 2025-06-13NANJING UNIV OF POSTS & TELECOMM
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411950766.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-06-13
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

In large-scale integrated circuit design, traditional clock tree structures have area and power consumption problems caused by clock path delay, poor clock offset control effect, and excessive buffer number.

Method used

The clock tree synthesis method based on two-partition four-branch H-tree is adopted, and the number of buffers is reduced through the GBC greedy clustering algorithm, combined with the buffer reselection bit algorithm with low clock offset and the long path delay optimization algorithm, the clock tree is built from top to bottom to optimize the clock offset and delay.

Benefits of technology

It effectively reduces the number of buffers, reduces the delay and offset of the clock tree, and improves the comprehensive performance and flexibility of the clock tree.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145987A_ABST
    Figure CN120145987A_ABST
Patent Text Reader

Abstract

The invention discloses an efficient clock tree synthesis method based on a two-partition four-branch tree, which belongs to the field of integrated circuits and comprises the following steps: clustering triggers in a clock tree by adopting a greedy strategy iteration selection point through a GBC greedy clustering algorithm; on the basis of bottom layer clustering, based on a buffer reselection algorithm of low clock skew, reselection is carried out on a buffer through a far-near collocation principle; a two-partition four-branch H-tree-like structure is adopted, an H-tree-like structure is constructed on the top layer of a clock tree, and a fixed number of buffers are evenly inserted into a long path; according to the method, the position of the buffer is adjusted through geometric Boolean operation, an insertable point of the buffer is generated by conducting expansion and grid segmentation on a free area, and then the position of the buffer is dynamically adjusted based on the property of a Manhattan rectangle. The clock tree design efficiency and precision are effectively improved, and the method is suitable for clock tree comprehensive design of a high-performance integrated circuit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of integrated circuits, and particularly relates to an efficient clock tree synthesis method based on a two-partition four-branch tree. Background Art

[0002] Clock Tree Synthesis (CTS) is a crucial step in the design of digital integrated circuits and plays a vital role in timing design. The main task of the clock tree is to distribute the clock signal from the source node to each flip-flop (FF) and logic module in the circuit and ensure the timing requirements for signal arrival. The optimization of the clock tree has a direct impact on the performance, power consumption, and area of the system. Especially in the design of high-frequency and large-scale integrated circuits, the performance of the clock tree determines the maximum operating frequency and signal synchronization of the entire system.

[0003] In traditional clock tree synthesis methods, the selection of the clock tree structure and the buffer (Buf) insertion strategy are the key factors determining the performance of the clock tree. Common clock tree structures include the H-tree and the fishbone tree, etc., each having different advantages and disadvantages.

[0004] The H-tree structure is one of the most common structures in clock tree synthesis. It has good symmetry, can effectively reduce clock skew and achieve uniform clock signal transmission. The advantages of the H-tree lie in its simple structure, easy implementation, and low clock skew in a symmetric layout, which can better meet the timing requirements of high-frequency circuits. However, the H-tree also has certain limitations, mainly reflected in its lack of flexibility for complex circuits. As the circuit scale increases, the depth of the H-tree structure increases, resulting in an increase in the delay of the clock tree and affecting the performance of the system.

[0005] The main feature of the fishbone tree is to optimize the propagation path of the clock signal through multiple shorter paths, thereby significantly reducing the total delay of the clock signal. Compared with the traditional H-tree structure, the fishbone tree can reduce the long-term cumulative delay in signal transmission to a certain extent. Especially in multi-level clock paths, it can effectively reduce the propagation delay of the clock and improve the timing performance. However, the design of the fishbone tree focuses on the optimization of path delay and does not particularly balance the clock arrival time of each branch on the clock tree, which may result in a less effective Skew control effect of the clock than the H-tree.

[0006] During the clock tree synthesis process, the insertion strategy of buffers also plays a crucial role in the performance of the clock tree. Generally, the goal of buffer insertion is to reduce the signal propagation delay and the skew of the clock tree. However, the selection of the number, location, and type of buffers is a complex multi-objective optimization problem, involving trade-offs in multiple aspects such as power consumption, area, clock delay, and clock skew. For this reason, many studies have proposed buffer insertion methods based on different optimization strategies, such as top-down insertion method, local optimization method, and constraint-based insertion method, etc. Summary of the Invention

[0007] In summary, through innovative algorithms and optimization strategies, the present invention proposes a new clock tree synthesis method based on a two-partition four-branch tree, aiming to improve the performance of the clock tree and solve the challenges faced by traditional clock tree structures in large-scale integrated circuit design. The purpose of the present invention is to overcome the deficiencies in aspects such as clock path delay, clock skew control, and the number of buffers in existing clock tree optimization methods, and provide a clock tree synthesis method based on a two-partition four-branch H-like tree. Considering the area and power consumption problems caused by too many buffers, the GBC greedy clustering method is used at the bottom layer to maximize the utilization rate of buffer fan-out as much as possible, thereby greatly reducing the number of buffers used. Considering the optimization of clock skew and clock delay, while constructing the H-like tree from top to bottom, a delay-optimized buffer insertion algorithm is combined to effectively control the clock delay under low clock skew. Through an optimization strategy that combines bottom-up and top-down, the three key indicators of clock tree synthesis are comprehensively improved. At the same time, compared with the traditional clock tree structure, the flexibility of optimizing different key indicators is improved.

[0008] The present invention adopts the following technical solutions to achieve the above-mentioned invention purposes: An efficient clock tree synthesis method based on a two-partition four-branch H-like tree, comprising the following steps: Step 1, through the GBC greedy clustering algorithm, using a greedy strategy to iteratively select points, cluster the flip-flops (FFs) in the clock tree to maximize the utilization rate of buffer fan-out, and minimize the number of buffers required to the greatest extent; Step 2, on the basis of the bottom-layer clustering, based on the buffer repositioning algorithm with low clock skew, reposition the buffers according to the principle of combining near and far to optimize the clock skew, and appropriately sacrifice the delay to balance the multi-objective requirements; Step 3, adopt a two-partition four-branch H-like tree structure, construct an H-like tree at the top layer of the clock tree, and ensure low clock skew while effectively reducing the delay on the path by uniformly inserting a fixed number of buffers on the long path; Step 4: Adjust the buffer positions using geometric Boolean operations. By dilating the free area and performing grid segmentation, the insertable points for the buffers are generated. Then, based on the properties of Manhattan rectangles, the buffer positions are dynamically adjusted to avoid buffer overlap while ensuring the optimization of local clock tree metrics.

[0009] As an efficient clock tree synthesis method based on a two-partition four-branch H-tree-like structure, the method for clustering clock tree flip-flops in Step 1 is as follows: Using the GBC greedy clustering algorithm, based on a greedy strategy, flip-flops in the clock tree are iteratively selected and clustered. During the clustering process, considering the connection relationships, fan-out degrees, and resource requirements of the buffers for the flip-flops, the fan-out utilization rate of the buffers is maximized, thus effectively reducing the number of required buffers. The flip-flops in each cluster are regarded as an independent clustering unit, reducing the algorithm complexity of subsequent synthesis.

[0010] As an efficient clock tree synthesis method based on a two-partition four-branch H-tree-like structure, the method for optimizing clock skew based on the buffer repositioning algorithm with low clock skew in Step 2 is as follows: Based on the flip-flop clustering, the buffer repositioning algorithm with low clock skew is used to reposition the selected buffers. Adopting the principle of "matching near and far", that is, according to the relative distances between the flip-flops and the buffers, the positions of the buffers are optimized and adjusted to minimize the clock skew to the greatest extent. During this process, the delay of some clock paths is appropriately sacrificed to balance the multi-objective requirements of clock skew control and clock path delay.

[0011] As an efficient clock tree synthesis method based on a two-partition four-branch H-tree-like structure, the method for constructing the clock tree using the two-partition four-branch H-tree structure in Step 3 is as follows: At the top layer of the clock tree, a two-partition four-branch H-tree-like structure is adopted to construct the main trunk of the clock tree. The H-tree-like structure effectively reduces the delay on long paths by uniformly inserting a fixed number of buffers on the clock paths. In the two-partition four-branch structure, by reasonably partitioning the clock paths and allocating symmetric buffers to control the skew of the clock signal, the timing performance of the clock tree is further improved.

[0012] As an efficient clock tree synthesis method based on a two-partition four-branch H-tree-like structure, the method for adjusting the buffer positions through geometric Boolean operations in Step 4 is as follows: Geometric Boolean operations are used to optimize and adjust the buffer positions. During the clock tree layout process, by dilating the free area and performing grid segmentation, the insertable points for the buffers are extracted. Based on the properties of Manhattan rectangles, by dynamically adjusting the positions of the buffers, buffer overlap is avoided while ensuring the optimization of various key metrics in the local area of the clock tree.

[0013] The present invention adopts the above technical solution and has the following beneficial effects: (1) The present invention proposes an efficient clock tree synthesis method based on a two-partition four-branch H-tree. Through the GBC greedy clustering algorithm, while satisfying the constraints, the number of buffers required for the underlying clock tree construction is reduced as much as possible, the algorithm complexity of the later synthesis is reduced, and the program running speed is improved.

[0014] (2) The present invention constructs a two-partition four-branch H-tree and combines it with a long path delay optimization algorithm to effectively optimize the global delay of the clock tree from top to bottom while controlling the clock deviation to be extremely low, thereby greatly improving the performance of clock tree synthesis. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The accompanying drawings are used to constitute a part of the specification to facilitate further understanding of the present invention, and are used together with the embodiments of the present invention to explain the present invention but do not constitute a limitation of the present invention; Figure 1 It is a flow chart of the efficient clock tree synthesis method based on a two-partition four-branch H-tree of the present invention; Figure 2 It is a schematic diagram of a two-partition four-branch class H-tree established by the present invention; Figure 3 It is a schematic diagram of a buffer reselection algorithm based on low clock offset of the present invention; Figure 4 It is a schematic diagram of the buffer position adjustment based on geometric Boolean operation of the present invention; Figure 5 middle Figure 5-1 This is the optimization curve for inserting BUF quantity taking case 10 as an example. Figure 5-2 is a piecewise function. DETAILED DESCRIPTION

[0016] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0017] In order to make the purpose, technical features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should belong to the protection scope of the present invention.

[0018] In the following description, many specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0019] The present invention will be further described below with reference to the accompanying drawings.

[0020] 1) GBC Greedy Clustering Algorithm First, determine the number of partition regions according to the given layout length and width and the number of FFs. Then, select the FF farthest from the center of the region in each region as the initial clustering point, greedily select adjacent FFs to join the clustering, and check the max_fanout and max_rc constraints. If not satisfied, set this point as a new class. Repeat the above steps until all points have been grouped into clusters, where FF is a flip-flop, max_fanout is the maximum fan-out, and max_rc is the maximum resistance capacitance.

[0021] 2) Two-Partition Four-Branch Class H-Tree Construction The traditional H-tree structure is single. Therefore, we constructed a class H-tree structure as shown in Figure 2 which can more flexibly adjust the position of the buffers while ensuring low clock skew. In fact, the construction of the class H-tree is inseparable from the buffer insertion algorithm for long paths. Through experiments, it can be obtained that inserting a fixed number of buffers in a certain path can effectively reduce the clock delay. On this basis, it is deduced from the mean value inequality that when buffers are inserted evenly in the same path, the lowest clock delay in the current path can be obtained. Taking the insertion of one buffer as an example, ensuring that the Manhattan distance from buffer A to buffer B is equal to the Manhattan distance from buffer B to buffer C can optimize the clock delay in the current path. Generally speaking, the construction of the class H-tree completes the optimization of the clock delay from top to bottom.

[0022] 3) Long Path Buffer Insertion Delay Optimization Algorithm Analyzing the definitions of buffer delay and wire network unit delay given in the problem, it can be obtained that inserting a fixed number of BUFs in a certain path can effectively reduce the time delay. The derivation of its insertion quantity and position is as follows: The time delay latency1 of a certain path without inserting a buffer: , in the above formula (2), r is the resistance per unit length, c is the capacitance per unit length, D is the Manhattan distance of the path between two points, k is a defined constant, and the time delay latency2 of a certain path when inserting n buffers: , in the above formula (3), r is the resistance per unit length, c is the capacitance per unit length, D is the Manhattan distance of the path between two points, k is a defined constant. According to the mean value inequality, when inserting BUFs in the same path, the lowest delay can be obtained under the current conditions when the BUFs are evenly distributed. Therefore, the method of evenly inserting BUFs is adopted. The following is the derivation of the optimal insertion number: If it is necessary to satisfy, the following relational expression exists: , in the above formula (4), k is a defined constant, D is the Manhattan distance of the path between two points, n is the number of buffers inserted in the path, and the optimal insertion number under a certain Manhattan distance D is solved : , since the insertion number variable n is discrete, the traversal method can be used to determine the number of BUFs inserted in the current path. The following Figure 5 is the optimization curve of the number of inserted BUFs taking case10 as an example ( Figure 5-1 ), Figure 5-2 is a piecewise function.

[0023] Therefore, in the subsequent clock tree synthesis algorithm, first calculate the piecewise function curve with the values given in the question, and use the look-up table method to adjust the number of buffer insertions in different distance segments.

[0024] The above specific implementation manners and embodiments are the specific supports for the technical idea proposed by the present invention, and the protection scope of the present invention cannot be limited thereby. Any equivalent change or equivalent modification made on the basis of the technical solution of the present invention according to the technical idea proposed by the present invention still belongs to the protection scope of the technical solution of the present invention.

Claims

1. An efficient clock tree synthesis method based on a two-partition four-branch tree, characterized in that: The following steps are involved: Step 1: cluster the triggers in the clock tree by using the GBC greedy clustering algorithm and adopting the greedy strategy to iteratively select points; Step 2: Based on the underlying clustering, the buffer reselection algorithm based on low clock offset is used to reselect the buffer according to the near-far matching principle. Step 3: Use a two-partition four-branch H-tree structure to build an H-tree at the top level of the clock tree, and insert a fixed number of buffers evenly on the long path; Step 4: Use geometric Boolean operations to adjust the buffer position. By dilating and meshing the free area, insertable points of the buffer are generated. Then, the buffer position is dynamically adjusted based on the Manhattan rectangle properties.

2. The efficient clock tree synthesis method based on a two-partition four-branch tree according to claim 1 is characterized in that: In the step 1, the GBC greedy clustering algorithm greedily iterates and selects appropriate points without specifying the number of initial clusters.

3. The efficient clock tree synthesis method based on a two-partition four-branch tree according to claim 2 is characterized in that: First, determine the number of divided areas according to the given length and width of the map and the number of FFs. Then select the FF farthest from the center of each area as the initial clustering point. Greedily select adjacent FFs to add to the cluster, and check the max_fanout and max_rc constraints. If they are not satisfied, set the point as a new class. Repeat the above steps until all points have been clustered.

4. The efficient clock tree synthesis method based on a two-partition four-branch tree according to claim 1 is characterized in that: In the step 2, the buffer relocation algorithm based on low clock offset optimizes the buffer position by matching the clustering results near and far, specifically including connecting the buffer far from the center point of the clock tree to the buffer close to the center point.

5. The efficient clock tree synthesis method based on a two-partition four-branch tree according to claim 1 is characterized in that: In the step 4, the application of geometric Boolean operation is to dilate the design area and then divide the design area into grids of the size of the buffer to generate insertable points of the buffer.

6. The efficient clock tree synthesis method based on a two-partition four-branch tree according to claim 5 is characterized in that: When inserting BUFs in the same path in step 4, the lowest delay under the current conditions can be obtained when BUFs are evenly distributed. Therefore, the optimal number of insertions under a certain Manhattan distance D is obtained by evenly inserting BUFs. : , the above formula (1) n is the variable for the number of buffers inserted, k is a constant, delay buf is the intrinsic delay of the buffer, due to the insertion of the number variable n It is discrete and uses traversal to determine the number of BUFs inserted in the current path.

Citation Information

Patent Citations

  • Realization method for register clustering in clock tree synthesis

    CN105930591A

  • Clustering-based H-shaped clock tree trunk node coordinate selection method and system

    CN114880983A

  • Clock tree establishment method considering low-voltage clock skew fluctuation optimization

    CN117272878A

  • Clock tree comprehensive optimization method based on H tree

    CN117764024A

  • Flip-flop clustering for integrated circuit design

    US20170228485A1