Efficient clock tree synthesis method based on complementary Manhattan ring

By adopting the complementary Manhattan ring and offset controllable fishbone tree structure in clock tree synthesis, the problems of increasing clock tree delay, poor clock offset control effect and excessive buffers in traditional clock tree synthesis method are solved, and efficient clock tree comprehensive performance is achieved.

CN119940284AActive Publication Date: 2025-05-06NANJING UNIV OF POSTS & TELECOMM
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510413913.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-05-06
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The traditional clock tree synthesis method has problems such as increasing clock tree delay, poor clock offset control effect and excessive buffer number in high-frequency, large-scale integrated circuit design.

Method used

The efficient clock tree synthesis method based on the complementary Manhattan ring is adopted, and the optimization strategy of combining bottom-up and top-down is used to form complementary ring clusters, merge cluster rings, build an offset controllable fish bone tree structure, and adjust the buffer position through geometric operations.

Benefits of technology

The offset at the bottom of the clock tree is reduced, the number of buffers required to build the clock tree is reduced, and the utilization rate of maximum fan-out of the buffer is improved, achieving the effect of extremely low clock deviation and the delay as close to the theoretical limit as possible.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940284A_ABST
    Figure CN119940284A_ABST
Patent Text Reader

Abstract

The invention discloses an efficient clock tree synthesis method based on a complementary Manhattan ring, which comprises the following steps: dividing a layout into a plurality of regions, pre-inserting a plurality of buffers into each region, preliminarily forming a clustering ring taking each buffer as a center, and enabling the size of the clustering ring of each buffer to be complementary with the distance from the buffer to the center of the region; circularly accessing each clustering ring, and searching clusters which can be combined until the number of the clusters is converged; constructing a clock tree by adopting an offset controllable fishbone tree structure; geometric operation is used for judging the overlapping condition of the buffers, a free area is obtained through calculation, the free area is subjected to expansion and grid division, the insertable positions of the buffers are obtained, and then the positions of the buffers are switched and adjusted based on the relation between the buffers at the upper level and the lower level. According to the method, through combination of multiple strategies such as clustering optimization and buffer position adjustment, low delay and a small number of buffers are realized on the basis of low offset, and the method is suitable for clock tree comprehensive design of a high-performance integrated circuit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of integrated circuit physical design, and in particular relates to an efficient clock tree synthesis method based on complementary Manhattan rings. Background Art

[0002] In digital integrated circuit design, Clock Tree Synthesis (CTS) is a key step and plays a vital role in timing design. The core task of the clock tree is to distribute the clock signal from the source node to each flip-flop (FF) and logic module in the circuit, while ensuring that the signal meets the corresponding timing requirements when it arrives. The optimization of the clock tree has a direct impact on the performance, power consumption and area of ​​the system, especially in the design scenario of high-frequency and large-scale integrated circuits. The performance of the clock tree determines the maximum operating frequency and signal synchronization of the entire system.

[0003] In traditional clock tree synthesis methods, the selection of clock tree structure and buffer (BUF) insertion strategy are key factors affecting clock tree performance. Common clock tree structures include H-tree and fishbone tree, each of which has different advantages and disadvantages.

[0004] The H-tree structure is a very common one in clock tree synthesis. It has good symmetry and can effectively reduce clock skew and achieve uniform transmission of clock signals. The advantage of the H-tree is that it has a simple structure and is easy to implement. In addition, when the layout is symmetrical, the clock skew is low, which can well meet the timing requirements of high-frequency circuits. However, the H-tree is not well adapted to complex circuits. As the scale of the circuit continues to expand, the depth of the H-tree structure will increase, which will gradually increase the delay of the clock tree and affect the system performance.

[0005] The outstanding feature of the fishbone tree is that it optimizes the propagation path of the clock signal with the help of multiple shorter paths, thereby significantly reducing the total delay of the clock signal. Compared with the traditional H-tree structure, the fishbone tree can reduce the long-term accumulated delay in the signal transmission process to a certain extent, especially in multi-level clock paths, it can effectively reduce the propagation delay of the clock and improve the timing performance. However, the design focus of the fishbone tree is on the optimization of path delay. On the branches of the clock tree, there is no special balance adjustment for the clock arrival time of each branch, which may cause the skew control effect of the clock to be inferior to that of the H-tree.

[0006] During the clock tree synthesis process, the buffer insertion strategy also plays a key role in the clock tree performance. Generally speaking, buffer insertion aims to reduce signal transmission delay and reduce the skew of the clock tree. However, the selection of the number, location and type of buffers is a complex multi-objective optimization problem involving trade-offs in power consumption, area, clock delay and clock skew. Summary of the invention

[0007] The present invention aims to solve one of the technical problems existing in the related art at least to a certain extent.

[0008] The purpose of the present invention is to provide an efficient clock tree synthesis method based on complementary Manhattan rings, which comprehensively improves the key indicators of clock tree synthesis through a combination of bottom-up and top-down optimization strategies, and improves the flexibility of optimizing different key indicators compared with the traditional clock tree structure.

[0009] In order to achieve the above-mentioned object, the present invention provides an efficient clock tree synthesis method based on complementary Manhattan ring, comprising: S100, dividing the map into multiple regions, pre-inserting a number of buffers into each region, and preliminarily forming a clustering ring centered on each buffer, wherein the size of the clustering ring of each buffer is complementary to the distance from the buffer to the center of the region; S200, looping through the clustering rings to find clusters that can be merged until the number of clusters converges; S300, constructing a clock tree using an offset-controllable fishbone tree structure; S400, using geometric operations to determine the overlap of the buffers and calculate the free area, by expanding and gridding the free area, obtaining the insertable position of the buffer, and then switching and adjusting the buffer position based on the relationship between the upper and lower buffers.

[0010] A further preferred technical solution of the present invention is that step S100 specifically includes: S110, determining the number of divided regions according to the given layout length and width, the number of triggers, the maximum fan-out, and the maximum resistance and capacitance load, and performing region division; S120, in each area, taking the center of the area as a reference, forming equidistant lines with different distances from the center, and then pre-inserting buffers uniformly on the equidistant lines; S130, for each buffer, considering its distance to the center of the region, determining a cluster radius complementary to the distance, so that the delay from the region to each trigger is close; S140 , allocating a trigger to each buffer within a set cluster radius to form a clustering ring.

[0011] Preferably, in step S200, each cluster ring is visited cyclically to find clusters that can be merged. The method adopted is: visit each cluster ring cyclically, continuously reallocate the triggers connected under the buffer with the least fan-out, find other buffers, reduce the number of buffers, and realize the merging of clusters.

[0012] Preferably, when the clock tree is constructed using the offset-controllable fishbone tree structure in step S300, a path optimal insertion method and a path delay adjustment method are used.

[0013] Preferably, the specific steps of step S300 are: S310, based on the path optimal insertion method, plan the shortest distance connection for the cluster ring farthest from the clock source, and then insert buffers evenly according to the distance to minimize the delay, use the delay as the standard delay, and use the connection path as the main path; S320, establishing a branch path based on the main path and connecting to other cluster rings; S330: Based on the path delay adjustment method, adjust the buffer position on the branch so that the delay of each clustering ring converges to the standard delay.

[0014] Preferably, when planning the shortest delay path for the cluster farthest from the clock source in step S310, a uniform insertion buffer is used, and a Manhattan distance The optimal number of insertion buffers The calculation formula is:

[0015] in, is the resistance per unit length, is the capacitance per unit length, is the intrinsic delay of the buffer, is the Manhattan distance of the path between two points; is the variable for the number of buffers inserted. is an integer, the calculated Round up and round down to select the best value.

[0016] Preferably, when adjusting the buffer position on the path where the delay needs to be adjusted in step S330, the path is changed from a straight line to a broken line to increase the delay.

[0017] Beneficial effects: The efficient clock tree synthesis method based on complementary Manhattan rings of the present invention reduces the offset of the bottom layer of the clock tree through complementary ring clustering; by merging clustering, while meeting performance requirements, the number of buffers required for bottom-level clock tree construction is reduced as much as possible, thereby improving the utilization rate of the maximum fan-out of the buffer. The present invention constructs an offset-controllable fishbone tree and combines it with a long path delay optimization algorithm to ensure that the clock deviation is extremely low while making the delay as close to the theoretical limit as possible and the number of buffers as small as possible, thereby achieving the joint optimization of the three and greatly improving the performance of clock tree synthesis. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flow chart of an efficient clock tree synthesis method based on complementary Manhattan rings of the present invention; Figure 2 is a schematic diagram of the present invention based on complementary ring clustering; Figure 3 It is a schematic diagram of the offset controllable fishbone tree structure established by the present invention. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments, and they should not be understood as limitations on the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. In the description of the present invention, it should be understood that the terms used are only for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0020] Combine the following Figure 1-Figure 3 The invention describes an efficient clock tree synthesis method based on complementary Manhattan ring.

[0021] Embodiment: The purpose of this embodiment is to overcome the deficiencies of existing clock tree optimization methods in terms of clock path delay, clock offset control, number of buffers, etc. Through innovative algorithms and optimization strategies, a new clock tree synthesis method based on complementary Manhattan rings is proposed, aiming to improve the performance of the clock tree and solve the challenges faced by traditional clock tree structures in large-scale integrated circuit design.

[0022] The overall process of this method is as follows Figure 1 As shown, including: S100, dividing the map into multiple regions, pre-inserting a number of buffers into each region, and preliminarily forming a clustering ring centered on each buffer, wherein the size of the clustering ring of each buffer is complementary to the distance from the buffer to the center of the region; S200, looping through the clustering rings to find clusters that can be merged until the number of clusters converges; S300, constructing a clock tree using an offset-controllable fishbone tree structure; S400, using geometric operations to determine the overlap of the buffers and calculate the free area, by expanding and gridding the free area, obtaining the insertable position of the buffer, and then switching and adjusting the buffer position based on the relationship between the upper and lower buffers.

[0023] Detailed description of each step: Step S100 is used to form a regional complementary ring cluster. Specifically, it includes: S110, determining the number of divided regions according to the given layout length and width, the number of triggers, the maximum fan-out, and the maximum resistance and capacitance load, and performing region division; S120, in each area, taking the center of the area as a reference, forming equidistant lines with different distances from the center, and then pre-inserting buffers uniformly on the equidistant lines; S130, for each buffer, according to the principle of matching near and far and complementary distances, determine the cluster radius so that the delay of the area to each trigger is close; like Figure 2 As shown, and are the distances from the two buffers to the center of the region, and are the cluster radii of the two buffers respectively, and the size of the trigger clustering ring of each buffer is complementary to the distance from the buffer to the center, expressed as:

[0024] S140 , allocating triggers to each buffer within the set cluster radius to form a clustering ring, thereby ensuring a small offset between the triggers at the bottom layer.

[0025] The pre-insertion buffer can quickly form the preliminary connection of the clustering ring and avoid the non-convergence after the clustering delay iteration due to the uncertainty of the ring center position.

[0026] Step S200 is used to merge clusters and reduce the number of cluster rings.

[0027] The method adopted is: based on trigger clustering, find buffers with small fan-outs, try to connect the triggers under them to other buffers, determine whether the constraints and performance requirements are met, and continuously loop the operation until the number of buffers converges or the maximum number of iterations is reached, thereby minimizing the number of buffers and improving the utilization of the maximum fan-out.

[0028] Step S300 is used to construct a clock tree.

[0029] The traditional fishbone tree is not easy to control the offset and has poor structural flexibility. Therefore, this implementation constructs Figure 3 The offset-controllable fishbone tree structure shown can more flexibly adjust the position of the buffer while ensuring low latency, thereby reducing the offset. The construction of the offset-controllable fishbone tree is based on the path optimal insertion method and the path delay adjustment method. The path optimal insertion method can be derived mathematically, while the path delay adjustment method is a better approach obtained after trying multiple methods. The combination of these two methods can achieve delay convergence and offset reduction.

[0030] The specific steps include: S310, based on the path optimal insertion method, using the offset controllable fishbone tree structure to plan the shortest distance connection for the clustering ring farthest from the clock source, and then inserting the buffer evenly according to the distance to minimize the delay, using the delay as the standard delay, and using the connection path as the main path; According to the definition of buffer delay and line network unit delay, inserting a certain number of buffers can effectively reduce the delay under a fixed path. The insertion quantity and position are derived as follows: Delay without buffer insertion , expressed as:

[0031] in, is the resistance per unit length, is the capacitance per unit length, is the Manhattan distance of the path between two points; insert When there are 1 buffer, the delay of a path , expressed as:

[0032] in, This is the delay of the buffer.

[0033] It can be obtained from the Lagrange multiplier method that when inserting buffers on the same path, the buffers are evenly distributed on the path, which can obtain the lowest delay under the current conditions. Therefore, the method of uniformly inserting buffers is adopted. The following is the derivation of the optimal number of insertions: First calculate :

[0034] Solve for a Manhattan distance The optimal number of insertions for:

[0035] Re-order get:

[0036] Since the number of variables is inserted is an integer, using The number of buffers inserted in the current path is determined by rounding up and down.

[0037] Therefore, in the subsequent clock tree synthesis algorithm, the piecewise function curve is first calculated using the numerical values ​​given in the question, and the number of buffer insertions in different distance segments is adjusted using a lookup table.

[0038] S320, establishing a branch path based on the main path and connecting to other cluster rings; S330, after completing the optimal insertion of the path, the delay needs to be converged according to the standard, and the delay of the branch path needs to be adjusted. According to the derivation in the above method, adjusting the position of the buffer can change the delay of the path. Therefore, the position of the buffer on the path that needs to adjust the delay is adjusted, so that the path changes from a straight line to a broken line, thereby increasing the delay.

[0039] Step S400 is used to adjust the buffer position and layout the clock tree.

[0040] During the clock tree layout process, geometric operations are performed on the plane area around the buffer to remove the space occupied by the trigger and other buffers to obtain the free area. The free area is expanded and segmented to extract the buffer insertion position. Different strategies are switched to select the insertion position according to the connection relationship of the buffer to ensure the optimization of key indicators of the clock tree in the local area.

[0041] In the above technical solution of this embodiment: (1) Considering the area and power consumption issues caused by too many buffers, an iterative merging and clustering method is adopted at the bottom layer to make full use of the maximum fan-out, thereby reducing the number of buffers.

[0042] (2) Considering the optimization of clock offset, a ring cluster is formed during the bottom-level clustering to ensure that the offset between each trigger in the cluster is within a certain threshold. In addition, the cluster radius is adjusted according to the distance complementary relationship, thus achieving effective control of clock delay under low clock offset.

[0043] (3) Considering the reduction of clock delay, while ensuring the offset is small, the longest distance is minimized and the delay of the farthest cluster after merging is used as the standard for global delay convergence, thereby reducing the average delay.

[0044] Through the combination of bottom-up and top-down optimization strategies, the three key indicators of the clock tree are comprehensively improved. Compared with the traditional clock tree structure, the flexibility of optimizing different key indicators is improved.

[0045] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An efficient clock tree synthesis method based on complementary Manhattan rings, characterized in that: include: S100, dividing the map into multiple regions, pre-inserting a number of buffers into each region, and preliminarily forming a clustering ring centered on each buffer, wherein the size of the clustering ring of each buffer is complementary to the distance from the buffer to the center of the region; S200, looping through the clustering rings to find clusters that can be merged until the number of clusters converges; S300, constructing a clock tree using an offset-controllable fishbone tree structure; S400, using geometric operations to determine the overlap of the buffers and calculate the free area, by expanding and gridding the free area, obtaining the insertable position of the buffer, and then switching and adjusting the buffer position based on the relationship between the upper and lower buffers.

2. The efficient clock tree synthesis method based on complementary Manhattan ring according to claim 1, characterized in that: Step S100 specifically includes: S110, determining the number of divided regions according to the given layout length and width, the number of triggers, the maximum fan-out, and the maximum resistance and capacitance load, and performing region division; S120, in each area, taking the center of the area as a reference, forming equidistant lines with different distances from the center, and then pre-inserting buffers uniformly on the equidistant lines; S130, for each buffer, considering its distance to the center of the region, determining a cluster radius complementary to the distance, so that the delay from the region to each trigger is close; S140 , allocating a trigger to each buffer within a set cluster radius to form a clustering ring.

3. The efficient clock tree synthesis method based on complementary Manhattan ring according to claim 1, characterized in that: In step S200, each cluster ring is visited cyclically to find clusters that can be merged. The method used is: cyclically visit each cluster ring, continuously reallocate the triggers connected under the buffer with the least fan-out, find other buffers, reduce the number of buffers, and realize the merging of clusters.

4. The efficient clock tree synthesis method based on complementary Manhattan ring according to claim 1, characterized in that: When the clock tree is constructed using the offset-controllable fishbone tree structure in step S300, the optimal path insertion method and the path delay adjustment method are used.

5. The efficient clock tree synthesis method based on complementary Manhattan ring according to claim 4, characterized in that: The specific steps of step S300 are: S310, based on the path optimal insertion method, plan the shortest distance connection for the cluster ring farthest from the clock source, and then insert buffers evenly according to the distance to minimize the delay, use the delay as the standard delay, and use the connection path as the main path; S320, establishing a branch path based on the main path and connecting to other cluster rings; S330: Based on the path delay adjustment method, adjust the buffer position on the branch so that the delay of each clustering ring converges to the standard delay.

6. The efficient clock tree synthesis method based on complementary Manhattan ring according to claim 5, characterized in that: When planning the shortest delay path for the cluster farthest from the clock source in step S310, a uniform buffer insertion method is adopted, and a certain Manhattan distance The optimal number of insertion buffers The calculation formula is: ; in, is the resistance per unit length, is the capacitance per unit length, is the intrinsic delay of the buffer, is the Manhattan distance of the path between two points; is the variable for the number of buffers inserted. is an integer, the calculated Round up and round down to select the best value.

7. The efficient clock tree synthesis method based on complementary Manhattan ring according to claim 5, characterized in that: When adjusting the buffer position on the path that needs to adjust the delay in step S330, the path is changed from a straight line to a broken line to increase the delay.

Citation Information

Patent Citations

  • Method of constructing hierarchical clock tree for integrated circuit

    CN112100971A

  • Method and device for constructing clock tree and chip

    CN118780238A

  • Clock tree synthesis for low power consumption and low clock skew

    US20060053395A1

  • Method for balanced-delay clock tree insertion

    US6698006B1