Efficient Clock Tree Synthesis Method Based on Complementary Manhattan Rings
Through the clock tree synthesis method based on complementary Manhattan rings, the clock path delay, clock offset and buffer number of clock trees are optimized, and the clock tree comprehensive challenges in large-scale integrated circuit design are solved, achieving more efficient clock tree performance and flexibility.
Patent Information
- Application Number
- CN202510413913.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing clock tree synthesis method has shortcomings in clock path delay, clock offset control and number of buffers in large-scale integrated circuit designs, especially in high-frequency circuits, where the adaptability and optimization flexibility of traditional structures are insufficient.
The efficient clock tree synthesis method based on complementary Manhattan rings is adopted, and the key indicators of the clock tree are optimized through the optimization strategy combining bottom-up and top-down, including layout area division, buffer pre-insertion, clustering merging, offset controllable fishbone tree structure construction and buffer position adjustment.
It reduces the offset of the clock tree, reduces the number of buffers, improves the comprehensive performance and flexibility of the clock tree, and meets the timing requirements of large-scale integrated circuits.
Smart Images

Figure CN119940284B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of integrated circuit physical design, and specifically relates to an efficient clock tree synthesis method based on complementary Manhattan rings. Background Art
[0002] In the design of digital integrated circuits, clock tree synthesis (CTS for short) is a crucial step and plays a vital role in timing design. The core task of the clock tree is to distribute the clock signal from the source node to each flip-flop (FF) and logic module in the circuit, while ensuring that the signal arrives to meet the corresponding timing requirements. The optimization of the clock tree directly affects the performance, power consumption, and area of the system. Especially in the design scenarios of high-frequency and large-scale integrated circuits, the performance of the clock tree determines the maximum operating frequency and signal synchronization of the entire system.
[0003] In traditional clock tree synthesis methods, the selection of the clock tree structure and the insertion strategy of buffers (BUF) are the key factors affecting the performance of the clock tree. Common clock tree structures include the H-tree and the fishbone tree, etc., each with its own advantages and disadvantages.
[0004] The H-tree structure is a very common one in clock tree synthesis. It has good symmetry, can effectively reduce clock skew, and achieve uniform transmission of the clock signal. The advantages of the H-tree are simple structure, easy to implement, and low clock skew when symmetrically laid out, which can well meet the timing requirements of high-frequency circuits. However, the H-tree has poor adaptability to complex circuits. As the circuit scale expands, the depth of the H-tree structure will increase, resulting in an increase in the delay of the clock tree and affecting the system performance.
[0005] The outstanding feature of the fishbone tree is to optimize the propagation path of the clock signal by means of multiple shorter paths, thereby significantly reducing the total delay of the clock signal. Compared with the traditional H-tree structure, the fishbone tree can reduce the long-term cumulative delay during signal transmission to a certain extent. Especially in multi-level clock paths, it can effectively reduce the propagation delay of the clock and improve the timing performance. However, the design focus of the fishbone tree is on the optimization of path delay, and there is no special balance adjustment for the clock arrival time of each branch on the clock tree, which may result in a worse clock skew control effect than the H-tree.
[0006] During the clock tree synthesis process, the insertion strategy of buffers also plays a crucial role in the performance of the clock tree. Generally speaking, buffer insertion aims to reduce signal propagation delay and lower the clock tree skew. However, the selection of the number, location, and type of buffers is a complex multi-objective optimization problem, involving trade-offs in multiple aspects such as power consumption, area, clock delay, and clock skew. Summary of the Invention
[0007] The present invention aims to solve at least one of the technical problems existing in the related art to a certain extent.
[0008] The object of the present invention is to provide an efficient clock tree synthesis method based on complementary Manhattan rings. Through an optimization strategy that combines bottom-up and top-down approaches, the key metrics of clock tree synthesis are comprehensively improved, and the flexibility of optimizing different key metrics is increased compared with traditional clock tree structures.
[0009] To achieve the above object, the present invention provides an efficient clock tree synthesis method based on complementary Manhattan rings, including:
[0010] S100. Divide the layout into multiple regions, pre-insert a number of buffers in each region, and initially form a clustering ring centered on each buffer. The size of the clustering ring of each buffer is complementary to the distance from the buffer to the center of the region;
[0011] S200. Circularly access each clustering ring, find clusters that can be merged until the number of clusters converges;
[0012] S300. Construct a clock tree using an offset-controllable fishbone tree structure;
[0013] S400. Use geometric operations to judge the overlapping situation of buffers and calculate the free area. By inflating and grid-dividing the free area, obtain the insertable positions of buffers, and then switch and adjust the buffer positions based on the relationship with the upper and lower level buffers.
[0014] A further preferred technical solution of the present invention is that step S100 specifically includes:
[0015] S110. Determine the number of divided regions according to the given layout length and width, the number of flip-flops, the maximum fan-out, and the maximum resistance-capacitance load, and perform region division;
[0016] S120. In each region, taking the center of the region as a reference, form equidistant lines with different distances from the center, and then uniformly pre-insert buffers on the equidistant lines;
[0017] S130. For each buffer, considering its distance from the center of the region, determine the cluster radius complementary to this distance, so that the delay from the region delay to each flip-flop is close;
[0018] S140. Assign triggers to each buffer within the set cluster radius to form a clustering ring.
[0019] Preferably, in step S200, each clustering ring is cyclically accessed to find clusters that can be merged. The method adopted is: cyclically access each clustering ring, continuously re - assign the triggers connected to the buffer with the least fan - out, search for other buffers, reduce the number of buffers, and achieve the merging of clusters.
[0020] Preferably, when constructing the clock tree using an offset - controllable fishbone tree structure in step S300, it is based on the path - optimal insertion method and the path - delay adjustment method.
[0021] Preferably, the specific steps of step S300 are as follows:
[0022] S310. Based on the path - optimal insertion method, plan the shortest - distance connection for the clustering ring farthest from the clock source, then uniformly insert buffers according to the distance to minimize the delay. Take this delay as the standard delay and take this connection path as the main path.
[0023] S320. Build branches on the basis of the main path and connect them to other clustering rings.
[0024] S330. Based on the path - delay adjustment method, adjust the positions of the buffers on the branches so that the delay of each clustering ring converges to the standard delay.
[0025] Preferably, when planning the shortest - delay path for the clustering farthest from the clock source in step S310, the method of uniformly inserting buffers is adopted. The number of buffers optimally inserted at a certain Manhattan distance The best number of buffers to insert The calculation formula is:
[0026]
[0027] Among them, is the resistance per unit length, is the capacitance per unit length, is the intrinsic delay of the buffer, is the Manhattan distance of the path between two points; is the variable of the number of inserted buffers. Since the inserted - number variable is an integer, the calculated is rounded up and down respectively for optimal selection.
[0028] Preferably, when adjusting the positions of the buffers on the path whose delay needs to be adjusted in step S330, the path is changed from a straight - line shape to a broken - line shape to increase the delay.
[0029] Beneficial effects: The efficient clock tree synthesis method based on complementary Manhattan rings of the present invention reduces the skew at the bottom layer of the clock tree through a complementary circular clustering method; by merging clusters, while meeting the performance requirements, the number of buffers required for constructing the clock tree at the bottom layer is minimized as much as possible, and the utilization rate of the maximum fan-out of the buffers is improved. By constructing a skew-controllable fishbone tree and combining a long-path delay optimization algorithm, while ensuring an extremely low clock skew, the delay is made as close to the theoretical limit as possible, and the number of buffers is as small as possible, realizing the common optimization of the three, and greatly improving the performance of clock tree synthesis. Brief Description of the Drawings
[0030] Figure 1 is a flowchart of the efficient clock tree synthesis method based on complementary Manhattan rings of the present invention;
[0031] Figure 2 is a schematic diagram of the complementary circular clustering of the present invention;
[0032] Figure 3 is a schematic diagram of the structure of the skew-controllable fishbone tree established by the present invention. Detailed Embodiment
[0033] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention, and they should not be construed as limiting the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention. In the description of the present invention, it should be understood that the terms used are only for the purpose of description and cannot be construed as indicating or implying relative importance.
[0034] The following is combined with Figures 1 - 3 to describe the efficient clock tree synthesis method based on complementary Manhattan rings provided by the present invention.
[0035] Embodiment: The purpose of this embodiment is to overcome the deficiencies in aspects such as clock path delay, clock skew control, and buffer quantity in existing clock tree optimization methods. Through innovative algorithms and optimization strategies, a new clock tree synthesis method based on complementary Manhattan rings is proposed, aiming to improve the performance of the clock tree and solve the challenges faced by traditional clock tree structures in large-scale integrated circuit design.
[0036] The overall process of this method is as Figure 1 shown and includes:
[0037] S100. Divide the layout into multiple regions, pre-insert several buffers in each region, initially form a clustering ring centered on each buffer, and the size of the clustering ring of each buffer is complementary to the distance from the buffer to the region center;
[0038] S200. Cyclically access each clustering ring, find clusters that can be merged until the number of clusters converges;
[0039] S300. Construct a clock tree using an offset-controllable fishbone tree structure;
[0040] S400. Use geometric operations to judge the overlap of buffers and calculate the available area. By expanding and meshing the available area, obtain the insertable positions of buffers, and then adjust the buffer positions based on the relationship with the upper and lower level buffers.
[0041] The following is a detailed description of each step:
[0042] Step S100 is used to form region-complementary circular clustering. Specifically, it includes:
[0043] S110. Determine the number of divided regions according to the given layout length and width, number of flip-flops, maximum fan-out, and maximum resistance-capacitance load, and perform region division;
[0044] S120. In each region, taking the region center as a reference, form equidistant lines with different distances to the center, and then uniformly pre-insert buffers on the equidistant lines;
[0045] S130. For each buffer, according to the principle of matching near and far and distance complementarity, determine the cluster radius so that the delay from the region delay to each flip-flop is close;
[0046] As Figure 2 shown, let and be the distances from two buffers to the region center respectively, and be the cluster radii of two buffers respectively, and the size of the flip-flop clustering ring of each buffer is complementary to the distance from the buffer to the center, expressed as:
[0047]
[0048] S140. Allocate flip-flops for each buffer within the set cluster radius range to form a clustering ring, ensuring a small offset between the bottom-level flip-flops.
[0049] Pre-inserting buffers can quickly form the initial connection of the clustering ring, avoiding non-convergence after delay iteration during clustering due to the uncertain position of the ring center.
[0050] Step S200 is used to merge clusters and reduce the number of clustering loops.
[0051] The method adopted is as follows: Based on the trigger clustering, find the buffers with few fan-outs, try to connect the triggers under them to other buffers, and judge whether the constraint requirements and performance requirements are met. Keep looping until the number of buffers converges or the maximum number of iterations is reached, minimizing the number of buffers and improving the utilization rate of the maximum fan-out.
[0052] Step S300 is used to construct a clock tree.
[0053] Traditional fishbone trees are not easy to control the offset and have poor structural flexibility. Therefore, in this implementation, a fishbone tree structure with controllable offset as shown in Figure 3 is constructed, which can more flexibly adjust the position of the buffer on the premise of ensuring low delay, realizing the reduction of the offset. The construction of the fishbone tree with controllable offset is based on the path-optimal insertion method and the path delay adjustment method. The path-optimal insertion method can be obtained by mathematical derivation, and the path delay adjustment method is the better practice obtained after trying various methods. The combination of these two methods can achieve delay convergence and offset reduction.
[0054] The specific steps include:
[0055] S310. Based on the path-optimal insertion method, use the fishbone tree structure with controllable offset to plan the shortest distance connection for the clustering loop farthest from the clock source, and then uniformly insert buffers according to the distance to minimize the delay. Take this delay as the standard delay and use this connection path as the main path;
[0056] According to the definition of buffer delay and wire unit delay, it can be obtained that under a fixed path, inserting a certain number of buffers can effectively reduce the delay. The derivation of the insertion quantity and position is as follows:
[0057] The delay in the case of not inserting a buffer , which is expressed as:
[0058]
[0059] Among them, is the resistance per unit length, is the capacitance per unit length, is the Manhattan distance of the path between two points;
[0060] In the case of inserting buffers, the delay of a certain path is expressed as:
[0061]
[0062] Among them, This is the delay of the buffer.
[0063] It can be obtained from the Lagrange multiplier method that when inserting buffers on the same path, the buffers are evenly distributed on the path, which can obtain the lowest delay under the current conditions. Therefore, the method of uniformly inserting buffers is adopted. The following is the derivation of the optimal number of insertions:
[0064] First calculate :
[0065]
[0066] Solve for a Manhattan distance The optimal number of insertions for:
[0067]
[0068] Re-order get:
[0069]
[0070] Since the number of variables is inserted is an integer, using The number of buffers inserted in the current path is determined by rounding up and down.
[0071] Therefore, in the subsequent clock tree synthesis algorithm, the piecewise function curve is first calculated using the numerical values given in the question, and the number of buffer insertions in different distance segments is adjusted using a lookup table.
[0072] S320, establishing a branch path based on the main path and connecting to other cluster rings;
[0073] S330, after completing the optimal insertion of the path, the delay needs to be converged according to the standard, and the delay of the branch path needs to be adjusted. According to the derivation in the above method, adjusting the position of the buffer can change the delay of the path. Therefore, the position of the buffer on the path that needs to adjust the delay is adjusted, so that the path changes from a straight line to a broken line, thereby increasing the delay.
[0074] Step S400 is used to adjust the buffer position and layout the clock tree.
[0075] During the clock tree layout process, geometric operations are performed on the plane area around the buffer to remove the space occupied by the trigger and other buffers to obtain the free area. The free area is expanded and segmented to extract the buffer insertion position. Different strategies are switched to select the insertion position according to the connection relationship of the buffer to ensure the optimization of key indicators of the clock tree in the local area.
[0076] In the above technical solution of this embodiment:
[0077] (1), Considering the area and power consumption problems caused by too many buffers, an iterative merging and clustering method is adopted at the bottom layer, so that the maximum fan-out is fully utilized, thereby reducing the number of buffers.
[0078] (2), Considering the optimization of clock skew, a ring-shaped cluster is formed during bottom-layer clustering to ensure that the skew between each flip-flop in the cluster is within a certain threshold, and according to the distance complementary relationship, the cluster radius size is adjusted to effectively control the clock delay under low clock skew.
[0079] (3), Considering the reduction of clock delay, while ensuring small skew, the method of minimizing the longest distance is adopted, and the delay of the farthest merged cluster is used as the standard for global delay convergence, reducing the average delay.
[0080] Through an optimization strategy that combines bottom-up and top-down, the three key metrics of clock tree synthesis are comprehensively improved. At the same time, compared with the traditional clock tree structure, the flexibility of optimizing different key metrics is improved.
[0081] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An efficient clock tree synthesis method based on complementary Manhattan rings, characterized in that, Including: S100: Divide the layout into multiple regions, pre-insert several buffers in each region, initially form a clustering ring centered on each buffer, and the size of the clustering ring of each buffer is complementary to the distance from the buffer to the region center; S200: Circularly access each clustering ring, find clusters that can be merged until the number of clusters converges; S300: Construct a clock tree using an offset-controllable fishbone tree structure; When constructing a clock tree using an offset-controllable fishbone tree structure, based on the path-optimal insertion method and the path delay adjustment method, specifically: S310: Based on the path-optimal insertion method, plan the shortest distance connection for the clustering ring farthest from the clock source, then evenly insert buffers according to the distance to minimize the delay, use this delay as the standard delay, and use this connection path as the main path; Among them, when planning the shortest delay path for the cluster farthest from the clock source, the method of uniformly inserting buffers is adopted, and a certain Manhattan distance The optimal number of inserted buffers The calculation formula is: ; Among them, is the resistance per unit length, is the capacitance per unit length, is the intrinsic delay of the buffer, is the Manhattan distance of the path between two points; is the variable of the number of inserted buffers. Since the variable of the inserted number is an integer, the calculated is rounded up and down respectively for optimal selection; S320: Establish branches on the basis of the main path to connect to other clustering rings; S330: Based on the path delay adjustment method, adjust the positions of the buffers on the branches so that the delay of each clustering ring converges to the standard delay; S400: Use geometric operations to judge the overlapping situation of the buffers and calculate the free area, obtain the insertable positions of the buffers by expanding and meshing the free area, and then adjust the buffer positions based on the relationship with the upper and lower buffers; 2. The high - efficiency clock tree synthesis method based on complementary Manhattan rings according to claim 1, wherein Step S100 specifically includes: S110: Determine the number of divided regions according to the given layout length and width, the number of flip-flops, the maximum fan-out, and the maximum resistance-capacitance load, and perform region division; S120: In each region, with the region center as the reference, form equidistant lines with different distances to the center, and then evenly pre-insert buffers on the equidistant lines; S130: For each buffer, consider its distance to the region center, determine the cluster radius complementary to this distance, so that the delay from the region delay to each flip-flop is close; S140: Allocate flip-flops for each buffer within the set cluster radius to form a clustering ring; 3. The high-efficient clock tree synthesis method based on a complementary Manhattan ring according to claim 1, wherein In step S200, when circularly accessing each clustering ring to find clusters that can be merged, the method used is: circularly access each clustering ring, continuously re-allocate the flip-flops connected under the buffer with the least fan-out, find other buffers, reduce the number of buffers, and achieve the merger of clusters; 4. The high - efficiency clock tree synthesis method based on complementary Manhattan rings according to claim 1, wherein When adjusting the positions of the buffers on the path that needs to adjust the delay in step S330, change the path from a straight line to a broken line to increase the delay.
Citation Information
Patent Citations
Method of constructing hierarchical clock tree for integrated circuit
CN112100971A
Method and device for constructing clock tree and chip
CN118780238A