Hierarchical clock tree synthesis method and system based on hybrid topological structure

Through a hierarchical clock tree synthesis method based on a hybrid topology structure, a balanced greedy search clustering algorithm and an adaptive H-tree construction algorithm are adopted, combined with a post-processing strategy, to solve the clock deviation and delay optimization problems in the existing CTS method and achieve efficient collaborative optimization of the clock network.

CN120850901APending Publication Date: 2025-10-28FUZHOU UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510945937.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

The existing CTS method is difficult to optimize local clock deviation and delay at the same time. The clock tree structure has a single topology, which makes it difficult to meet the comprehensive design requirements of different regions and multiple levels. It also has problems such as high power consumption and difficult timing analysis.

Method used

A hierarchical clock tree synthesis method based on a hybrid topology structure is adopted, including a balanced greedy search clustering algorithm, an adaptive H-tree construction algorithm and a post-processing strategy. Clock skew and delay are optimized through local clustering, the topology is dynamically adjusted, and buffers are inserted to optimize clock network performance.

Benefits of technology

Under limited resource constraints, clock skew, clock delay, and buffer insertion number are optimized collaboratively to improve the overall performance and design efficiency of the clock network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120850901A_ABST
    Figure CN120850901A_ABST
Patent Text Reader

Abstract

The invention provides a hierarchical clock tree synthesis method and system based on a hybrid topological structure, and aims to collaboratively optimize clock skew and clock delay in a chip clock network. The method comprises the following steps: firstly, introducing a balanced greedy search clustering algorithm, clustering registers at the initial stage of clock network construction, and reducing local clock skew and delay; and generating a bottom layer clock tree in each cluster by adopting a bounded deviation construction method so as to ensure the clock synchronization quality in the cluster. In a top-layer clock tree construction stage, the invention provides a self-adaptive H-tree topology construction algorithm, and the H-tree structure is dynamically adjusted according to the position relationship among all layers of clusters, so that a global clock distribution path is optimized under multiple layers. Besides, in order to further improve the CTS quality, the invention further designs a clock network optimization strategy based on buffer post-processing, and local buffer insertion and position adjustment are carried out for the problems of fan-out unevenness and load overrun so as to eliminate potential illegal paths.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of integrated circuit design and electronic automation technology, and in particular to a hierarchical clock tree synthesis method and system based on hybrid topology. Background Technology

[0002] With the continuous evolution of integrated circuit manufacturing processes, chip size and integration density have increased dramatically, leading to a continuous increase in the scale and hierarchical complexity of on-chip clock networks. As a crucial step in ensuring the quality of synchronization signal distribution, Clock Tree Synthesis (CTS) places higher demands on the balance optimization between clock skew and clock delay. However, existing CTS methods generally face the following problems in practical applications: difficulty in simultaneously optimizing local clock skew and delay, and a single clock tree structure topology that is difficult to meet the synthesis design requirements of different regions and multiple levels. Although some have proposed clock trees combining mesh and tree topologies, these suffer from high power consumption and difficulty in timing analysis. Summary of the Invention

[0003] In view of this, the purpose of this invention is to provide a hierarchical clock tree synthesis method and system based on a hybrid topology, so as to achieve collaborative optimization of clock skew and clock delay in the chip clock network, thereby improving the overall performance and design efficiency of CTS.

[0004] To achieve the above objectives, the present invention adopts the following technical solution: a hierarchical clock tree synthesis method based on a hybrid topology, comprising the following steps:

[0005] Step 1: The balanced greedy search clustering algorithm is used to divide all clock pins into multiple local clusters based on the spatial distribution and load characteristics of the clock pins. Within each local cluster, the bounded bias clock tree (BST) construction algorithm is run to reduce intra-cluster clock bias and delay.

[0006] Step 2: Using the BST root node as the clock pin at the top level, apply the adaptive H-tree construction algorithm to dynamically adjust the topology of each layer, thereby achieving coordinated optimization of the deviation and delay of the global clock network.

[0007] Step 3: Fine-tune buffer fanout and load violations through post-processing strategies to eliminate design rule violations and further improve clock network performance, thereby optimizing clock skew, clock delay and buffer insertion number under limited resource constraints.

[0008] In a preferred embodiment, in step 1, the balanced greedy search clustering algorithm selects a register as the center of the first cluster from the given set of clock pins, and records the cluster center position and fan-out; then, based on the size of the design region and the clustering factor, a maximum running cluster distance upper limit is calculated, as follows:

[0009]

[0010] Among them, Dis max The calculated maximum radius is represented by Width and Length, which represent the current wiring area size. Total cap Max represents the total capacitance of the clock pin. f The value represents the number of clock pins that the smallest buffer in the LIB library can drive, and α represents the weighting coefficient; Formula (2) considers the calculation of the cost for each cluster:

[0011] C (i,j) =dis i,j +β×(Max f -C size ) 2 (2)

[0012] Among them, C i,j This indicates the cost calculation for adding the current clock pin to the cluster, dis i,j Max represents the Manhattan distance between the current node and the cluster center. f C represents the maximum number of driveable pins of the buffer. size This indicates the current cluster size, and β is the weighting coefficient. Each pin will be added to the cluster with the lowest cost.

[0013] For each subsequent clock pin node, calculate its Manhattan distance to each cluster center. If the fanout of a cluster does not reach the set upper limit and the distance is within the allowable range, add the clock pin to this cluster and update the center point position, taking the center of all clock pins in the cluster as the new cluster center. If no suitable cluster can be found, create a new cluster with the current clock pin as the cluster center. At the same time, merge unbalanced clusters with the nearest cluster, while ensuring that the merged cluster still meets the fanout constraint.

[0014] In a preferred embodiment, in step 2, the adaptive H-tree construction algorithm dynamically selects the horizontal and vertical extension lengths of each H-type structure during the construction process and adaptively divides the subtree according to the clock pin distribution contained in the current subtree; the adaptive H-tree construction algorithm greedily selects the optimal combination of delay and routing resources under the current division at each division stage.

[0015] First, determine the depth of the adaptive H-tree using the following formula:

[0016]

[0017] Where D represents the depth of the adaptive H-tree, and N represents the number of clock pins in the chip; after determining the depth of the H-tree according to formula (3), each layer of the H-tree is constructed from bottom to top. In the construction of each layer, the clock tree is divided into several regions, and the optimal solution of each layer of the H-tree is calculated by the cost function (4). The process is iterated until the maximum depth is reached; where t and s represent the clock delay and clock deviation of the currently constructed clock tree, respectively. max and s max The maximum clock delay and maximum clock deviation are set, with δ1 and δ2 being weighting coefficients.

[0018] In a preferred embodiment, in step 3, the clock network load optimization strategy divides the clock network with load violations into a first subnet and a second subnet by inserting buffers; checks whether the connection from each node in the clock network to its child node exceeds a preset maximum load constraint value, and takes measures to reduce the load size of connections exceeding this threshold; when a load violation is detected, the number and location of buffers to be inserted at appropriate positions are calculated based on the distance between the current node and its child nodes; the insertion of buffers aims to divide long wires into multiple shorter segments; the number of segments n is used to find the optimal solution through iterative loops to minimize the total delay; the clock delay TotalDelay of each segment is calculated by the following formula:

[0019]

[0020] Among them, R unit and C unit These represent unit resistance and unit capacitance, respectively; dis is the Manhattan distance between the current node and its child nodes; and Buf is the distance between the current node and its child nodes. delay Indicates the internal delay of the buffer;

[0021] Continuously adjust the number of segments n and find the optimal value that minimizes the total delay.

[0022] In a preferred embodiment, net fanout optimization involves first calculating and removing nodes in non-compliant clock nets that require adjustment, then selecting the nearest clock net that meets the net fanout requirements, and reassigning these nodes to it.

[0023] This invention also provides a hierarchical clock tree synthesis system based on a hybrid topology, which runs the aforementioned hierarchical clock tree synthesis method based on a hybrid topology.

[0024] Compared with the prior art, the present invention has the following advantages: it can coordinately optimize clock skew, clock delay and buffer insertion number under limited resource constraints. Attached Figure Description

[0025] Figure 1This is a flowchart of the algorithm of a preferred embodiment of the present invention;

[0026] Figure 2 The BGSR clustering algorithm flow is a preferred embodiment of the present invention. Detailed Implementation

[0027] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0028] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0029] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations according to this application; as used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise; furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0030] A hierarchical clock tree synthesis method based on hybrid topology, refer to Figure 1-2 , specifically including the following steps:

[0031] Step 1: To reduce the running time of the algorithm and decrease the size of the clock pins during the operation of the top-level algorithm, a balanced greedy search clustering algorithm is used to divide all clock pins into multiple local clusters based on the spatial distribution and load characteristics of the flip-flops. Within each cluster, a bounded bias clock tree construction algorithm is run, and buffers are inserted appropriately to reduce the clock bias and delay within the cluster.

[0032] Step 2: To reduce the overall skewness of the constructed clock tree, the top-level clock topology adopts an H-tree topology to achieve balanced driving of multiple clock regions and low clock skewness. However, traditional H-tree construction methods only pre-define a fixed geometric structure based on chip size and the number of clock pins, failing to dynamically adjust the clock topology according to the actual geometric distribution of clock pins. This can easily lead to clock delay imbalance, redundant trace lengths, or low buffer clock efficiency. This strategy dynamically selects the horizontal and vertical extension lengths of each H-type structure during the construction process and can adaptively partition according to the clock pin distribution contained in the current subtree. Specifically, at each partitioning stage, the algorithm greedily selects the optimal combination of delay and routing resources under the current partition, thereby ensuring that the global clock network achieves a better balance between clock skewness, clock delay, and the number of buffers.

[0033] Step 3: To address the issues of buffer fan-out and load violations, this invention proposes a comprehensive post-processing optimization strategy. By inserting buffers into the network, the non-compliant wires are segmented to reduce the load on a single segment, and the optimal number of segments is selected based on iterative optimization to minimize the total delay. Simultaneously, for network fan-out violations, the fan-out structure is adjusted and the load is balanced by inserting new buffers or reallocating nodes to adjacent compliant networks, thereby improving the timing performance of clock signal transmission and network stability.

[0034] Specifically, the balanced greedy clustering algorithm:

[0035] First, the algorithm selects a register from the given set of clock pins as the center of the first cluster, recording the cluster center position and fan-out. Then, based on the size of the design region and the clustering factor, it calculates an upper limit for the maximum running cluster distance, as follows:

[0036]

[0037] In formula (1) Dis max The calculated maximum radius is represented by Width and Length, which represent the current wiring area size. Total cap Max represents the total capacitance of the clock pin. f This represents the number of clock pins that the smallest buffer in the LIB library can drive, and α represents the weighting coefficient. The following formula takes into account the cost calculation for each cluster:

[0038] C (i,j) =dis i,j +β×(Max f -C size ) 2 (2)

[0039] Among them, C i,j This indicates the cost calculation for adding the current clock pin to the cluster, dis i,j Max represents the Manhattan distance between the current node and the cluster center. f C represents the maximum number of driveable pins of the buffer. size This indicates the current cluster size, and β is the weighting coefficient. Each pin will be added to the cluster with the lowest cost.

[0040] like Figure 2As shown, for each subsequent clock pin node, the Manhattan distance to each cluster center is calculated. If the fanout of a cluster does not reach the set upper limit, and the distance is within the allowable range, the clock pin is added to this cluster, and the center point position is updated (the center of all clock pins in the cluster is taken as the new cluster center). If no suitable cluster can be found, a new cluster is created, with the current clock pin as the cluster center. Simultaneously, to balance the load driven by the underlying buffer, the algorithm merges unbalanced clusters with the nearest cluster, while ensuring that the merged cluster still meets the fanout constraint. This improves cluster utilization and further reduces clock skew.

[0041] Specifically, the adaptive H-tree construction:

[0042] To reduce the overall skewness of the constructed clock tree, the top-level clock topology adopts an H-tree topology to achieve balanced driving of multiple clock regions and low clock skewness. However, traditional H-tree construction methods only pre-define a fixed geometric structure based on chip size and the number of clock pins, failing to dynamically adjust the clock topology according to the actual geometric distribution of clock pins. This can easily lead to clock delay imbalance, redundant trace lengths, or low buffer clock efficiency. This strategy dynamically selects the horizontal and vertical extension lengths of each H-type structure during the construction process and can adaptively partition according to the clock pin distribution contained in the current subtree. Specifically, at each partitioning stage, the algorithm greedily selects the optimal combination of delay and routing resources under the current partition, thereby ensuring that the global clock network achieves a better balance between clock skewness, clock delay, and the number of buffers.

[0043] First, determine the depth of the adaptive H-tree using the following formula.

[0044]

[0045] Where D represents the depth of the adaptive H-tree, and N represents the number of clock pins in the chip. This strategy determines the depth of the H-tree according to formula (3) and then constructs each layer of the H-tree from bottom to top. In the construction of each layer, the clock tree is divided into several regions, and the optimal solution of each layer of the H-tree is calculated by the cost function (4). This process iterates until the maximum depth is reached. Where t and s represent the clock delay and clock skew of the currently constructed clock tree, respectively. max and s max The maximum clock delay and maximum clock deviation are set, with δ1 and δ2 being weighting coefficients.

[0046] Specifically, the clock network load optimization strategy of this invention divides clock networks with load violations into two subnets (Network 1 and Network 2) by inserting a buffer. This division effectively reduces the load of each subnet, thereby reducing clock latency and optimizing the timing characteristics of the entire clock tree. This method can effectively improve the signal integrity of the clock network, thus making the clock distribution network more efficient and robust.

[0047] Specifically, the algorithm checks whether the connections from each node to its child nodes in the clock network exceed a preset maximum load constraint value. For connections exceeding this threshold, measures must be taken to reduce their load size to ensure that timing constraints are not violated. When a load violation is detected, the algorithm calculates the number and location of buffers to insert at appropriate positions based on the distance between the current node and its child nodes. Buffer insertion aims to divide long wires into multiple shorter segments, thereby reducing the load size of each segment. Specifically, the number of segments, n, is iteratively calculated to find the optimal solution that minimizes the total delay. The clock delay of each segment is calculated using the following formula:

[0048]

[0049] Among them, R unit and C unit These represent unit resistance and unit capacitance, respectively; dis is the Manhattan distance between the current node and its child nodes; and Buf is the distance between the current node and its child nodes. delay This indicates the internal delay of the buffer.

[0050] To minimize the total load, the algorithm continuously adjusts the number of segments, n, and searches for the optimal value that minimizes the total delay. This process is achieved through iterative loops, thus ensuring the timing constraints of the clock tree while optimizing signal transmission delay. Through this load optimization strategy, long wires in the clock tree are effectively segmented, significantly reducing the load on the wires and consequently reducing signal transmission delay, thereby improving the performance and reliability of the clock tree.

[0051] Specifically, this invention provides an in-depth analysis of the fan-out distribution of clock networks and proposes improvement measures for clock networks with fan-out violations. One effective method is to insert buffers at appropriate locations to adjust the fan-out, or to improve the fan-out distribution by reallocating nodes. In some cases, if a network with compliant fan-out exists near the violating network, the nodes of the violating sub-network can be reallocated. If no compliant network exists near the violating network, a new buffer is inserted to resolve the violation. Specifically, first, the nodes that need adjustment in the violating clock network are calculated and removed. Then, the nearest clock network that meets the fan-out requirements is selected, and these nodes are reallocated to it. Through this optimization process, the node distribution of each sub-network can be made more uniform, avoiding excessive concentration or dispersion of nodes, thereby achieving a balance in signal transmission load.

Claims

1. A hierarchical clock tree synthesis method based on a hybrid topology, characterized in that, Includes the following steps: Step 1: The balanced greedy search clustering algorithm is used to divide all clock pins into multiple local clusters based on the spatial distribution and load characteristics of the clock pins. Within each local cluster, the bounded bias clock tree (BST) construction algorithm is run to reduce intra-cluster clock bias and delay. Step 2: Using the BST root node as the clock pin at the top level, apply the adaptive H-tree construction algorithm to dynamically adjust the topology of each layer, thereby achieving coordinated optimization of the deviation and delay of the global clock network. Step 3: Fine-tune buffer fanout and load violations through post-processing strategies to eliminate design rule violations and further improve clock network performance, thereby optimizing clock skew, clock delay and buffer insertion number under limited resource constraints.

2. The hierarchical clock tree synthesis method based on a hybrid topology according to claim 1, characterized in that, In step 1, the balanced greedy search clustering algorithm selects a register from the given set of clock pins as the center of the first cluster, and records the cluster center position and fan-out. Then, based on the size of the design region and the clustering factor, it calculates an upper limit of the maximum running cluster distance, as follows: Among them, Dis max The calculated maximum radius is represented by Width and Length, which represent the current wiring area size. Total cap Max represents the total capacitance of the clock pin. f The value represents the number of clock pins that the smallest buffer in the LIB library can drive, and α represents the weighting coefficient; Formula (2) considers the calculation of the cost for each cluster: C (i,j) =dis i,j +β×(Max f -C size ) 2 (2) Among them, C i,j This indicates the cost calculation for adding the current clock pin to the cluster, dis i,j Max represents the Manhattan distance between the current node and the cluster center. f C represents the maximum number of driveable pins of the buffer. size This indicates the current cluster size, and β is the weighting coefficient. Each pin will be added to the cluster with the lowest cost. For each subsequent clock pin node, calculate its Manhattan distance to each cluster center. If the fanout of a cluster does not reach the set upper limit and the distance is within the allowable range, add the clock pin to this cluster and update the center point position, taking the center of all clock pins in the cluster as the new cluster center. If no suitable cluster can be found, create a new cluster with the current clock pin as the cluster center. At the same time, merge unbalanced clusters with the nearest cluster, while ensuring that the merged cluster still meets the fanout constraint.

3. The hierarchical clock tree synthesis method based on a hybrid topology according to claim 1, characterized in that, In step 2, the adaptive H-tree construction algorithm dynamically selects the horizontal and vertical extension lengths of each H-type structure during the construction process and adaptively divides the subtree according to the clock pin distribution contained in the current subtree; the adaptive H-tree construction algorithm greedily selects the optimal combination of delay and routing resources under the current division in each division stage. First, determine the depth of the adaptive H-tree using the following formula: Where D represents the depth of the adaptive H-tree, and N represents the number of clock pins in the chip; after determining the depth of the H-tree according to formula (3), each layer of the H-tree is constructed from bottom to top. In the construction of each layer, the clock tree is divided into several regions, and the optimal solution of each layer of the H-tree is calculated by the cost function (4). The process is iterated until the maximum depth is reached; where t and s represent the clock delay and clock deviation of the currently constructed clock tree, respectively. max and s max The maximum clock delay and maximum clock deviation are set, with δ1 and δ2 being weighting coefficients.

4. The hierarchical clock tree synthesis method based on a hybrid topology according to claim 1, characterized in that, In step 3, the clock net load optimization strategy divides clock nets with load violations into a first subnet and a second subnet by inserting buffers. It checks whether the connections from each node to its child nodes in the clock net exceed a preset maximum load constraint value. For connections exceeding this threshold, measures are taken to reduce their load. When a load violation is detected, the number and location of buffers to insert are calculated based on the distance between the current node and its child nodes. The insertion of buffers aims to divide long wires into multiple shorter segments. The number of segments, n, is iteratively searched to find the optimal solution that minimizes the total delay. The clock delay (TotalDelay) of each segment is calculated using the following formula: Among them, R unit and C unit These represent unit resistance and unit capacitance, respectively; dis is the Manhattan distance between the current node and its child nodes; and Buf is the distance between the current node and its child nodes. delay Indicates the internal delay of the buffer; Continuously adjust the number of segments n and find the optimal value that minimizes the total delay.

5. The hierarchical clock tree synthesis method based on a hybrid topology according to claim 4, characterized in that, Net fanout optimization: First, calculate and remove the nodes that need adjustment in the non-compliant clock nets. Then, select the nearest clock net that meets the net fanout requirements and reassign these nodes to it.

6. A hierarchical clock tree synthesis system based on a hybrid topology, characterized in that, Run the hierarchical clock tree synthesis method based on hybrid topology as described in any one of claims 1-5 above.

Citation Information

Cited By

  • Clock tree synthesis method and computer readable storage medium

    CN121168368A

  • Clock tree synthesis method based on collaborative optimization of clustering and reinforcement learning

    CN121351757A