Clock tree synthesis method based on synchronous concurrence and hierarchical clustering
By adopting a clock tree synthesis method based on synchronous concurrency and hierarchical clustering in the digital integrated circuit design, using the KmeansPlus algorithm and multi-threading technology, the problem of insufficient comprehensive speed and accuracy of large-scale digital integrated circuit clock tree is solved, and the rapid and accurate clock tree synthesis and design process is achieved.
Patent Information
- Application Number
- CN202510101996.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The prior art is difficult to quickly and accurately perform clock tree synthesis of large-scale digital integrated circuits, resulting in time-consuming and low quality design processes.
The clock tree synthesis method based on synchronous concurrency and hierarchical clustering is adopted, and coarse clustering and fine clustering are performed through the KmeansPlus algorithm, combined with multi-thread acceleration to achieve fast and accurate clock tree synthesis.
It improves the running speed and accuracy of the clock tree synthesis, reduces the time-consuming design process, and avoids the trap of local optimal solutions.
Smart Images

Figure CN120030979A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of integrated circuits, and in particular relates to a clock tree synthesis method based on synchronous concurrency and hierarchical clustering. Background Art
[0002] Clock tree synthesis is an important step in digital integrated circuit design. By quickly and accurately synthesizing and optimizing the clock tree, the speed and reliability of digital integrated circuit design can be improved. Digital integrated circuit clock tree synthesis refers to connecting the laid-out registers to the clock source through clustering and balancing. Clustering is to divide the laid-out registers into multiple groups to form leaf nodes, which converge to the corresponding buffers. The buffers are clustered and finally converge to the root node (clock source). Balancing is the optimization balance of the clock tree. While ensuring that the number of loads carried by the buffer does not exceed the maximum fan-out and the load capacitance does not exceed the maximum load capacitance, the clock skew (Skew) and delay (Latency) from the clock to each child node and the number of buffers N are reduced as much as possible. In digital circuits, clock tree synthesis has a great impact on circuit performance. It determines the stability and reliability of circuit clock propagation. Clock tree synthesis is one of the key steps in digital circuit design, which is conducive to ensuring the correct function and performance of the circuit, thereby shortening product development time and reducing costs and improving product quality.
[0003] Digital EDA tools are usually used to design and verify the correctness of digital circuits. These tools can help engineers simulate, test and optimize the performance of circuits. However, as the scale of integrated circuits in digital chips becomes larger and larger, and the complexity of digital circuits increases, the speed and stability of clock tree synthesis have become a big problem. EDA tools cannot adapt to the requirements of large-scale integrated circuit clock tree synthesis, resulting in poor quality of clock tree synthesis, repeated digital chip design processes, and the entire process is very time-consuming. The key point of clock tree synthesis is clustering. The clustering algorithm is a method for optimizing buffers. All registers are clustered into several groups and a clock buffer is assigned to each group. EDA tools use efficient, stable and accurate clustering algorithms to greatly improve the speed and efficiency of clock tree synthesis, thereby accelerating digital chip development and shortening the time to market for products.
[0004] The traditional Kmeans algorithm randomly selects K initial cluster centers, then assigns each data point to the nearest cluster center, and iteratively calculates each cluster center until the cluster center no longer changes significantly or the maximum number of iterations is reached. The selection of the initial cluster in Kmeans will affect the final clustering results, causing the algorithm to fall into a local optimal solution. And with the increase of circuit elements in VLSI, the traditional Kmeans method of serial clustering will result in huge time overhead in processing huge registers, and a fast and accurate method is urgently needed to achieve high-quality clustering. Summary of the invention
[0005] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a clock tree synthesis method based on synchronous concurrency and hierarchical clustering. Starting from reading the layout file, the number of coarse clustering target groups and the threshold of the number of single-group registers are obtained according to the total number of registers and the number of available threads. Then, the total registers insts, the number of coarse clustering target groups n, the maximum number of iterations T, the random seed value seed and other information are input to the KmeansPlus algorithm to perform coarse clustering grouping of registers, so that the number of registers in each group is lower than the threshold of the number of registers in a single group. The number of coarse clusters in each group is combined with the maximum fanout max_fanout of the buffer used Buffer and the scaling ratio ratio to calculate the number of fine clustering target groups divide_number, and the target group number divide_number, the maximum number of iterations T, the random seed value seed and other information are input into the synchronous concurrent KmeansPlus algorithm for fine clustering. This method combines the advantages of multi-threaded acceleration and the intelligent selection of the center point of KmeansPlus to achieve accurate and fast clock tree synthesis from the register level. Compared with the traditional clock tree synthesis method, the running speed and accuracy of the clock tree synthesis are improved.
[0006] The present invention adopts the following technical solutions to achieve the above-mentioned invention object:
[0007] The clock tree synthesis method based on synchronous concurrency and hierarchical clustering includes the following steps:
[0008] Step S1, select whether the clustering object at this level is a register or a buffer according to the hierarchy, read in the physical information of all registers and clock sources in the layout file, start clustering at the register level as the first level and build a tree, and obtain the target grouping number and single group register number threshold according to the total number of registers and the number of available threads; for the level read in as the buffer in the hierarchical clustering iteration step, the same is true, and the target grouping number and single group buffer number threshold are obtained first.
[0009] Step S2: According to the single group register quantity threshold, the total registers insts, the target group quantity n, the maximum number of iterations T, and the random seed value seed are input to the KmeansPlus algorithm to perform coarse clustering and grouping of the read registers, and it is determined whether each group result of the coarse clustering is less than the single group quantity threshold. If it is less than the single group quantity threshold, the result is returned for the next step of fine clustering, otherwise the coarse clustering is performed again;
[0010] Step S3, extracting the grouping results described in step S2, obtaining the number of groups divide_number according to the number of registers in each group and the maximum fan-out max_fanout of the buffer to be used after fine clustering and the scaling ratio ratio, and inputting the register information of each group, the target number of groups divide_number, the maximum number of iterations T, the random seed value seed and other physical information into the synchronous and concurrent KmeansPlus algorithm for fine clustering;
[0011] Step S4, extract the fine clustering convergence point in step S3, perform overlap judgment and line length judgment on the convergence point, determine that after the coordinates of the convergence point are combined with the buffer size, there is no overlap with other registers of the lower level, and the line length of all leaf nodes connected to the convergence point does not exceed the maximum line length; use the buffer Buffer to instantiate the convergence point, set the sub-units of the buffer to all leaf nodes connected to the convergence point; the buffer results obtained by fine clustering of all groups at this level are used as the next level input of the hierarchical clock tree synthesis.
[0012] Step S5, extract the result of the first-level synthesis in step S4 as the input of the coarse clustering in the second-level clock tree synthesis, iterate steps S1-S4 to obtain the second-level synthesis result; perform bottom-up hierarchical clustering from registers to clock sources, and connect the buffers of leaf nodes and aggregation nodes step by step until the buffers of the last level converge to the clock source.
[0013] As a further optimization scheme of the clock tree synthesis method based on synchronous concurrency and hierarchical clustering, the specific method of obtaining the number of coarse clustering target groups and the threshold of a single group of registers according to the number of registers and the number of available threads in step S1 is as follows: the number of coarse clustering target groups is obtained according to the number of available threads and the number of groups processed by each thread, and a reasonable grouping threshold is calculated according to the size of the registers in the read layout information and the number of coarse clustering target groups to achieve a smaller number of coarse clustering iterations, and the number of iterations when fine clustering is performed on the coarse clustering results is also controlled within a reasonable range.
[0014] As a further optimization scheme of the clock tree synthesis method based on synchronous concurrency and hierarchical clustering, step S2 inputs the total registers insts, the target number of groups n, the maximum number of iterations T, and the random seed value seed into the KmeansPlus algorithm according to the threshold of the number of single-group registers to perform coarse clustering on the read-in registers. Only when all the grouping results are less than the threshold of the number of single-group registers can fine clustering be performed. Otherwise, all coarse clustering needs to be performed again. The threshold of the number of single-group registers and the target number of coarse clustering groups are adjusted until the grouping threshold is met.
[0015] As a further optimization scheme of the clock tree synthesis method based on synchronous concurrency and hierarchical clustering, step S3 uses the number of registers insts_number after coarse clustering, the maximum fan-out max_fanout of the buffer Buffer used by the instantiated convergence point after clustering, and the scaling ratio ratio to obtain the number of fine clustering groups divide_number, and inputs the number of fine clustering groups divide_number, the maximum fan-out max_fanout, the number of iterations T, and the random seed value seed into the synchronous concurrent KmeansPlus algorithm, and performs fine clustering on the registers after coarse clustering to obtain the convergence point of each group of clusters.
[0016] As a further optimization scheme of the clock tree synthesis method based on synchronous concurrency and hierarchical clustering, step S4 calculates whether the buffer and the connected register overlap based on the convergence point coordinates obtained by fine clustering and the size of the buffer used, and calculates whether the Manhattan distance between the buffer and the connected register exceeds the maximum line length max_net. If overlap occurs or the maximum line length is exceeded, the convergence point position needs to be fine-tuned.
[0017] As a further optimization scheme of the clock tree synthesis method based on synchronous concurrency and hierarchical clustering, step S5 extracts the result of the first-level synthesis in step S4 as the input of the coarse clustering in the second-level clock tree synthesis, and iterates steps S1-S5 to obtain the second-level synthesis result until the convergence point is connected to the clock source, that is, the complete clock tree synthesis is achieved. For the first-level clustering, the leaf nodes are registers, and the leaf nodes of the second-level and above clustering are buffers. The target grouping number, single-group register threshold, and read-in register information in steps S1-S5 can be replaced with buffers accordingly.
[0018] A computer-readable storage medium stores a computer program, which executes the steps of the above-mentioned path timing prediction method when the computer program is running.
[0019] The present invention adopts the above technical solution and has the following beneficial effects:
[0020] (1) The present invention proposes a clock tree synthesis method based on synchronous concurrency and hierarchical clustering. The synchronous concurrent KmeansPlus method is used to perform complete clock tree synthesis on the post-layout registers. According to the method of hierarchically constructing the clock tree from bottom to top, the registers / buffers are coarsely clustered and grouped using the KmeansPlus method at each level. The coarse clustering results that are less than the threshold of the number of registers in a single group enter the thread pool through the fine clustering task. The synchronous concurrent KmeansPlus algorithm is applied to finely cluster each group of registers / buffers to obtain the group of convergence points, which are instantiated as convergence buffers. The leaf nodes such as the registers / buffers at this level are connected to the convergence buffers to construct a tree. The buffers of all groups at this level are used as the input of the next level of coarse clustering. The clock tree is constructed from bottom to top until the clock source, thereby realizing fast and accurate clock tree synthesis in digital circuits and improving the speed of the digital chip design process.
[0021] (2) The present invention adds synchronous and concurrent multi-threaded processing to the traditional KmeansPlus algorithm, and locks the fine clustering results as shared data to achieve data protection; the multi-threaded synchronous and concurrent KmeansPlus algorithm maximizes resource utilization compared to traditional serial clustering, and realizes multiple groups of clustering in the same time period through parallel computing when performing fine clustering on large-scale registers / buffers, accelerating the bottom-up hierarchical clustering clock tree. Compared with the traditional Kmeans method, it improves the overall running speed of the clock tree while adding intelligent selection of the center point to avoid falling into the local optimal solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings are used to constitute a part of the description to facilitate a further understanding of the present invention, and are used together with the embodiments of the present invention to explain the present invention but do not constitute a limitation of the present invention;
[0023] Figure 1 It is a flow chart of a clock tree synthesis method based on synchronous concurrency and hierarchical clustering of the present invention;
[0024] Figure 2 is a schematic diagram of the synchronous concurrent KmeansPlus algorithm of the present invention;
[0025] Figure 3 It is a detailed schematic diagram of the present invention regarding coarse clustering and fine clustering;
[0026] Figure 4 It is a schematic diagram of the thread pool solution in the synchronous concurrent KmeansPlus algorithm of the present invention. DETAILED DESCRIPTION
[0027] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0028] In order to make the purpose, technical solution and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work should belong to the protection scope of the present invention.
[0029] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the purpose of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0030] The present invention is further described below in conjunction with the accompanying drawings. The clock tree synthesis method based on synchronous concurrency and hierarchical clustering is as follows Figure 1 As shown, it includes S1 to S5.
[0031] S1. According to the read-in layout file, read the location information of the layout area, registers and clock sources and convert them into coordinates, read the physical information such as unit resistance, unit capacitance, unit line length delay, and read the constraint information such as the maximum fanout max_fanout of the buffer Buffer and the maximum path delay. For the first-level registers of the clock tree comprehensive hierarchical clustering, the target group number is obtained based on the number of registers, the number of threads that can be used, and the number of groups that each thread is expected to process. Then, the single group register number threshold is obtained by combining the number of registers and the target group number. The threshold is expressed as formula (1):
[0032] insts_each=insts_number / n (1)
[0033] Where insts_each is the threshold value of the number of registers in a single group, insts_number is the total number of registers, and n is the number of coarse clustering target groups. For buffers at other levels of the clock tree, the same is true for calculating the threshold value of the number of buffers in a single group.
[0034] For example, for registers with an order of magnitude of 100k, 8 threads can be used, each thread processes the clustering of two groups of registers / buffers in fine clustering, the target number of groups is 16, and the grouping threshold is set to 20,000. After testing, this setting can meet the requirements through one coarse clustering, and the time overhead is much lower than that of the traditional serial KmeansPlus method.
[0035] S2. Input the total register / buffer coordinates and name information of the level input insts, the target grouping number n, the maximum number of iterations T, and the seed value seed for randomly generating points into the KmeansPlus algorithm for coarse clustering to obtain the registers / buffers after preliminary grouping of coarse clustering.
[0036] S3. Determine the number of registers / buffers in a single group after coarse clustering. If the number of each group is lower than a threshold of the number of registers / buffers in a single group, return the result for the next fine clustering. Otherwise, perform coarse clustering again until the threshold is met.
[0037] S4. Obtain the number of groups divide_number according to the number of registers / buffers in each group in the thread pool, the maximum fan-out max_fanout of the buffer Buffer used by the instantiated clustering convergence point, and the scaling ratio ratio, as shown in formula (2).
[0038] divide_number=insts_number / (max_fanout×ratio) (2)
[0039] By using the scaling ratio and the maximum fan-out max_fanout, the number of leaf nodes connected to each group of convergence points in fine clustering does not exceed the maximum fan-out, and the number of convergence point connections is less than or equal to the maximum fan-out by adjusting the scaling ratio. When the scaling ratio is 1, the maximum fan-out of each buffer can be fully utilized. The number of registers / buffers insts, divide_number, maximum number of iterations T, and seed value of the randomly generated point after coarse clustering in the thread pool are input into the synchronous and concurrent KmeansPlus algorithm for fine clustering. The specific method of the Kmeanplus algorithm is:
[0040] (1) Randomly select a unit from the register / buffer to be clustered as the center point, and calculate the Manhattan distance between the remaining devices and the center point, see formula (3):
[0041] D 1,l =|X 1 -X l |+|Y 1 -Y l | (3)
[0042] Where D 1,l It is the sum of the distance differences between the center point 1 and other unselected cells l along the X-coordinate and Y-coordinate directions.
[0043] (2) Put the Manhattan distance between each register and the center point into the set Ω:
[0044] Ω={D 1,2 ,D 1,3 ,...,D 1,n} (4)
[0045] Where n represents the total number of units to be clustered, D 1,2 is the Manhattan distance between the central unit 1 and other units 2.
[0046] (3) Find the larger Manhattan distance. Formula (5) uses all the calculated Manhattan distances divided by the total number of distances to get the probability distribution. The farther the distance, the easier it is to be selected as the new center point. p A new center point is selected according to probability.
[0047]
[0048] D p ={D′ 1,2 ,D′ 1,3 ,...,D′ 1,n} (6)
[0049] (4) Repeat operations (1)-(3) until k center points are found, where k is the number of groups to be clustered.
[0050] (5) After finding k center points, traverse all remaining units and use formula (7) to calculate the distance between their two points, and summarize the units with closer distances into the group of corresponding center points:
[0051]
[0052] Among them, m represents the unit selected as the center point, and l represents the remaining units.
[0053] The results of coarse clustering enter the multi-thread through the fine clustering task. For the example of 100k registers given above, divide_number is calculated for each group of coarse clustering results. Each of the 8 threads uses the KmeansPlus algorithm to process two groups of registers / buffers. The clustering convergence point information and coordinates are placed in a structure container. After each thread completes two groups of clustering, it is destroyed to release resources. The process is repeated until all grouped registers / buffers are clustered.
[0054] S5. In the hierarchical clock tree synthesis, after the fine clustering of registers / buffers at a certain level is completed, the convergence point needs to be instantiated, and the information and coordinates of each convergence point in the structure container are read to determine whether there is overlap after the instantiation of the convergence point and whether the line length does not exceed the maximum line length. Then, each convergence point is instantiated as a buffer, and the child of the buffer is set to the leaf node connected to the convergence point. All buffers at this level are used as input for the coarse clustering of the next level until the number of buffers at a certain level is 1, indicating that the construction of the hierarchical clock tree synthesis has reached the clock source at this time, and the clock source is connected to the last level of buffer, that is, the child of the clock source is set to the buffer. Finally, the DFS algorithm is used to propagate forward from the clock source to each leaf node, and the delay Latency and global deviation Skew of each path are calculated to determine the quality of clock tree synthesis. The delay and deviation formulas are as follows:
[0055]
[0056] N is the number of registers and buffers in each path in the clock network. A path contains N-1 buffers and 1 register. delay is the buffer delay value, net_delay k,k+1 is the line length delay between unit k and unit k+1.
[0057]
[0058] Skew is the global deviation, which is the difference between the latency of the path with the longest latency and the latency of the path with the shortest latency.
[0059] The above specific implementation methods and examples are specific supports for the technical ideas proposed in the present invention, and cannot be used to limit the protection scope of the present invention. Any equivalent changes or equivalent modifications made on the basis of the technical scheme of the present invention in accordance with the technical ideas proposed in the present invention still fall within the protection scope of the claims of the present invention.
Claims
1. A clock tree synthesis method based on synchronous concurrency and hierarchical clustering, characterized in that: The steps include: Step S1, reading the physical information of all registers and clock sources according to the layout file, and obtaining the number of coarse clustering target groups and the threshold of the number of single group registers according to the number of registers and the number of available threads; Step S2, according to the single group register quantity threshold, use the KmeansPlus algorithm to perform coarse clustering grouping on the read registers, and determine whether the number of registers in each group is less than the single group register quantity threshold. If it is less than the threshold, return the result for fine clustering, otherwise re-perform coarse clustering until the grouping result is less than the single group register quantity threshold; Step S3, extracting the grouping result in step S2, combining the number of registers in the group and the maximum fan-out max_fanout of the buffer Buffer used by the instantiated convergence point to obtain the number of fine clustering groups divide_number, and applying the synchronous and concurrent KmeansPlus algorithm to all register groups for fine clustering; Step S4, extracting the coordinates of the fine clustering convergence point in step S3, combining the size of the buffer used to instantiate the convergence point, determining whether there is an overlap with the connected sub-register, and determining the maximum line length between the convergence point coordinates and the connected register. If the above requirements are met, then use the buffer Buffer to instantiate the group convergence point after each fine clustering, and all buffers Buffer at this level are used as inputs for the next level of clustering; Step S5, extract the buffer instantiated in step S4 as the first-level comprehensive result, and input all the instantiated buffers of this level as the leaf nodes of the second level into the KmeansPlus algorithm for second-level coarse clustering, and iterate steps S1 to S5 to perform bottom-up hierarchical clustering from registers to clock sources, connecting leaf nodes and instantiated aggregation nodes step by step until the last level nodes converge to the clock source.
2. The clock tree synthesis method based on synchronous concurrency and hierarchical clustering according to claim 1, characterized in that: The specific method of step S1 to obtain the target group number and the single group register threshold according to the number of registers and the number of available threads is: calculate the target group number according to the number of available threads and the number of groups processed by each thread, and then calculate the single group register threshold according to the total number of registers and the target group number.
3. The clock tree synthesis method based on synchronous concurrency and hierarchical clustering according to claim 1, characterized in that: In step S2, register insts, a single group of register quantity threshold n, the number of iterations T, and a random seed value seed are input into the KmeansPlus method for coarse clustering.
4. The clock tree synthesis method based on synchronous concurrency and hierarchical clustering according to claim 3, characterized in that: The KmeansPlus algorithm used in the coarse clustering in step S2 is as follows: establish a random generator gen; randomly select an inst from a large number of insts as the center point; calculate the Manhattan distance of the remaining insts from the center point, find the inst with the farthest Manhattan distance from the center point as the new center point, until n center points are found; traverse the remaining insts and group them into the nearest center point as a group; return the constructed group.
5. The clock tree synthesis method based on synchronous concurrency and hierarchical clustering according to claim 1, characterized in that: The step S3 uses the coarse clustering grouping result of step S2, and determines the number of fine clustering groups divide_number according to the number of registers in each group, the maximum fan-out max_fanout of the buffer used Buffer, and the scaling ratio ratio, and inputs the single group of registers insts, the number of groups divide_number, the number of iterations T, and the random seed value seed into the synchronous concurrent KmeansPlus method for fine clustering. The registers are added to the thread pool through the fine clustering task for multi-threaded processing, and the clustering return value is set as shared data, and the shared data is locked to protect the data.
6. The clock tree synthesis method based on synchronous concurrency and hierarchical clustering according to claim 1, characterized in that: The step S4 extracts the fine clustering convergence point of step S3 as the position of the instantiated buffer Buffer, and before instantiation, combines the coordinates of the point and the buffer size to determine whether there is overlap with the connected register. If overlap occurs, it is avoided by fine-tuning the position of the convergence point; and calculates the line length between the convergence point and the coordinates of the connected register, and determines whether each path exceeds the maximum line length constraint, so as to avoid a larger global skew after synthesis.
7. The clock tree synthesis method based on synchronous concurrency and hierarchical clustering according to claim 1, characterized in that: The step S4 extracts the convergence points obtained by each group of KmeansPlus method fine clustering in step S3, and instantiates the convergence points as buffers Buffer as the clustering results of the clock tree at this level, and the buffers Buffer obtained from all fine clustering results at this level are used as leaf nodes of the next level of the clock tree.
8. The clock tree synthesis method based on synchronous concurrency and hierarchical clustering according to claim 1, characterized in that: The step S5 extracts the instantiated buffer of step S4 as the first-level result of the hierarchical clock tree synthesis. All the instantiated buffers at this level are input as leaf nodes of the second-level clock tree. Starting from the second-level leaf nodes to the previous level of the clock source, all convergence points are buffers Buffer, indicating that the input of coarse clustering and fine clustering starting from the second level is buffer Buffer. At this time, each level of information can be replaced with buffer information accordingly.
9. The clock tree synthesis method based on synchronous concurrency and hierarchical clustering according to claim 1, characterized in that: The step S5 performs coarse clustering and fine clustering on each level of leaf nodes to obtain a sink node through hierarchical clock tree synthesis, until the number of sink nodes is 1, indicating that the buffer after instantiation of the sink node is a subunit of the clock source, and the clock tree synthesis is completed by connecting the clock source.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program executes the steps of the clock tree synthesis method according to any one of claims 1 to 9 when running.
Citation Information
Patent Citations
Realization method for register clustering in clock tree synthesis
CN105930591A
Method for realizing robust clock tree comprehensive algorithm for near threshold
CN112257378A
Clock tree establishment method considering low-voltage clock skew fluctuation optimization
CN117272878A
Method and apparatus for generating a variation-tolerant clock-tree for an integrated circuit chip
US20080168412A1