A Clustering-Based Clock Tree Synthesis Method and System
The clustering-based clock tree synthesis method addresses the challenges of clock skew and buffer insertion in large-scale ICs by optimizing clock tree construction with K-means and MiniBatch K-means, resulting in improved clock tree quality and design compliance.
Patent Information
- Application Number
- CN202510587367.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-08
AI Technical Summary
Traditional clock tree synthesis algorithms are difficult to balance clock skew, clock delay and number of buffers, resulting in high computational complexity, path redundancy and design rules violations.
Using the K-means clustering method and its variants, combining the constraints of clock delay, clock offset and number of buffers, the clock tree is built from the bottom up, the number and position of clusters are dynamically adjusted, and the clock tree is optimized using the DME algorithm to add position overlap constraints and maximum RC constraints to ensure that the clock tree meets the design requirements.
It improves the comprehensive quality of the clock tree, reduces clock delay and offset, reduces the number of buffer usage, reduces circuit complexity and power consumption, improves algorithm efficiency and applicability, and provides a reliable design foundation.
Smart Images

Figure CN120106007B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of integrated circuit design and electronic design automation, and in particular to a clustering-based clock tree synthesis method and system. Background Art
[0002] In recent years, the development of the integrated circuit (IC) field has been increasingly rapid, and the chip design requirements have been continuously increasing. With the progress of integrated circuit manufacturing processes, the number of transistors in a chip has increased year by year, and the number of logic gates included in the chip has increased sharply. A very large scale integration (VLSI) chip integrates tens of millions or even tens of billions of transistors, and thousands of nets need to be connected; at the same time, each net has hundreds or even more wiring schemes, which makes the wiring problem extremely complex. When designing a chip, clock tree synthesis (CTS) is a very important step in the physical design process. It interconnects the modules, standard cells, and input / output units distributed in the chip core according to the logical relationship, and the formed clock network must meet the design rules. The quality of the CTS result directly affects the function, performance, and stability of the entire chip.
[0003] Traditional CTS algorithms often cannot balance the goals of running time and CTS quality, and the final CTS scheme determined has cases of not avoiding short circuits and violating design rules, so it is difficult to solve the CTS problem of multi-instantiated modules. Existing CTS algorithms can be roughly divided into three categories: CTS algorithms based on heuristic algorithms, such as L-shaped pattern routing in the shortest path, although they run relatively fast, due to the greediness of the heuristic algorithm itself, many better CTS strategies will be missed, which may lead to larger clock delays or uneven clock skews; CTS algorithms based on search algorithms, such as the maze algorithm, although they can discretize the problem space, due to ignoring the concurrency of CTS strategies, not only are they prone to repeatedly searching a large number of similar strategies, but also a large number of redundant paths still need to be searched under multiple endpoints, resulting in an excessive number of buffer insertions; CTS algorithms based on graph theory algorithms, such as network flow and Steiner tree algorithms, although they have a seemingly more systematic and sophisticated theoretical system, do not make full use of the geometric information in the CTS process, making the CTS not flexible enough. In summary, existing methods are difficult to achieve a balance among clock skew, clock delay, and the number of buffer insertions. Summary of the Invention
[0004] Objective of the Invention: The objective of the present invention is to provide a clustering-based clock tree synthesis method and system. By adding constraints such as clock delay, clock skew, and the number of buffers to the K-means clustering method and its variants, an optimal clock tree insertion scheme is obtained, solving the problems of high computational complexity, path redundancy, short circuits, and design rule violations in traditional CTS algorithms when dealing with large-scale designs.
[0005] Technical Solution: The clustering-based clock tree synthesis method described in the present invention includes the following steps:
[0006] (1) Read the sizes, quantities, and positions of the registers in the circuit;
[0007] (2) Store the coordinates of the registers in a coordinate set, perform clustering on the coordinate set to generate several clusters to be processed, and clear the coordinate set; adjust each cluster to be processed according to the constraint conditions until the constraint conditions are met, and use the position of the clustering center in each cluster to be processed as the position of the buffer at the first level to be inserted;
[0008] The constraint conditions include the maximum fan-out number limit constraint and the maximum RC constraint. For a cluster to be processed that does not meet the constraint conditions, remove the buffer that is farthest from the clustering center and reassign the cluster to be processed. If it cannot be assigned, store the coordinates of the buffer in the coordinate set; the maximum RC constraint is: calculate , if it is not greater than the maximum RC limit, then the maximum RC constraint is met, is the Manhattan distance between the buffer and the clustering center, and r and c are the resistance value and capacitance value of the connection line between the buffer and the clustering center respectively;
[0009] (3) Store the coordinates of the buffers at the first level in the coordinate set, and repeat step (2) to obtain the positions of the buffers at the second level;
[0010] (4) Use the DME algorithm to build the clock tree for the buffers at the second level until the clock tree synthesis of the circuit is completed, obtaining the positions of the buffers from the third level to the Nth level;
[0011] (5) Output the clock tree synthesis scheme, including the number, position, and connection relationship of the buffers.
[0012] Further, in step (2), a position overlap constraint is also included. If the buffer overlaps with the register position, move the position of the buffer. Specifically, it includes the following steps:
[0013] The register is a first rectangle with size The buffer is a second rectangle with size Expand outward based on the vertex in the preset direction of the first rectangle to form a size of The third rectangle, which contains the first rectangle and coincides with the first vertex angle of the first rectangle, and the first vertex angle is located at the diagonal position in the preset direction of the first rectangle;
[0014] Determine whether the vertex in the preset direction of the second rectangle is inside the third rectangle. If it is inside the third rectangle, the register and the buffer overlap, and move the vertex in the preset direction of the second rectangle to the safe line segment on the third rectangle;
[0015] The method for selecting the safe line segment on the third rectangle is as follows:
[0016] Determine whether there are other registers around this register. The third rectangle formed according to this register is called the target third rectangle, and the third rectangles formed according to other registers around this register are called neighboring third rectangles. If they intersect, whether there are other registers around this register; Remove the part that intersects with the neighboring third rectangle on the boundary of the third rectangle, and remove the part that coincides with the first rectangle to obtain the safe line segment.
[0017] Further, in step (2), if the number of coordinates in the coordinate set is greater than the maximum fan-out number, the MinibatchKmeans clustering algorithm is used for clustering, otherwise the K-means clustering algorithm is used for clustering.
[0018] Further, in the Minibatch Kmeans clustering algorithm, the value of K in the clustering is:
[0019] ;
[0020] where is the total number of coordinates in the coordinate set, is the maximum fan-out number.
[0021] Further, in the K-means clustering algorithm, the elbow method is used to determine the value of K.
[0022] Further, in step (4), determine whether the position of the buffer satisfies the position overlap constraint. After satisfying the position overlap constraint, determine whether the maximum RC constraint is satisfied. If not, continue to insert the buffer and continue to detect the position overlap constraint and the maximum RC constraint until the constraints are satisfied.
[0023] The clustering-based clock tree synthesis system described in the present invention includes:
[0024] A data reading unit for reading the size, quantity, and position of registers in a circuit;
[0025] The first-level buffer insertion unit is used to store the coordinates of the registers into a coordinate set, cluster the coordinate set to generate several clusters to be processed, and clear the coordinate set; adjust according to the constraint conditions in each cluster to be processed until the constraint conditions are met, and use the positions of the cluster centers in each cluster to be processed as the positions of the inserted first-level buffers;
[0026] The constraint conditions include the maximum fan-out number limit constraint and the maximum RC constraint. For the clusters to be processed that do not meet the constraint conditions, the buffer farthest from the cluster center is removed and the clusters to be processed are reallocated. If it cannot be allocated, the coordinates of the buffer are stored in the coordinate set; the maximum RC constraint is: calculate , if it is not greater than the maximum RC limit, then the maximum RC constraint is met, is the Manhattan distance between the buffer and the cluster center, and r and c are the resistance value and capacitance value of the line connecting the buffer and the cluster center respectively;
[0027] The second-level buffer insertion unit is used to store the coordinates of the first-level buffers into the coordinate set, and repeat the first-level buffer insertion unit to obtain the positions of the second-level buffers;
[0028] The subsequent buffer insertion unit is used to build the clock tree for the second-level buffers through the DME algorithm until the clock tree synthesis of the circuit is completed, and obtain the positions of the third-level to the Nth-level buffers;
[0029] The clock tree synthesis scheme output unit is used to output the clock tree synthesis scheme, including the number, positions and connection relationships of the buffers.
[0030] Further, in the first-level buffer insertion unit, if the number of coordinates in the coordinate set is greater than the maximum fan-out number, the Minibatch Kmeans clustering algorithm is used for clustering, otherwise the K-means clustering algorithm is used for clustering;
[0031] In the Minibatch Kmeans clustering algorithm, the value of K in the clustering is:
[0032] ;
[0033] where is the total number of coordinates in the coordinate set, is the maximum fan-out number;
[0034] In the K-means clustering algorithm, the elbow method is used to determine the value of K.
[0035] The electronic device described in the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded into the processor, the clock tree synthesis method based on clustering described above is implemented.
[0036] The computer-readable storage medium described in the present invention stores a computer program. When the computer program is executed by a processor, the clock tree synthesis method based on clustering described above is implemented.
[0037] Beneficial effects: Compared with the prior art, the advantages of the present invention are as follows:
[0038] (1) Improve the quality of clock tree synthesis: By grouping registers through the K-means clustering method and its variants, and constructing the clock tree bottom-up in combination with the DME algorithm, it can effectively calculate the Manhattan distance between adjacent points and extend the rectangular area along the 45-degree direction, determine the overlapping line segments to balance the clock skew at the upper level, thereby reducing the clock delay and clock skew as a whole, and ensuring the stability and synchronization of the clock signal. Dynamically adjust the number of clusters during the clustering process to ensure that the fan-out within each cluster does not exceed the maximum allowable value, and further optimize the distribution of clusters through post-processing methods, making the insertion positions of buffers more reasonable, thereby reducing the number of buffers used and the complexity and power consumption of the circuit;
[0039] (2) Improve the algorithm efficiency and applicability: Use the MiniBatch K-means clustering method to perform preliminary clustering on a large number of registers, significantly reducing the computational cost, especially suitable for scenarios with limited memory or excessive data volume, and improving the applicability of the algorithm in large-scale integrated circuit design. Use the MiniBatch K-means clustering with faster clustering speed in the initial clustering stage to ensure the clustering accuracy; while switch to the K-means clustering combined with the elbow method in the subsequent stage, without fixing the value of K, use the elbow method to obtain the most suitable value of K, further optimize the clustering result, balance the clustering time and clustering accuracy, and improve the efficiency and effect of the algorithm;
[0040] (3) Optimize the design constraint conditions: Add constraints on conditions such as clock delay, clock skew, and buffer number in each step of the clustering process, and perform judgments and optimizations on the clustering results through post-processing methods such as maximum fan-out constraint, maximum RC constraint, and position overlap constraint to ensure that the finally generated clock tree meets all design constraint conditions and avoid problems such as short circuits and violations of design rules. By dynamically adjusting the number and position of clusters, and performing constraint and position overlap judgments when building each level of the clock tree, the construction of the clock tree is made more flexible and can better adapt to different design requirements and scenarios;
[0041] (4) Lay the foundation for subsequent design: The optimized design information output includes register location information, the location information of inserted buffers, and the net connection relationship part, which is stored in the form of a text file, providing a reliable basis for the verification and adjustment of subsequent digital circuit designs, and helping to improve the efficiency and quality of the entire chip design. The clock tree scheme generated by this method not only performs excellently in the CTS stage, but also provides a solid foundation for subsequent physical designs, promoting the connection and overall optimization of the design process, and having broad application prospects and practical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a flowchart of the clock tree synthesis method of the present invention.
[0043] Figure 2 It is a flowchart for generating buffer locations through Kmeans clustering.
[0044] Figure 3 It is a schematic diagram of position overlap in an embodiment of the present invention.
[0045] Figure 4 It is a schematic diagram of the principle for determining position overlap in an embodiment of the present invention.
[0046] Figure 5 It is a schematic diagram of the buffer insertion result at the first level in an embodiment of the present invention.
[0047] Figure 6 It is a schematic diagram of the buffer inserted through DME in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0049] As Figure 1 shown, the clustering-based clock tree synthesis method includes the following steps.
[0050] S1: Write a program for reading input files in C++ language.
[0051] Use a C++ program to read the input and constraint files to obtain the register connection information and signal topology of the circuit. The content obtained from the input file includes the unit, floorplan, coordinates of the clock source CLK, sizes of the flip-flops (FF) and buffers (buffer) FF_cell and buffer_cell, and coordinates of all flip-flops FF_instance; the content obtained from the constraint file includes the resistance value r, capacitance value c, maximum rc value limit, maximum fan-out limit Max_fanout, and delay data of a single buffer, etc. Finally, visualize the extracted information to verify whether the circuit information is accurately extracted.
[0052] S2: Perform Kmeans clustering on the FF coordinates and insert the first-level buffers.
[0053] As Figure 2 shown, to perform step S2, it is first necessary to construct three key data sets, namely the coordinate set, the set of clusters to be processed, and the clustering result set. Enter the coordinate information of all flip-flops in the circuit into the coordinate set completely. Subsequently, judge the scale of the coordinate set. If the number of coordinates it contains exceeds the preset maximum fan-out limit, immediately start the Minibatch Kmeans clustering algorithm to perform clustering tasks on all coordinates in the coordinate set. After clustering, store the obtained clustering results (i.e., multiple clusters) into the set of clusters to be processed uniformly. At the same time, clear all coordinate data in the coordinate set for subsequent operations.
[0054] Next, carefully examine each cluster in the set of clusters to be processed. Focus on checking the number of coordinates in each cluster. For those clusters whose number of coordinates within the cluster still exceeds the maximum fan-out limit, use a post-processing method based on a priority queue for optimization. Specifically, according to the Manhattan distance, sequentially select the flip-flop farthest from the clustering center from each over-limit cluster and try to find a new cluster that can accommodate it. If a suitable new cluster is successfully found, move the flip-flop to the new cluster; if no suitable new cluster can be found after traversing all possibilities, put the flip-flop back into the coordinate set. During the process of moving flip-flops out or in, update the position of the clustering center of the relevant cluster in real time to ensure that the number of coordinates in each cluster does not exceed the constraint limit of the maximum fan-out.
[0055] After the above post-processing is completed, enter the maximum RC constraint determination phase. For each cluster in the cluster set to be processed, calculate its RC value and compare it with the preset maximum RC value. If the RC value of the cluster is less than or equal to the maximum RC value, it indicates that the cluster meets the design requirements, and it is removed from the cluster set to be processed and stored in the clustering result set; otherwise, if the RC value of the cluster exceeds the maximum RC value, the point with the farthest Manhattan distance from the clustering center needs to be removed from the cluster, put it back into the coordinate set, and then update the clustering center of the cluster. Calculate and determine again whether the RC value of the cluster meets the constraints.
[0056] After the RC value meets the conditions, it is also necessary to perform position overlap determination and correction on the clustering center of the cluster. Figure 3 A buffer with overlapping positions is selected for effect display. The red color represents the initial position of the overlapping buffer. After position overlap determination and correction, the insertion position of the buffer is transferred to the non-overlapping green position. The position overlap determination and correction are as follows Figure 4As shown, during the overlapping determination process, for all rectangles determined by the sizes of registers and buffers, the position overlapping determination can be performed by comparing the relative positions of a certain vertex of the rectangle. In the following discussion, the lower left vertex is used. The shape of the register is like rectangle A1B1C1D1, its length is denoted as a, and its width is denoted as c. The shape of the buffer is like rectangle EFGH, its length is denoted as b, and its width is denoted as c. That is, the registers and buffers have the same width. First, determine whether there is an overlap. Rectangle A1B1C1D1 already exists. Starting from the position of point B1, move b units to the left horizontally and then c units up vertically to obtain vertex I1. Starting from the position of point B1, move b units to the left horizontally and then c units down vertically to obtain vertex J1. Construct rectangle I1J1K1D1. At this time, just determine whether the position of the lower left vertex of the buffer is inside rectangle I1J1K1D1. If it is inside, it means there is an overlap and movement is required. To minimize the movement distance to reduce the deviation, try to select the position of the safe line segment on the side of rectangle I1J1K1D1 that will not cause an overlap. At this time, it is necessary to find other possible registers near rectangle A1B1C1D1. Starting from the position of point I1, move a units to the left horizontally and then c units up vertically to obtain vertex L. Starting from the position of point J1, move a units to the left horizontally and then c units down vertically to obtain vertex M. Starting from the position of point D1, move b units to the right horizontally and then c units up vertically to obtain vertex O. Starting from the position of point K1, move b units to the right horizontally and then c units down vertically to obtain vertex N. First, find all registers whose lower left vertices are inside rectangle LMNO, such as A2B2C2D2. Then, in the same way, construct a rectangle similar to rectangle I1J1K1D1, such as I2J2K2D2. Then, the intersecting line segment between rectangle I2J2K2D2 and rectangle I1J1K1D1 is the non-placeable position. By judging all the registers inside rectangle LMNO, the safe line segment on rectangle I1J1K1D1 can be obtained. When the lower left vertex of the buffer is moved to the safe line segment, the non-overlap condition is satisfied. To minimize the deviation as much as possible, select a point on the safe line segment with the smallest difference in distance from the original position after movement for placement. If none of the rectangles I1J1K1D1 satisfy, select another nearest register and perform the same operation until a non-overlapping safe position is found.
[0057] Since the Minibatch Kmeans clustering algorithm has significant advantages in processing large-scale point sets, in the later stage of the initial clustering, when the number of remaining registers in the coordinate set is small, in order to further improve the clustering accuracy, it is necessary to switch the clustering method. The present invention adopts K-means clustering combined with the elbow method as the subsequent clustering algorithm. The criterion for determining the switch of the clustering method is: when the number of remaining coordinates in the coordinate set is greater than the maximum fan-out limit, continue to use the Minibatch Kmeans clustering; when the number of coordinates is less than or equal to the maximum fan-out limit, switch to the clustering method of K-means combined with the elbow method. After the clustering method is switched, except for the difference in the clustering algorithm itself, the remaining post-processing processes (such as the optimization operation based on the priority queue) and the position overlap determination operation are the same as before, ensuring the coherence and stability of the entire clustering process. The insertion result of the first-level buffer of this example is as Figure 5 shown.
[0058] S3: Perform Kmeans clustering on the coordinates of the first-level buffer and insert them into the second-level buffer.
[0059] This operation is basically the same as S2, except that the points initially stored in the coordinate set are changed to the coordinates of the first-level buffer.
[0060] S4: Build the subsequent clock tree for the obtained second-level buffer through the DME algorithm.
[0061] After obtaining the coordinates of the second-level buffer, the delay deviation between different registers in the direct clustering to clk may be very large. The DME algorithm balances the two buffers. First, use the greedy algorithm to classify the coordinates of the second-level buffer, then use the DME algorithm to obtain the possible placement positions, and then perform overlap determination and movement on the possible positions. On the premise of meeting the non-overlap condition, minimize the error of the delay offset as much as possible. Detect the maximum RC constraint. If it is not satisfied, insert an additional buffer. The inserted buffer will divide the original length d into two segments e and f, where e + f = d. The distances e and f in the two lines d1 and d2 will affect the total delay. The selected position hopes that e1 2 + f1 2 The delay brought plus its corresponding upper-level delay and e2 2 + f2 2 The sum of the delay brought plus its corresponding upper-level delay has as small a difference as possible. Then detect whether there is overlap, and repeat the overlap detection and RC constraint determination steps until all conditions are met. As Figure 6 shown, where the blue rectangle is the second-level buffer and the green rectangle is the position of the next-level buffer.
[0062] S5: Output the obtained optimal clock tree scheme.
[0063] After completing the optimization process, the system will output the optimized circuit clock tree synthesis information to a file. The output content covers key parts such as register position information, the positions of inserted buffers, and the optimized signal topology and net connection relationships. These output information have good readability and standardization, can provide detailed data support for subsequent circuit verification work, and are also convenient for designers to carry out iterative design to further optimize circuit performance.
[0064] The clustering-based clock tree synthesis system described in the present invention includes:
[0065] A data reading unit for reading the sizes, quantities, and positions of registers in the circuit;
[0066] A first-level buffer insertion unit for storing the coordinates of the registers into a coordinate set, clustering the coordinate set to generate several clusters to be processed, clearing the coordinate set; adjusting according to the constraint conditions in each cluster to be processed until the constraint conditions are met, and taking the position of the clustering center in each cluster to be processed as the position of the first-level buffer;
[0067] The constraint conditions include the maximum fan-out number limit constraint and the maximum RC constraint. For a cluster to be processed that does not meet the constraint conditions, the buffer farthest from the clustering center is removed and the cluster to be processed is reallocated. If it cannot be allocated, the coordinates of the buffer are stored in the coordinate set; the maximum RC constraint is: calculate , if it is not greater than the maximum RC limit, then the maximum RC constraint is satisfied, is the Manhattan distance between the buffer and the clustering center, and r and c are the resistance value and capacitance value of the connection line between the buffer and the clustering center respectively;
[0068] A second-level buffer insertion unit for storing the coordinates of the first-level buffer into the coordinate set and repeating the first-level buffer insertion unit to obtain the position of the second-level buffer;
[0069] A subsequent buffer insertion unit for building a clock tree for the second-level buffer through the DME algorithm until the clock tree synthesis of the circuit is completed to obtain the positions of the third-level to the Nth-level buffers;
[0070] A clock tree synthesis scheme output unit for outputting a clock tree synthesis scheme, including the quantity, position, and connection relationship of the buffers.
[0071] The electronic device described in the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, the clustering-based clock tree synthesis method described above is implemented.
[0072] The computer-readable storage medium according to the present invention stores a computer program, and when the computer program is executed by a processor, it implements the clock tree synthesis method based on clustering as described above.
[0073] The computer-readable storage medium may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other medium that can store program code in the form of instructions or data structures and can be accessed by a computer.
[0074] The processor is used to execute the computer program stored in the memory to implement each step in the methods involved in the above embodiments.
Claims
1. A clustering-based clock tree synthesis method, characterized in that, It includes the following steps: (1) Read the size, quantity, and location of the registers in the circuit; (2) Store the coordinates of the registers into a coordinate set, perform clustering on the coordinate set to generate several clusters to be processed, and clear the coordinate set; Adjust according to the constraint conditions in each cluster to be processed until the constraint conditions are met, and use the location of the cluster center in each cluster to be processed as the location of the inserted first-level buffer; The constraint conditions include the maximum fan-out number limit constraint and the maximum RC constraint. For the cluster to be processed that does not meet the constraint conditions, the buffer that is farthest from the clustering center is removed and the cluster to be processed is reallocated. If it cannot be allocated, the coordinates of the buffer are stored in the coordinate set; the maximum RC constraint is: calculate , if it is not greater than the maximum RC limit, then the maximum RC constraint is satisfied, is the Manhattan distance between the buffer and the clustering center, and r and c are the resistance value and capacitance value of the line connecting the buffer and the clustering center respectively; (3) Store the coordinates of the first-level buffer into the coordinate set, and repeat step (2) to obtain the location of the second-level buffer; (4) Build the clock tree for the second-level buffer through the DME algorithm until the clock tree synthesis of the circuit is completed, and obtain the locations of the third-level to Nth-level buffers; (5) Output the clock tree synthesis scheme, including the quantity, location, and connection relationship of the buffers; In step (2), a location overlap constraint is also included. If the buffer overlaps with the register in location, move the location of the buffer. Specifically, it includes the following steps: The register is a first rectangle with a size of , and the buffer is a second rectangle with a size of . Based on the vertex in the preset direction of the first rectangle, it expands outward to form a third rectangle with a size of . The third rectangle contains the first rectangle and coincides with the first top angle of the first rectangle, and the first top angle is located at the diagonal position in the preset direction of the first rectangle; Determine whether the vertex in the preset direction of the second rectangle is inside the third rectangle. If it is inside the third rectangle, the register and the buffer overlap, and move the vertex in the preset direction of the second rectangle to the safe line segment on the third rectangle; The selection method of the safe line segment on the third rectangle is as follows: Determine whether there are other registers around this register. The third rectangle formed according to this register is called the target third rectangle, and the third rectangles formed according to other registers around this register are called neighboring third rectangles. If they intersect, there are other registers around this register; Remove the part that intersects with the neighboring third rectangle and the part that coincides with the first rectangle on the boundary of the third rectangle to obtain the safe line segment.
2. The clustering-based clock tree synthesis method according to claim 1, wherein In step (2), if the number of coordinates in the coordinate set is greater than the maximum fan-out number, use the Minibatch Kmeans clustering algorithm for clustering, otherwise use the K-means clustering algorithm for clustering.
3. The clustering-based clock tree synthesis method according to claim 2, wherein In the MinibatchKmeans clustering algorithm, the value of K in the clustering is: ; wherein is the total number of coordinates in the coordinate set, is the maximum fan-out number.
4. The clustering-based clock tree synthesis method according to claim 2, wherein Use the elbow method to determine the value of K in the K-means clustering algorithm.
5. The clustering-based clock tree synthesis method according to claim 1, wherein In step (4), determine whether the location of the buffer meets the location overlap constraint. After meeting the location overlap constraint, determine whether the maximum RC constraint is met. If not, continue to insert buffers and continue to detect the location overlap constraint and the maximum RC constraint until the constraints are met.
6. A clustering-based clock tree synthesis system, characterized in that It includes: A data reading unit for reading the size, quantity, and location of the registers in the circuit; A first-level buffer insertion unit for storing the coordinates of the registers into a coordinate set, performing clustering on the coordinate set to generate several clusters to be processed, and clearing the coordinate set; Adjust according to the constraint conditions in each cluster to be processed until the constraint conditions are met, and use the location of the cluster center in each cluster to be processed as the location of the inserted first-level buffer; The constraint conditions include the maximum fan-out number limit constraint and the maximum RC constraint. For the cluster to be processed that does not meet the constraint conditions, the buffer that is farthest from the clustering center among them is removed and the cluster to be processed is reallocated. If it cannot be allocated, the coordinates of the buffer are stored in the coordinate set; the maximum RC constraint is: calculate , if it is not greater than the maximum RC limit, the maximum RC constraint is satisfied, is the Manhattan distance between the buffer and the clustering center, and r and c are the resistance value and capacitance value of the line connecting the buffer and the clustering center respectively; A second-level buffer insertion unit for storing the coordinates of the first-level buffer into the coordinate set and repeating the first-level buffer insertion unit to obtain the location of the second-level buffer; A subsequent buffer insertion unit is used to construct a clock tree for the second-level buffer by using a DME algorithm until the clock tree synthesis of the circuit is completed to obtain the positions of the third-level to N-th-level buffers; A clock tree synthesis solution output unit, used to output a clock tree synthesis solution, including the number and location of buffers and their connection relationship; The first-level buffer insertion unit also includes a position overlap constraint. If the buffer and the register position overlap, the position of the buffer is moved, which specifically includes the following steps: The register is a first rectangle with a size of The buffer is a second rectangle with a size of Based on the vertex in the preset direction of the first rectangle, it expands outward to form a third rectangle with a size of The third rectangle contains the first rectangle and coincides with the first vertex angle of the first rectangle. The first vertex angle is located at the diagonal position in the preset direction of the first rectangle; Determine whether the vertex of the second rectangle in the preset direction is inside the third rectangle. If it is inside the third rectangle, the register and the buffer overlap, and move the vertex of the second rectangle in the preset direction to the safety line segment on the third rectangle; The method for selecting the safety line segment on the third rectangle is: Determine whether there are other registers around the register. The third rectangle formed by the register is called the target third rectangle, and the third rectangle formed by other registers around the register is called the neighboring third rectangle. If they intersect, there are other registers around the register. On the boundary of the third rectangle, remove the part that intersects with the neighboring third rectangle, and remove the part that overlaps with the first rectangle to obtain a safety line segment.
7. The clustering-based clock tree synthesis system according to claim 6, wherein In the first-level buffer insertion unit, if the number of coordinates in the coordinate set is greater than the maximum fan-out number, the Minibatch Kmeans clustering algorithm is used for clustering, otherwise the K-means clustering algorithm is used for clustering; In the Minibatch Kmeans clustering algorithm, the K value in the cluster is: ; wherein is the total number of coordinates in the coordinate set, is the maximum fan-out number; The elbow method is used to determine the K value in the K-means clustering algorithm.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is loaded into a processor, the clustering-based clock tree synthesis method according to any one of claims 1 to 5 is implemented.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the clustering-based clock tree synthesis method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Clock tree processing method and system, electronic equipment and medium
CN118350340A
Machine-learning based prediction method for iterative clustering during clock tree synthesis
US11244099B1