Clock Tree Synthesis Optimization Method and System Based on Hierarchical Region Partitioning and Buffer Insertion

By adopting the methods of layered area division and buffer insertion in the clock tree comprehensive optimization, the problems of clock skew, delay and buffer insertion redundancy in complex circuits are solved, and the stability and efficient allocation of clock signals are achieved, which improves the reliability and efficiency of the design.

CN119886040BActive Publication Date: 2025-07-01NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510373332.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-01
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

The existing comprehensive clock tree optimization technology is difficult to efficiently deal with the problems of clock skew, clock delay, buffer insertion redundancy and low optimization efficiency in complex circuits.

Method used

The clock tree comprehensive optimization method based on hierarchical area division and buffer insertion is adopted. The bottom and middle buffers are inserted through rectangular cutting and diamond mesh division, and the top buffer is inserted according to the H-tree structure. Combining the clustering algorithm and dynamic clock optimization strategy, the allocation of clock signals and buffer position are optimized.

Benefits of technology

Quickly identify and eliminate clock delays and clock offsets, reduce buffer count, optimize clock signal allocation, reduce design power consumption, improve clock signal quality, and improve circuit simulation and verification efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119886040B_ABST
    Figure CN119886040B_ABST
Patent Text Reader

Abstract

The present invention discloses a clock tree synthesis optimization method and system based on hierarchical region division and buffer insertion. This method performs clock tree synthesis optimization by inserting buffers hierarchically. First, all registers are used as basic units to cut the circuit into rectangular regions and insert bottom-layer buffers at the geometric center points. Then, the upper-layer buffers are used as basic units, and the circuit region is divided by using a diamond grid division method and middle-layer buffers are inserted at the geometric center points. Finally, top-layer buffers are inserted according to the H-tree structure. After inserting the buffers, the position of the buffers is dynamically adjusted to minimize the delay and skew control of the global clock tree, effectively reduce the number of buffers, lower the power consumption, and improve the design efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to digital circuit simulation and optimization technologies, and particularly to a clock tree synthesis optimization method and system based on hierarchical region partitioning and buffer insertion. Background Art

[0002] In digital integrated circuit design, the clock signal serves as the benchmark for data transmission, and its stability and reliability are crucial for the functionality and performance of the entire system. The clock signal is usually the signal with the largest fan-out, the longest transmission distance, and the fastest operating speed in the entire chip. Therefore, the generation and distribution process of the clock signal has a decisive impact on the design quality of the chip. If the clock signal fails to meet the critical timing requirements under the worst conditions, it may cause data errors to be incorrectly latched into registers, resulting in the failure and instability of the system functionality.

[0003] In modern chip design, clock tree synthesis (CTS) is a key step in the physical design phase, especially in the ASIC (Application Specific Integrated Circuit) process based on a standard cell library. The design goal of the clock tree is to ensure that when the clock signal reaches all clock receivers (such as registers, memories, etc.) from the clock source, it meets the timing constraints and avoids problems such as clock skew and delay. The design process of the clock tree usually includes the selection of the root node, the distribution of the clock signal, and the insertion of clock drivers.

[0004] The clock signal starts from the source point (root node), passes through multiple distribution nodes, and finally reaches registers or other clock terminals, forming a complete clock tree. The root node of the clock tree is usually located at the clock source position of the chip, and the distribution nodes and leaf nodes are respectively associated with the transmission process of the clock signal. To ensure the effective transmission of the clock signal, it is often necessary to insert buffers between the distribution nodes and the leaf nodes to reduce the attenuation and delay during the transmission of the clock signal.

[0005] Existing clock tree synthesis optimization techniques include hierarchical clock tree synthesis methods, which divide the design of the clock tree into multiple levels and distribute clock signals in a hierarchical manner. Although it can effectively reduce the computational complexity, when optimizing clock skew and clock delay, the hierarchical method may lead to poor global clock synchronization and large local clock skew. Especially when the chip design complexity is high, insufficient optimization at the low level may cause uneven distribution of clock signals, resulting in large clock skew. In addition, traditional buffer insertion methods are based on local optimal strategies, such as greedy algorithms. Under certain conditions, they may insert too many buffers to optimize clock skew or delay, but fail to consider the need for global optimization, resulting in unnecessary buffer insertion. The insertion of too many buffers not only increases the power consumption of the chip but also occupies more layout space, which may lead to area waste. Especially in designs with limited resources, excessive buffer insertion will reduce the design efficiency. Summary of the Invention

[0006] Object of the Invention: The object of the present invention is to provide a clock tree synthesis optimization method and system based on hierarchical region division and buffer insertion, which can identify and optimize timing problems and system function failures and instabilities that may be caused by clock signal distribution in the circuit, and thus achieve stable design of the circuit, and solve the technical problems that traditional clock tree synthesis methods cannot efficiently handle clock skew, clock delay, redundant buffer insertion, and low optimization efficiency in complex circuits.

[0007] Technical Solution: The clock tree synthesis optimization method based on hierarchical region division and buffer insertion described in the present invention includes the following steps:

[0008] Step 1: Extract circuit information, including register connection information and signal topological structure;

[0009] Step 2: Use all registers as basic units, divide the circuit area by the rectangular cutting method, and insert bottom-layer buffers at the geometric center points of each first region obtained after division;

[0010] Step 3: Use the upper-layer buffers as basic units, divide the circuit area by the rhombus grid division method, and insert middle-layer buffers at the geometric center points of each second region obtained after division; Repeat Step 3 until n layers of middle-layer buffers are inserted; n is the number of layers of middle-layer buffers;

[0011] Wherein the upper-layer buffer of the nth layer of middle-layer buffers is the (n - 1)th layer of middle-layer buffers, and the upper-layer buffer of the first layer of middle-layer buffers is the bottom-layer buffer;

[0012] Step 4: Insert m layers of top-layer buffers according to the H-tree structure; m is the number of layers of top-layer buffers;

[0013] Step 5: Output the circuit information after inserting buffers, including the position information of the bottom buffers, middle buffers, and top buffers, the signal topology after inserting buffers, and the connection relationship between registers and buffers.

[0014] Further, in Step 2, use the hierarchical clustering algorithm to divide each first region into several point sets according to the maximum fan-out limit. The points in the point sets are registers; calculate the geometric center of each point set and insert a bottom buffer.

[0015] Further, in Steps 2 and 3, perform overlap detection on each bottom buffer and middle buffer. If there is an overlap, adjust the position of the geometric center and then insert the bottom buffer and middle buffer, including the following steps:

[0016] Detect whether the edges of the bottom buffer and the middle buffer intersect with other buffers and / or registers. If they intersect, move the position of the geometric center in the opposite direction of the intersecting edge, and then perform overlap detection again until any edge of the bottom buffer and the middle buffer does not intersect with other buffers and / or registers.

[0017] Further, in Step 3, taking the upper-layer buffer as the basic unit, divide the circuit area using the diamond grid division method, and insert a middle buffer at the geometric center point of each obtained second region, including:

[0018] Calculate the ratio of the number of upper-layer buffers to the maximum fan-out. Round up this ratio to obtain the middle-layer division number. Perform diamond division according to the middle-layer division number to obtain several second regions. Calculate whether the number of basic units in each second region meets the maximum fan-out limit. If it does not meet, re-divide the basic unit farthest from the geometric center of the second region until each second region meets the maximum fan-out limit; calculate the sum of the net_rc values between the basic units in each second region and the geometric center , if exceeds the maximum net_rc value limit within this second region, re-divide the basic unit farthest from the geometric center of the second region until each second region meets the maximum net_rc value limit; insert a middle buffer at the geometric center point of each second region;

[0019] where the net_rc value is , where D is the Manhattan distance between the basic unit and the geometric center, and r and c are the resistance value and capacitance value of the line connecting the basic unit and the geometric center, respectively.

[0020] Further, in Step 3, the number range of the nth-layer middle buffer is [30, 40].

[0021] Further, in step 4, for any two top-level buffers q1 and q2 connected to any one top-level buffer p at the k-th layer, calculate the net_rc values from p to q1 and q2, denoted as net_rc1 and net_rc2, and calculate the maximum net_rc values from q1 and q2 to the top-level buffers at the (k - 2)-th layer they are connected to, denoted as c1 and c2; where k = 3, 4, 5, …, m;

[0022] If net_rc1 + c1 > net_rc2 + c2, then adjust the position of p towards the direction of q1. If net_rc1 + c1 < net_rc2 + c2, then adjust the position of p towards the direction of q2. If net_rc1 + c1 = net_rc2 + c2, then do not adjust the position of p;

[0023] where the net_rc value is , where D is the Manhattan distance between the buffers, and r and c are the resistance value and capacitance value of the wire connection between the buffers respectively.

[0024] Further, in step 4, for any one top-level buffer p at the k-th layer, any one top-level buffer q1 connected to p, and any one top-level buffer q3 connected to q1 at the (k - 2)-th layer, where k = 3, 4, 5, …, m; if the abscissa of q3 is between the abscissas of p and q1, then adjust the position of q1 towards the direction of p along the abscissa. If the ordinate of q3 is between the ordinates of p and q1, then adjust the position of q1 towards the direction of p along the ordinate.

[0025] The clock tree synthesis optimization system based on hierarchical region division and buffer insertion according to the present invention includes:

[0026] A circuit information extraction unit, configured to extract circuit information, including register connection information and signal topology structure;

[0027] A bottom-layer buffer insertion unit, configured to use all registers as basic units, divide the circuit region by using the rectangular cutting method, and insert bottom-layer buffers at the geometric center points of each first region obtained after division;

[0028] A middle-layer buffer insertion unit, configured to use the upper-layer buffers as basic units, divide the circuit region by using the rhombic grid division method, and insert middle-layer buffers at the geometric center points of each second region obtained after division; repeat step 3 until n layers of middle-layer buffers are inserted; n is the number of layers of middle-layer buffers;

[0029] where the upper-layer buffer of the n-th layer of middle-layer buffers is the (n - 1)-th layer of middle-layer buffers, and the upper-layer buffer of the first layer of middle-layer buffers is the bottom-layer buffer;

[0030] A top - level buffer insertion unit for inserting m - layer top - level buffers according to the H - tree structure; m is the number of layers of the top - level buffers.

[0031] A clock tree synthesis optimization result output unit for outputting circuit information after buffer insertion, including the position information of the bottom - layer buffers, middle - layer buffers, and top - layer buffers, the signal topology structure after buffer insertion, and the connection relationship between the registers and the buffers.

[0032] The electronic device of the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, the clock tree synthesis optimization method based on hierarchical region division and buffer insertion as described above is implemented.

[0033] The computer - readable storage medium of the present invention stores a computer program. When the computer program is executed by the processor, the clock tree synthesis optimization method based on hierarchical region division and buffer insertion as described above is implemented.

[0034] Advantages: Compared with the prior art, the advantages of the present invention are as follows: (1) The present invention can quickly identify and eliminate clock delay and clock skew, especially ensuring the stability of the clock signal in complex circuits and meeting strict timing requirements. When inserting at the bottom layer and middle layer, a clustering algorithm, rectangular cutting, and diamond grid division strategy are adopted to precisely adjust the structure of the clock tree at each layer, effectively reducing the number of buffers, optimizing the distribution of the clock signal, ensuring a reduction in design power consumption and improving the quality of the clock signal; when inserting at the top layer, a dynamic clock optimization strategy is combined with an improved H - tree structure, enabling the present invention to adapt to different - scale design requirements, flexibly adjust the buffer positions, optimize the clock tree at each level, and improve the flexibility and versatility of the design. (2) The optimization results provided by the present invention, including relevant parameters such as buffer insertion positions, connection relationships, and clock tree distribution structures, can provide accurate bases for subsequent design verification and simulation, significantly improving the efficiency of circuit simulation and verification and ensuring design stability. Description of the Drawings

[0035] Figure 1 It is a flowchart of the clock tree synthesis optimization method of the present invention.

[0036] Figure 2 It is a schematic diagram of the rectangular cutting algorithm in an embodiment of the present invention.

[0037] Figure 3 It is a schematic diagram of the insertion result of the bottom - layer buffer in an embodiment of the present invention.

[0038] Figure 4 It is a schematic diagram of the principle of the position adjustment function in an embodiment of the present invention.

[0039] Figure 5 Schematic diagram of the rhombus grid division structure in the embodiment of the present invention.

[0040] Figure 6 Schematic diagram of the superiority structure analysis of the rhombus grid division in the embodiment of the present invention.

[0041] Figure 7 Schematic diagram of the analysis of the clock skew optimization principle in the embodiment of the present invention.

[0042] Figure 8 Schematic diagram of the analysis of the clock delay optimization principle in the embodiment of the present invention. Detailed implementation manners

[0043] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0044] As Figure 1 shown, the clock tree synthesis optimization method based on hierarchical region division and buffer insertion includes the following steps.

[0045] S1: Read circuit information through a C++ language program.

[0046] Read the input file through a C++ language program, and extract the register connection information and signal topology structure of the circuit. The information read from the input file includes the unit length unit, the circuit area size range floorplan, the clock source coordinates clock root, the sizes of the registers FF and buffers buffer, and the coordinates of all register instances FF instance. The information read in addition to the circuit information includes the resistance value r and capacitance value c of the internal wiring net of the chip per unit length, the maximum net_rc value limit, the maximum fanout limit Max_fanout, and the delay data of a single buffer, etc. The net_rc value is: , where in this formula refers to the net_rc value between unit 1 and unit 2. The unit can be a buffer or a register. Here, unit 1 and 2 are just examples; is the Manhattan distance between two units, and the unit of the resistance value r is Ω / um, and the unit of the capacitance value c is pF / um. After reading the input information, visualize the extracted information to check whether the circuit information is accurately extracted.

[0047] S2: Divide the circuit area into rectangles and insert underlying buffers.

[0048] According to the circuit information extracted in step S1, use all registers as basic units, and adopt a rectangular cutting algorithm to divide the entire circuit area, as Figure 2As shown, the division will continue until the set indicator terminates. This indicator can be adjusted according to the circuit scale, for example, divided into six equal parts or eight equal parts. After the division, a preliminary model of the circuit is formed. The improved hierarchical clustering algorithm is used to group within each region according to the maximum fan-out limit, forming several point sets. Each point set is traversed, and the sum of the x and y coordinates of each point in the point set is calculated, and then the geometric center point is obtained by dividing by the size of the point set. The points in the point set are the basic units. The generated geometric center points may coincide with other points or exceed the region boundary. Therefore, after inserting the underlying buffer at the geometric center, the overlap rate is detected. If the overlap rate is 0, the underlying buffer is inserted; otherwise, the position is adjusted through the position adjustment function to ensure that the finally generated points are valid. Finally, a buffer is inserted at each corrected geometric center point. These operations provide strong support for minimizing the number of inserted buffers and optimizing clock skew.

[0049] As Figure 3 shown, this circuit structure schematic diagram shows the layout of the entire circuit region after inserting the underlying buffer. The red dots in the figure represent registers (FF), the blue dots represent the inserted buffers (BUF), and the green pentagrams are the clock sources (CLK). This step does not connect the clock source, only showing the position information of the clock source; the entire region is continuously divided using the rectangular region cutting algorithm, then grouped using the hierarchical clustering algorithm, and finally a buffer is inserted at the geometric center and the overlap rate is detected. As Figure 4 an example, the green rectangles represent buffers, and the blue rectangles represent registers. When detecting the overlap rate, it will be based on whether the upper, lower, left, and right sides of the rectangular dimensions of the buffers and registers intersect to detect whether there is an overlap in the region between all the registers and buffers in each point set. If it is found that the upper side of a buffer intersects, it will be adjusted downward by the preset step length through the position adjustment function; if the upper side still intersects after re-detection, it will be moved downward by twice the step length. The adjustment strategies for the other intersecting sides are the same as those for the upper side intersection until it is adjusted until none of the upper, lower, left, and right sides intersect, and then the final position of the inserted buffer will be determined.

[0050] S3: Divide the circuit into diamond-shaped regions and insert middle-layer buffers.

[0051] In this step, since the middle-layer division requires multi-level division, the first-level middle-layer division is named , and the last-level middle-layer division is named . During the process of inserting buffers into the middle layer, the buffers inserted at the bottom layer need to be used as the basic units, and then the entire circuit region is divided into diamond shapes, with Figure 5For example, the entire area is divided into several rhombuses and triangles at the edges, where the red dots represent the geometric center positions to be inserted into the middle - layer buffer. The reason for using the rhombus grid division is as Figure 6 shown, Figure 6 the rectangle shown in (a) in Figure 6 and the rhombus shown in (b) in

[0052] have equal areas. The blue dots represent the geometric centers, and the black triangles represent the basic units within the area. Since the calculation of the overall circuit delay is based on the Manhattan distance, as shown by the red line in the figure, compared with the rectangle division, using the rhombus division can make the maximum Manhattan distance from the points within the area to the geometric center smaller, thus achieving lower clock delay and clock skew.

[0053] When using the improved rhombus grid division algorithm to divide the circuit area with the bottom - layer buffer as the basic unit, it is necessary to calculate the ratio of the number of bottom - layer buffers to the maximum fan - out, then round up. The middle layer is initially divided according to this value, and it is determined whether the number of basic units within a single area meets the maximum fan - out limit. If not, the basic unit farthest from the geometric center is found and re - allocated until the number of basic units in each area meets the maximum fan - out limit. Then, the sum of the net_rc values between the geometric center and each basic unit is compared with the maximum net_rc value. If it exceeds the limit, the basic unit farthest from the geometric center is re - allocated. Finally, a buffer is inserted at the geometric center to detect the overlap rate. If the overlap rate is 0, the buffer is inserted into the middle layer; otherwise, it is adjusted through the position adjustment function. Since the basic unit has changed from a register to a bottom - layer buffer, the position adjustment function will modify the size parameters until all the limits are met and then insert it. After the buffers in the layer are inserted, the layer will be divided with the buffers inserted in as the basic units. This step will be iterated until the buffers in the

[0054] S4: Insert the top - layer buffer based on the H - tree structure.

[0055] In this step, since the top - layer division also requires multi - level division, the first - layer top - layer division is named and the last - layer top - layer division is named 。The H-tree structure is a symmetric and recursively fractal clock distribution structure. The top-level buffer is inserted through the H-tree structure. After the H-tree structure is generated, the buffers in the H-tree structure are dynamically adjusted. When processing to the k-th layer (k ≥ 3), the dynamic adjustment function is called to finely adjust the position of each buffer in the current layer. Since the buffers in the k-th layer are generated at the geometric center of the buffers in the (k - 1)-th layer and the clock skew is small, it is necessary to consider the net_rc value from the buffers in the k-th layer to the buffers in the (k - 2)-th layer. For example, if there is a buffer p in the k-th layer, and there are two buffers q1 and q2 in the (k - 1)-th layer connected to p, calculate the maximum net_rc values from q1 and q2 in the (k - 1)-th layer to the buffers in the (k - 2)-th layer, denoted as c1 and c2, and then add the net_rc values from p to q1 and q2, denoted as net_rc1 and net_rc2. Compare the values of net_rc1 + c1 and net_rc2 + c2. If net_rc1 + c1 is larger, adjust in the direction of q1; if net_rc1 + c2 is larger, adjust in the direction of q2; if the two are equal, do not adjust the position of p. For Figure 7 example, after generating the basic H-tree structure, as Figure 7 shown in (a) below, where the points all represent the top-level buffers, and different colors represent the buffers in different layers of the top level. It can be seen that the distances from buf_1 to the orange points, namely buf_2 and buf_3, are the same, and the distances to buf_4, buf_5, buf_6, and buf_7 are also the same. However, the distance deviation from buf_1 to buf_8 and buf_9 is relatively large, which will cause a large clock skew. Therefore, it is necessary to dynamically adjust the position of the buffers in the layer of buf_2 to balance the clock skew. The schematic diagram after adjustment is as Figure 7 shown in (b) below. Using such a structure to insert the top-level buffer can ensure that the clock skew of the top level is 0.

[0056] After inserting the top-level buffer, a dynamic clock optimization strategy is also needed to continue dynamically adjusting the position of the top-level buffer to optimize the global clock skew and clock delay. This method determines whether to adjust the buffer position by judging the coordinate information. For example, there is a buffer b1 in the k-th layer, the buffer in the (k - 1)-th layer connected to this buffer is b2, and the buffer in the (k - 2)-th layer connected to b2 is b3. If the x coordinate of b3 is between the abscissas of b1 and b2, then b2 needs to be adjusted along the x coordinate towards b1. If the y coordinate of b3 is between the ordinates of b1 and b2, then b2 needs to be adjusted along the y coordinate towards b1. Taking the Figure 8 schematic diagram in (a) below as an example, buf_11 is A buffer inserted in the layer, buf_12 and buf_13 are The buffer inserted in the layer, buf_10 is The buffer inserted in the layer, because Buf_11 of the layer is inserted at the center of the distance between buf_12 and buf_13. Therefore, directly inserting buf_1 may cause the transmission route to repeat, resulting in an increase in the overall delay, or may also cause an increase in the overall clock skew. Therefore, the position of buf_11 needs to be dynamically adjusted according to the positions of buf_12 and buf_13, as Figure 8 shown in (b) of, to avoid the problems of increased clock delay and aggravated clock skew caused by detouring. This step still needs to perform an overlap rate detection before determining the final position of inserting the top-layer buffer, and connect the clock source after the buffer insertion is completed.

[0057] S5: Output the optimized circuit design information.

[0058] After the optimization is completed, the optimized circuit clock tree synthesis information is output to a file, including register position information, buffer insertion position information, optimized signal topology, and network connection relationships between registers and buffers and between different buffers. The output information facilitates subsequent circuit verification and iterative design.

[0059] The method of the present invention is verified through specific experiments below.

[0060] In this embodiment, a traditional heuristic algorithm: Greedy Heuristic is selected for comparative testing with the method of the present invention. In the test, three different circuits (case1, case2, case3) are selected. The three circuits are all circuit files containing 100K - 200K unevenly distributed FFs, and all FFs are of the same type and size. The experimental results are compared by means of scientific argumentation to verify the real effect of the method of the present invention.

[0061] Table 1: Operating result table of the method of the present invention

[0062]

[0063] Table 2: Operating result table of the clock tree synthesis optimization method based on the greedy algorithm

[0064]

[0065] The results show that the traditional method can achieve the expected clock tree synthesis and optimization effects in circuits of smaller scale. However, in medium-scale and larger-scale circuits, its running time increases and the optimization accuracy rate significantly decreases. While the clock tree synthesis and optimization method based on hierarchical region division and buffer insertion effectively optimizes clock delay and clock skew within a reasonable algorithm running time, and significantly reduces the number of buffer insertions. For example, in case 3, the number of register insertions is reduced to 3,293, which is 276 less than that of the heuristic algorithm. At the same time, this method performs excellently in optimizing skew and latency. Among them, in case 3, the skew is reduced by about 7 picoseconds and the latency is reduced by about 10 picoseconds. In case 2, the skew is reduced by about 2 picoseconds, comprehensively improving the comprehensive performance and design reliability of the clock tree.

[0066] The clock tree synthesis and optimization system based on hierarchical region division and buffer insertion described in the present invention includes:

[0067] A circuit information extraction unit, which is used to extract circuit information, including register connection information and signal topological structure;

[0068] A bottom-layer buffer insertion unit, which uses all registers as basic units, divides the circuit area by using the rectangular cutting method, and inserts bottom-layer buffers at the geometric center points of each first area obtained after division;

[0069] A middle-layer buffer insertion unit, which uses the upper-layer buffers as basic units, divides the circuit area by using the diamond grid division method, and inserts middle-layer buffers at the geometric center points of each second area obtained after division; until n layers of middle-layer buffers are inserted; n is the number of layers of middle-layer buffers;

[0070] Among them, the upper-layer buffer of the nth layer of middle-layer buffers is the (n - 1)th layer of middle-layer buffers, and the upper-layer buffer of the first layer of middle-layer buffers is the bottom-layer buffer;

[0071] A top-layer buffer insertion unit, which inserts m layers of top-layer buffers according to the H-tree structure; m is the number of layers of top-layer buffers;

[0072] A clock tree synthesis and optimization result output unit, which is used to output the circuit information after inserting buffers, including the position information of bottom-layer buffers, middle-layer buffers and top-layer buffers, the signal topological structure after inserting buffers, and the connection relationship between registers and buffers.

[0073] The electronic device described in the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements the clock tree synthesis and optimization method based on hierarchical region division and buffer insertion.

[0074] The computer-readable storage medium according to the present invention stores a computer program, and when the computer program is executed by a processor, the clock tree synthesis optimization method based on hierarchical region division and buffer insertion is implemented.

[0075] The computer-readable storage medium may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory or any other medium that can be used to store program code in the form of instructions or data structures and can be accessed by a computer.

[0076] The processor is configured to execute the computer program stored in the memory to implement each step in the method described in the above embodiments.

Claims

1. A clock tree synthesis optimization method based on hierarchical region partitioning and buffer insertion, characterized in that: The steps include: Step 1, extracting circuit information, including register connection information and signal topology; Step 2, taking all registers as basic units, dividing the circuit area by using a rectangular cutting method, and inserting a bottom buffer at the geometric center point of each first area obtained after the division; Step 3, using the upper layer buffer as the basic unit, dividing the circuit area using the diamond grid division method, inserting the middle layer buffer at the geometric center point of each second area obtained after the division; repeating step 3 until n layers of middle layer buffers are inserted; n is the number of layers of the middle layer buffer; The upper buffer of the nth layer middle buffer is the n-1th layer middle buffer, and the upper buffer of the 1st layer middle buffer is the bottom buffer; Step 4, insert the m-layer top buffer according to the H-tree structure; m is the number of layers of the top buffer; Step 5, output the circuit information after the buffer is inserted, including the location information of the bottom buffer, the middle buffer and the top buffer, the signal topology after the buffer is inserted, and the connection relationship between the register and the buffer.

2. The clock tree comprehensive optimization method based on hierarchical region partitioning and buffer insertion according to claim 1, characterized in that: In step 2, a hierarchical clustering algorithm is used to divide each first region into a number of point sets according to the maximum fan-out number limit, where the points in the point set are registers; the geometric center of each point set is calculated and inserted into the bottom buffer.

3. The clock tree comprehensive optimization method based on hierarchical region partitioning and buffer insertion according to claim 1, characterized in that: In step 2 and step 3, overlap detection is performed on each bottom layer buffer and middle layer buffer. If there is overlap, the position of the geometric center is adjusted and then the bottom layer buffer and the middle layer buffer are inserted, including the following steps: Check whether the edges of the bottom buffer and the middle buffer intersect with other buffers and / or registers. If so, move the position of the geometric center in the opposite direction of the intersecting edges, and then re-perform overlap detection until any edge of the bottom buffer and the middle buffer does not intersect with other buffers and / or registers.

4. The clock tree comprehensive optimization method based on hierarchical region partitioning and buffer insertion according to claim 1, characterized in that: In step 3, the above-layer buffer is used as a basic unit, the circuit area is divided by using a diamond grid division method, and the middle-layer buffer is inserted at the geometric center point of each second area obtained after the division, including: Calculate the ratio of the number of buffers in the previous layer to the maximum fan-out number, round up the ratio to get the middle-layer division number, perform diamond division according to the middle-layer division number to get several second areas, calculate whether the number of basic units in each second area meets the maximum fan-out number limit, if not, re-divide the basic unit farthest from the geometric center of the second area until each second area meets the maximum fan-out number limit; calculate the sum of the net_rc values ​​between the basic unit and the geometric center in each second area ,like If the maximum net_rc value limit in the second region is exceeded, the basic unit farthest from the geometric center of the second region is re-divided until each second region meets the maximum net_rc value limit; a middle buffer is inserted at the geometric center point of each second region; The net_rc value is , where D is the Manhattan distance between the basic unit and the geometric center, r and c are the resistance and capacitance values ​​of the line between the basic unit and the geometric center, respectively.

5. The clock tree comprehensive optimization method based on hierarchical region partitioning and buffer insertion according to claim 1, characterized in that: In step 3, the number of buffers in layer n is in the range of [30,40].

6. The clock tree comprehensive optimization method based on hierarchical region partitioning and buffer insertion according to claim 1, characterized in that: In step 4, for any two k-1th layer top buffers q1 and q2 connected to any k-th layer top buffer p, the net_rc values ​​from p to q1 and q2 are calculated, denoted as net_rc1 and net_rc2, respectively, and the maximum net_rc values ​​from q1 and q2 to the k-2th layer top buffer connected to them are calculated, denoted as c1 and c2; where k=3,4,5,…,m; If net_rc1+c1>net_rc2+c2, adjust the position of p toward q1; if net_rc1+c1<net_rc2+c2, adjust the position of p toward q2; if net_rc1+c1=net_rc2+c2, do not adjust the position of p; The net_rc value is , where D is the Manhattan distance between buffers, r and c are the resistance and capacitance of the wires between buffers, respectively.

7. The clock tree comprehensive optimization method based on hierarchical region partitioning and buffer insertion according to claim 1, characterized in that: In step 4, for any k-th layer top buffer p, any k-1-th layer top buffer q1 connected to p, and any k-2-th layer top buffer q3 connected to q1, where k=3, 4, 5, …, m; if the horizontal coordinate of q3 is between the horizontal coordinates of p and q1, the position of q1 is adjusted along the horizontal coordinate in the direction of p; if the vertical coordinate of q3 is between the vertical coordinates of p and q1, the position of q1 is adjusted along the vertical coordinate in the direction of p.

8. A clock tree synthesis optimization system based on hierarchical region partitioning and buffer insertion, characterized in that: include: A circuit information extraction unit, used to extract circuit information, including register connection information and signal topology; A bottom buffer insertion unit is used to divide the circuit area by using a rectangular cutting method using all registers as basic units, and insert a bottom buffer at the geometric center point of each first area obtained after the division; The middle-layer buffer insertion unit is used to divide the circuit area by using the upper-layer buffer as the basic unit using the diamond grid division method, and insert the middle-layer buffer at the geometric center point of each second area obtained after the division; repeat step 3 until n layers of middle-layer buffers are inserted; n is the number of layers of the middle-layer buffers; The upper buffer of the nth layer middle buffer is the n-1th layer middle buffer, and the upper buffer of the 1st layer middle buffer is the bottom buffer; A top-level buffer insertion unit, used for inserting m-level top-level buffers according to the H-tree structure; m is the number of layers of the top buffer; The clock tree synthesis optimization result output unit is used to output the circuit information after the buffer is inserted, including the location information of the bottom buffer, the middle buffer and the top buffer, the signal topology structure after the buffer is inserted, and the connection relationship between the register and the buffer.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into a processor, the clock tree synthesis optimization method based on hierarchical region partitioning and buffer insertion according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the clock tree synthesis optimization method based on hierarchical region partitioning and buffer insertion according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Clock tree construction method and system, chip, electronic equipment and storage medium

    CN116757150A

  • Clock tree comprehensive optimization method based on H tree

    CN117764024A