Clock tree synthesis method for realizing fast convergence of time sequence
Through the coordinated design technology of manual insertion buffer and clock tree, the clock tree structure is optimized, combined with multimode multiangle static timing analysis, the problem of timing convergence difficulties and long design time in the traditional clock tree comprehensive solution is solved, and the rapid timing convergence and high-quality clock tree design are achieved.
Patent Information
- Application Number
- CN202510282589.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-24
AI Technical Summary
The traditional clock tree comprehensive solution has problems such as difficulty in timing convergence and long design time.
The backbone of the clock structure is formed by manually inserting the buffer, combined with the clock tree collaborative design technology (CCOPT), the structural parameters of the clock tree are optimized, and the timing is checked through multimode multiangle static timing analysis (MMMC) to check whether the timing converges and make corresponding adjustments.
The rapid convergence of timing is achieved, reducing the number of iterations and comprehensive time of clock tree design, and the resulting clock tree has small clock deviation and good timing results.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of digital integrated circuit back-end design, and particularly relates to a clock tree synthesis method for achieving fast timing convergence. Background Art
[0002] With the advent of the post-Moore era, people not only have an increasing demand for integrated circuits, but also put forward more stringent requirements for the integration, performance, and power consumption of integrated circuits. These developments undoubtedly pose higher requirements for the timing of chips. Among the entire digital circuit back-end physical design, the step most closely related to timing is clock tree synthesis (CTS). Especially today when the process size is continuously shrinking, the proportion of interconnect delay is increasing, making it more difficult to obtain relatively small clock skew after clock tree synthesis.
[0003] The impact of clock tree synthesis on chip design is so great, not only because it is a key step in the back-end design process of integrated circuit design, but also because it is closely related to the timing of the entire design. To obtain a high-quality clock tree, it is far from enough to intervene only during clock tree synthesis. It is also necessary to start from each step of the back-end design process to find the most reasonable clock network design method, so as to reduce the power consumption of the clock network as much as possible on the premise of meeting the design requirements and achieving timing convergence, and further optimize the performance of the chip. Thus, researching the back-end design technology of integrated circuits under deep nano process nodes has great significance and research value for reducing key problems of chips, shortening the product design cycle, improving chip stability, and meeting project design goals. Generally speaking, traditional clock tree synthesis schemes have problems of relatively difficult timing convergence and long design time.
[0004] The information disclosed in this background art section is only intended to increase the understanding of the overall background of the present invention, and should not be regarded as an admission or any form of suggestion that this information constitutes prior art already known to those of ordinary skill in the art. Summary of the Invention
[0005] The purpose of the present invention is to provide a clock tree synthesis method for achieving fast timing convergence, aiming to reduce the number of timing violations existing in the back-end design process, so as to solve the problems of difficult timing convergence and long design time existing in current traditional clock tree synthesis schemes.
[0006] To achieve the above purpose, the present invention provides the following technical solutions:
[0007] A clock tree synthesis method for achieving fast timing convergence, comprising the following steps:
[0008] S1. After completing the layout of the macro cell, obtain the step-by-step information of the module position of the macro cell, and plan the general direction of the clock tree in advance;
[0009] S2. By means of manually inserting buffers, form the backbone of the clock structure, place the last-level buffer of the clock tree backbone at the center of the physical position of the module, and then complete the placement of the chip, and enter the clock tree synthesis;
[0010] S3. Based on the manually placed clock tree trunk and branches, configure the relevant parameters of the clock concurrent optimization (CCOPT) technology; then perform clock tree synthesis based on the clock backbone structure, and finally check whether the timing in the clock tree synthesis stage converges; if so, perform the routing task; if not, it is necessary to adjust the position and driving size of the last-level buffer of the clock tree trunk;
[0011] S4. After the routing is completed, check whether the timing converges through multi-mode multi-corner static timing analysis (MMMC); if so, perform sign-off; if not, it is necessary to make adjustments, specifically including: the distribution positions of the clock tree trunk and branches, re-layout or adjust the structural parameters in the clock tree synthesis stage.
[0012] Preferably, in S2, when forming the backbone of the clock structure by manually inserting buffers, the inserted buffers are high-driving force buffers, specifically one or more of BUF*D24*, BUF*D48*, and BUF*D64*.
[0013] Preferably, in S3, the clock signal CLK of the clock tree trunk part is generated by the top-level clock unit, the CLK signal is transmitted along the Trunk drive chain, and reaches the multi-source driver MBUF through the buffers of each branch, and then is transmitted through the clock tree branches, so that the clock signal reaches each register evenly.
[0014] Preferably, in S3, at different branch points of the clock tree trunk, clock trees with the required number of levels are generated according to the distribution range of the registers, and then the CCOPT technology is used for clock tree synthesis to check whether the timing meets the requirement: the clock skew of the clock tree is less than 0.3 ns.
[0015] Preferably, in S4, before performing MMMC, it is necessary to prepare the routed netlist, the extracted parasitic parameter file, the clock constraint file, and the working environment setting file.
[0016] Compared with the prior art, the present invention has the following beneficial effects: The clock tree synthesis method for realizing fast timing convergence of the present invention is conducive to obtaining a clock tree design with high quality, small clock skew, and meeting timing requirements. To a certain extent, it can reduce the number of iterations in clock tree design, the time spent on clock tree synthesis, obtain a smaller clock skew, and good timing results, which has good practical significance. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a schematic flowchart of the method of the present invention;
[0018] Figure 2 is a schematic diagram in the layout of the distributed clock tree + CCOPT in the method of the present invention;
[0019] Figure 3 is a schematic structural diagram of the distributed clock tree + CCOPT in the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The technical solutions of the present invention will be described clearly and completely below. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work shall fall within the protection scope of the present invention.
[0021] Refer to the attached Figure 1 , the present invention proposes a clock tree synthesis method for realizing fast timing convergence, which specifically includes the following steps:
[0022] Step 1: Extract all the registers of the same module and obtain their position information;
[0023] Step 2: Form the main trunk of the clock structure by manually inserting buffers; connect the buffers and inverters through which the clock signal enters from the clock port, place the last-stage inverter and the divided clock in the middle of the entire module, set the states of the inverters, buffers, and divided-frequency registers with connection relationships to fixed, and set the interconnect lines between them to dont touch so that their states do not change during placement; then complete the layout of the chip and enter clock tree synthesis;
[0024] Step 3: Configure the relevant parameters of the CCOPT technology based on the manually placed clock tree main trunk and branches; then perform clock tree synthesis based on the clock main trunk structure, and finally check whether the timing in the clock tree synthesis stage converges; if so, perform the routing task; if not, it is necessary to adjust the placement position and driving size of the last-stage buffer of the clock tree main trunk;
[0025] Step 4. After the wire winding is completed, check whether the timing converges through MMMC; if it does, perform sign-off; if not, adjustments are required, mainly by adjusting the placement position of the last-level register in the main part of the clock tree to achieve the goal.
[0026] The clock tree synthesis method for achieving fast timing convergence proposed by the present invention designs the clock tree based on a distributed clock tree structure in the clock tree synthesis stage, uses the CCOPT engine, and at the same time applies a more accurate effective skew to evaluate the impact of clock skew. Finally, the timing is checked through MMMC to ensure that the sign-off standard is met.
[0027] The distributed clock tree means that after manually guiding the main trunk of the clock signal tree from the clock port to the specified optimal physical position, branching starts from this optimal position, and timing violations are caused by excessive skew. The schematic diagram of its structure and the schematic diagram of the physical position are as shown in the appendix Figure 2 and 3 shown.
[0028] In this embodiment, the clock signal CLK is generated by the top-level clock unit. The CLK signal propagates along the Trunk main drive chain, and reaches the multi-source driver MBUF through the buffers of each level of branch Branch, and then is transmitted through the clock tree branches, so that the clock signal reaches each register evenly. The main trunk and branches are formed by manually inserting buffer chains. These buffers have high driving force, and one or more of BUF*D24*, BUF*D48*, BUF*D64* or other buffers with high driving force can be selected; the entire clock network uses high-level metal traces. The high-level metal line is wide, has strong electromigration rate, and has strong conductivity to drive a large number of loads.
[0029] The CCOPT technology optimizes the clock path and the logic path simultaneously in the clock tree synthesis stage, directly optimizes the transmitted clock, and includes influencing factors such as on-chip effects and gated clocks. That is, the launch clock L, the capture clock C, and the combinational logic delay D are all taken as optimization objects. When the clock period T > L + D + C, the setup time convergence is satisfied.
[0030] Based on the layout, create a distributed clock main trunk and perform clock tree synthesis. It is necessary to configure the parameters of the CCOPT technology and the FCHT clock structure parameters, and then use the ccopt_design command to complete the clock tree synthesis and check the timing; if the timing is satisfied, perform wire winding and check the timing through MMMC to ensure that the sign-off standard is met; if the timing is not satisfied, the parameters need to be adjusted.
[0031] In the layout stage, the backbone and branches of the clock structure are generated by manually inserting buffers, and the attributes of these inserted buffers are set to fixed. The specific distribution of the backbone and branches is set according to the chip requirements. The driving unit of the backbone uses a buffer with high driving ability. In this embodiment, the driving unit of BUF*D24* is selected. At the same time, high-level metal traces are used to drive a large number of loads. The ecoAddRepeater command is used to insert buffers, the add_ndr and create_route_type are used to create routing rules, and finally the ccopt_design command is used to complete the entire clock tree synthesis.
[0032] In the clock tree stage, the cluster operator is used to quickly allocate clock sink points, saving the running time of clock tree synthesis, which is beneficial to iterative design and obtaining a high-quality clock tree; the useful skew is used to better balance the clock skew of each sink point and more accurately evaluate the impact of clock skew. The relevant commands are: set cluster true; set useful_skew true, and the tool will call the relevant operators for optimization during clock tree synthesis.
[0033] After clock tree synthesis, it is judged whether the timing of the design in this stage converges. If it converges, routing is performed. After routing, MMMC is used to judge whether the timing of the entire design converges to complete the timing sign-off work. The main purpose of MMMC is to analyze and solve the timing convergence problem from multiple working modes and process perspectives earlier before the physical layout, ensuring the reliability of the design timing under different process, voltage, and temperature (PVT) conditions. MMMC is beneficial to discovering problems earlier, reducing design risks, and improving the reliability of physical design.
[0034] The foregoing description of specific exemplary embodiments of the present invention is for purposes of illustration and exemplification. These descriptions are not intended to limit the invention to the precise forms disclosed, and obviously, many changes and variations are possible in light of the above teachings. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the invention and its practical applications, so that those skilled in the art can implement and utilize various different exemplary embodiments of the invention, as well as various different selections and changes. The scope of the invention is intended to be defined by the claims and their equivalents.
Claims
1. A clock tree synthesis method for achieving rapid timing convergence, characterized in that: The following steps are involved: S1. After completing the layout of the macro cell, obtain the step-by-step information of the module location of the macro cell and plan the general direction of the clock tree in advance; S2. Form the clock structure trunk by manually inserting buffers, place the last level buffer of the clock tree trunk at the physical center of the module, and then complete the chip layout and enter the clock tree synthesis; S3, based on the manually placed clock tree trunk and branches, configure the relevant parameters of the CCOPT technology; Then, clock tree synthesis is performed based on the clock trunk structure, and finally the timing of the clock tree synthesis stage is checked to see if it has converged; if so, the routing task is performed; if not, the position and driver size of the last level buffer of the clock tree trunk need to be adjusted; S4. After the routing is completed, check whether the timing is converged through MMMC; if yes, sign-off is performed; if not, adjustments are required, including: the distribution position of the clock tree trunk and branches, re-layout or structural parameters adjusted during the clock tree synthesis stage.
2. The clock tree synthesis method for achieving rapid timing convergence according to claim 1, characterized in that: In S2, a buffer is inserted manually to form the backbone of the clock structure. The buffer inserted is a high-driving buffer, specifically one or more of BUF*D24*, BUF*D48*, and BUF*D64*.
3. The clock tree synthesis method for achieving rapid timing convergence according to claim 1, characterized in that: The clock signal CLK of the clock tree trunk in S3 is generated by the top-level clock unit. The CLK signal is transmitted along the Trunk driver chain, driven by the buffers of each level of branches to reach the multi-source driver MBUF, and then transmitted through the clock tree branches, so that the clock signal reaches each register evenly.
4. The clock tree synthesis method for achieving rapid timing convergence according to claim 1, characterized in that: The clock tree trunk in S3 generates the required number of clock trees at different branch points according to the distribution range of registers, and then uses CCOPT technology to perform clock tree synthesis to check whether the timing is met.
5. The clock tree synthesis method for achieving rapid timing convergence according to claim 1, characterized in that: Before making MMMC in S4, you need to prepare the netlist after routing, the extracted parasitic parameter file, the clock constraint file and the working environment setting file.