Optimization method for removing buffer congestion based on tree dynamic programming algorithm
The buffer deletion problem in timing repair is optimized through the tree dynamic programming algorithm, which solves the problem of inaccurate estimation of delay results when adding buffers to repair timing, and achieves the effect of reducing local utilization and wiring congestion without affecting timing results.
Patent Information
- Application Number
- CN202111506477.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-12-10
AI Technical Summary
When the timing optimization tool adds buffer to repair timing, the connection delay in the delay result can only be estimated, which may cause some buffers to be added.
Using a method based on a tree dynamic programming algorithm, the tree structure extraction and tree dynamic programming buffer deletion algorithm are used to traverse the tree structure from the bottom up, and combined with the nonlinear delay model to recalculate the unit delay, and filter out the scheme that meets the timing requirements and deletes the most buffers.
Without destroying the existing timing results, local utilization is optimized, wiring congestion and short circuits are reduced, and timing repair efficiency is improved.
Smart Images

Figure CN114386351B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of chip design, and in particular relates to a method for optimizing buffer congestion removal based on a tree-shaped dynamic programming algorithm. Background Art
[0002] As the design requirements of integrated circuits increase, power chips usually need to compress the area as much as possible in order to reduce costs, and the chip utilization is high. During the iterative process of chip timing optimization, many buffers are added to fix the hold time violations of the chip, which aggravates the problem of excessive local utilization and causes short circuits. At this stage, it is too late to re-plan the area. The floorplan part has been basically determined due to electro-migration (EM) checks and voltage drop (IRdrop) checks, and it is difficult to modify. In order to continue with other timing repairs to ultimately achieve timing convergence, local utilization can be optimized in this late design stage without destroying the existing timing results.
[0003] For similar problems of excessive local utilization in designs, consider deleting some of the buffers added to fix hold time violations. These buffers may be redundant, and deleting them will not have any impact on the timing results. The main reason for redundant buffers is that timing repair is divided into many rounds, and the repair of multiple timing violations is carried out together, such as hold time violations, noise violations (crosstalk), transition time violations (transition), etc. During multiple rounds of optimization, redundant buffers may appear. In addition, when the timing optimization tool adds buffers to repair timing, it can only estimate the connection delay (netdelay) in the delay result, which is not the exact connection delay result. Therefore, some buffers may be added extra. Summary of the invention
[0004] The technical problem to be solved by the present invention is: to provide a method for optimizing buffer congestion removal based on a tree-shaped dynamic programming algorithm, so as to solve the problem that when a timing optimization tool adds a buffer to repair the timing, the net delay in the delay result can only be estimated, but not the exact result of the net delay; therefore, some technical problems such as the addition of multiple buffers may be caused.
[0005] Technical solution of the present invention:
[0006] A method for optimizing buffer congestion removal based on a tree-shaped dynamic programming algorithm, comprising:
[0007] Step 1: Tree structure extraction: According to the requirements of the tree dynamic programming algorithm, the buffers to be deleted are counted as tree nodes; any non-buffer cell, if there is a buffer in its fan-out that repairs the timing violation, is taken as the root. Starting from the root, trace downward until no buffer exists in the fan-out;
[0008] Step 2: Tree-shaped dynamic programming buffer deletion algorithm: For a tree structure, all nodes are buffers. When traversing these nodes, there are two possible solutions, one is to delete the buffer, and the other is to keep the buffer. Traverse the tree from bottom to top. If a buffer C fans out two buffers: A and B; then the solutions for the buffers of A and B are merged into four, and these four solutions are represented by the name of the buffer that is finally deleted: 0, A, B, AB; after traversal, all solutions are gathered at the root node of the tree, and the solution that meets the timing requirements and deletes the most buffers is selected and returned.
[0009] After buffer C, the schemes further evolved into eight types: 0, A, B, AB, C, AC, BC, and ABC, and the schemes were screened.
[0010] The reference cell, delay value, transition time and capacitance of the cell are described after the node name.
[0011] Since the worst timing situation may occur at different timing corners, all timing situations are reported in the distributed multi scenarios analysis (DMSA) mode; the timing situations of all units connected under the same root belong to a timing group.
[0012] The delay calculation after buffer deletion is completed based on the calculation method of non-linear delay model (NLDM).
[0013] The unit delay is related to the input transition time and the output load. After the buffer is deleted, the unit delay of the upper and lower levels will change accordingly. The delay recalculation method after deleting the buffer is:
[0014] For a tree structure: ABC, where B is a buffer, if the solution is to delete B, then the output load of A will be
[0015] loadofnet(AB)+inputCofbufferB
[0016] becomes:
[0017] loadofnet(AB)+outputloadofbufferB;
[0018] Thus, the delay of A and the output transition time (output transition) of A are calculated by looking up the table with unchanged input transition and changed output load. Since the output transition of a cell is the same as the input transition of the cell it fans out from, when B is deleted:
[0019] A's output transition = C's input transition;
[0020] At this time, for C, the outputload remains unchanged, and the inputtransition changes, so the delay of C after the change is obtained; at this time, the delay of A and C is greater than the result when B is not deleted; therefore, for a timing path, if there is a slack x on the hold time path, and the unit delay of bufferB on the path is y, the unit delay change value of the previous and next stages A and C is z; if x+z>y, B is deleted. In the cell timing lib, there are tables of unit delays of different units regarding inputtransition and outputload, and the unit delay is obtained by looking up the table.
[0021] When inputtransition and outputload are not on the preset grid points, the interpolation calculation method is applied. Let Z be the delay value, X be the inputtransition, and Y be the outputload. Z, X, and Y satisfy the following relationship:
[0022] Z=A+B·X+C·Y+D·X·Y
[0023] ABCD is set as a variable; for the non-grid inputtransition m, there is a smaller inputtransition node value α1, and a larger node value α2; for the non-grid outputload n, there is a smaller outputload node value β1, and a larger node value β2; assuming that α1, β1 correspond to delay values Z11, α2, β1 correspond to delay values Z21, α1, β2 correspond to delay values Z12, α2, β2 correspond to delay values Z22; then:
[0024] Z11=A+B·α1+C·β1+D·α1·β1
[0025] Z12=A+B·α1+C·β2+D·α1·β2
[0026] Z21=A+B·α2+C·β1+D·α2·β1
[0027] Z22=A+B·α2+C·β2+D·α2·β2
[0028] Obtain the value of ABCD, and bring in the inputtransition and outputload at the current non-grid point to obtain the approximate solution Zmn at the non-grid point
[0029] Zmn=A+B·m+C·n+D·m·n.
[0030] The screening method for the scheme includes: after calculating and deleting a buffer N, if the output transition of N's previous level will exceed the maximum transition time (maxtransition), the scheme is directly deleted; if a sink has no timing path under the current timing group, the buffer deletion scheme for the sink is deleted.
[0031] Beneficial effects of the present invention:
[0032] The present invention is a method for optimizing the congestion problem in timing repair based on a tree-shaped dynamic programming algorithm. In multiple rounds of iterations of timing optimization in the back-end design of power chips, if there is a problem of excessive local utilization or wiring congestion, returning to the area planning and layout planning stages may result in insufficient design time. It can be considered to delete the redundant buffers added during the timing optimization process without affecting the timing results; this method can reduce the local utilization, leaving space for subsequent timing repair iterations, and reduce wiring congestion and short circuits, solving the problem that when the timing optimization tool adds buffers to repair the timing, the connection delay (netdelay) in the delay result can only be estimated, not the exact connection delay result; therefore, it may cause technical problems such as adding more buffers. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is the principle diagram of the tree-shaped dynamic programming algorithm of the present invention;
[0034] Figure 2 A schematic diagram of a traversal process for a specific implementation method;
[0035] Figure 3 It is a schematic diagram of the tree structure extraction algorithm;
[0036] Figure 4 The overall flow chart of the algorithm is shown in Figure 2. DETAILED DESCRIPTION
[0037] The present invention proposes a buffer deletion congestion optimization method based on a tree-shaped dynamic programming algorithm. The timing analysis tool used is PrimeTime of Synopsys, the PR tool is ICCompiler of Synopsys, and the RC extraction tool is StarRC. Dynamic programming is a method of solving complex problems by decomposing the original problem into relatively simple sub-problems. The tree-shaped dynamic programming algorithm is an algorithm for dynamic programming on a tree structure. In the buffer structure of chip repair timing, the tree structure is very common and is suitable for optimization using a tree-shaped dynamic programming algorithm.
[0038] In a tree structure, using a tree dynamic programming algorithm is better than a greedy algorithm. The idea of the greedy algorithm is to input all buffer names and timing conditions into the algorithm, determine the timing conditions after deleting each buffer, and delete it if no new timing violations occur. If new timing violations occur, do not delete it. Determine the conditions of each buffer in turn, and eventually get all deletable buffers. However, the greedy algorithm will take the best solution under the current circumstances in each step of selection, which may cause the early operations to affect the later optimization in the tree structure. For example, deleting the buffers on the common path first affects the deletion of buffers on multiple branch paths. From a global perspective, the effect is not as good as the result obtained by the tree dynamic programming.
[0039] The optimization process of the tree dynamic programming algorithm can be divided into the following three steps:
[0040] Tree structure extraction: According to the requirements of the tree dynamic programming algorithm, the buffers that need to be deleted should be counted as nodes of the tree. Then any non-buffer cell can be used as the root if there is a buffer in its fan-out that repairs the timing violation. Starting from the root, trace downward until there is no buffer in the fan-out. In addition, in order to facilitate the calculation of timing conditions in the main algorithm, the reference cell, delay value, transition time, capacitance (load), etc. of the cell should be described after the node name. Note that since the worst timing situation may occur at different timing corners, all timing conditions should be reported in the distributed multi scenarios analysis (DMSA) mode. In addition, the timing conditions of all cells connected to the same root should belong to a timing group (timinggroup).
[0041] Delay calculation: The delay calculation after the buffer is deleted in this design is completed based on the calculation method of the non-linear delay model (NLDM). The unit delay is related to the input transition time and the output load. After the buffer is deleted, the unit delay of its upper and lower levels will also change accordingly. After deleting the buffer, the recalculation method for the delay is as follows:
[0042] For a tree structure: ABC, where B is a buffer, if the solution is to delete B, then the output load of A will be
[0043] loadofnet(AB)+inputCofbufferB
[0044] becomes:
[0045] loadofnet(AB)+outputloadofbufferB.
[0046] Therefore, the delay of A and the output transition time (output transition) of A can be calculated by looking up the table with unchanged input transition and changed output load. Since the output transition of a cell is the same as the input transition of the cell it fans out from, it can be considered that when B is deleted:
[0047] A's output transition = C's input transition.
[0048] At this time, for C, the outputload remains unchanged, and the inputtransition changes, so the delay after the change of C can also be calculated. Usually, the delay of A and C is greater than the result when B is not deleted. Therefore, for a timing path, if there is a slack x on the hold time path, and the unit delay of bufferB on the path is y, the unit delay of the previous and next stages A and C changes by z. If x+z>y, then B can be considered to be deleted.
[0049] In addition, in the cell timing lib, there will be a table of input transition and output load for different units, and the unit delay is obtained by looking up the table. However, in actual application, there are many cases where input transition and output load are not on the preset grid points. In this case, in the present invention, the interpolation calculation method can be applied, the same as the "report_delay_calculation" command in PT, assuming Z as the delay value, X as the input transition, and Y as the output load. Z, X, and Y satisfy the following relationship:
[0050] Z=A+B·X+C·Y+D·X·Y
[0051] Set ABCD as variables. For the non-grid inputtransition - m, there will be a smaller inputtransition node value α1, and a larger node value α2. Similarly, for the non-grid outputload - n, there will be a smaller outputload node value β1, and a larger node value β2. Assume that the delay value corresponding to α1, β1 is Z11, α2, β1 is Z21, α1, β2 is Z12, α2, β2 is Z22. Then:
[0052] Z11=A+B·α1+C·β1+D·α1·β1
[0053] Z12=A+B·α1+C·β2+D·α1·β2
[0054] Z21=A+B·α2+C·β1+D·α2·β1
[0055] Z22=A+B·α2+C·β2+D·α2·β2
[0056] Obtain the value of ABCD, and substitute the inputtransition and outputload at the current non-grid point to obtain the approximate solution Zmn at the non-grid point.
[0057] Zmn=A+B·m+C·n+D·m·n
[0058] Tree-shaped dynamic programming buffer deletion algorithm: For a tree structure, all nodes are buffers. When traversing these nodes, there are two possible solutions, one is to delete the buffer, and the other is to keep the buffer. Traversing the tree from bottom to top, if a buffer C fans out two buffers: A and B, then the solutions for the buffers of A and B can be combined into four. These four solutions can be represented by the name of the buffer that is finally deleted: 0, A, B, AB. Then after buffer C, the solutions further evolve into eight:
[0059] 0,A,B,AB,C,AC,BC,ABC.
[0060] There are some screening methods for these solutions: 1. Calculate the output transition of the previous level of N after deleting a buffer N. If it will exceed the maximum transition time (maxtransition), delete the solution directly. 2. If a sink has no timing path under the current timing group, delete the buffer deletion solution for the sink.
[0061] After traversal, all solutions are gathered at the root node of the tree, and the solution with no violation of the maintenance time and the largest number of buffers deleted is selected. Figure 1 The structure of the tree dynamic programming algorithm is explained. The root node of the extracted tree structure is input into the algorithm. If it is not a sink, it is recursively passed to the next level k.sub[0]. Until the sink node is found, its information is converted into the solution format Z through the solutions_sink function. Then regress upward. If there are other sinks at the same level, they are also converted into solution format, and the two solutions are merged through the solutions combine function. If the traversed node at this time is a buffer, the solutions deletebuffer function can be used to add the solution after the buffer is deleted to Z. If the traversed node at this time is the root node, all solutions are screened through the solutions filter function, and the solution that meets the timing requirements and deletes the most buffers is selected and returned.
[0062] Advantages of the present invention:
[0063] The tree-structured dynamic programming algorithm is used to delete redundant buffers in the back-end design of power chips, reducing the local utilization of the chip and preparing for subsequent timing optimization.
[0064] The buffers that can be considered for deletion are extracted in a tree structure to facilitate optimization using a tree structure algorithm.
[0065] The algorithm combines the NLDM model to recalculate the unit delay, making the estimation in the algorithm more accurate. The same interpolation calculation method as the PT command is used to estimate, making the delay value at non-grid points more accurate.
[0066] The redundant buffers of the repair timing are deleted to reduce the local utilization of the chip without destroying the existing timing optimization results, so as to ensure that the chip design can be completed within a limited area.
Claims
1. A method for optimizing buffer congestion removal based on a tree-shaped dynamic programming algorithm, characterized in that: It includes: Step 1: Tree structure extraction: According to the requirements of the tree dynamic programming algorithm, the buffers to be deleted are counted as tree nodes; any non-buffer cell, if its fan-out contains a buffer that repairs the timing violation, is taken as the root. Starting from the root, trace downward until no buffer exists in the fan-out; Step 2: Tree-shaped dynamic programming buffer deletion algorithm: For a tree structure, all nodes are buffers. When traversing these nodes, there are two possible solutions: one is to delete the buffer, and the other is to keep the buffer. Traverse the tree from bottom to top. If a buffer C fans out two buffers: A and B; then the solutions for the buffers A and B are merged into four. These four solutions are represented by the names of the buffers that are finally deleted: 0, A, B, AB; after traversing, all solutions are gathered at the root node of the tree, and the solution that meets the timing requirements and deletes the most buffers is selected and returned; The method further comprises: After bufferC, the schemes further evolved into eight: 0, A, B, AB, C, AC, BC, and ABC; the schemes were screened; Among them, the screening method for the plan includes: after calculating and deleting a bufferN, if the outputtransition of N's previous level will exceed the maximum transition time maxtransition, then the plan is directly deleted; if a sink has no timing path under the current timing group, then the buffer deletion plan for the sink is deleted.
2. The method for optimizing buffer congestion removal based on a tree-shaped dynamic programming algorithm according to claim 1, characterized in that: After the node name, the reference cell, delay value delay, transition time transition and capacitance load of the cell are described.
3. The method for optimizing buffer congestion removal based on a tree-shaped dynamic programming algorithm according to claim 2, characterized in that: Since the worst timing situation may occur at different timing corners, all timing situations are reported in the distributed multiscenarios analysis (DMSA) mode; the timing situations of all units connected under the same root belong to a timing group timinggroup.
4. The method for optimizing buffer congestion removal based on a tree-shaped dynamic programming algorithm according to claim 1, characterized in that: The delay calculation after buffer deletion is completed based on the calculation method of non-linear delay model (NLDM).
5. The method for optimizing buffer congestion removal based on a tree-shaped dynamic programming algorithm according to claim 4, characterized in that: The unit delay is related to the input transition time inputtransition and the output load outputload. After the buffer is deleted, the unit delay of the upper and lower levels will change accordingly. The delay recalculation method after deleting the buffer is: For a tree structure: ABC, where B is a buffer, if the solution is to delete B, then the output load of A changes from the original loadofnet(AB)+inputCofbufferB to: loadofnet(AB)+outputloadofbufferB; Therefore, the delay of A and the output transition time of A are calculated by looking up the table with the unchanged input transition and the changed output load. Since the output transition of a cell is the same as the input transition of the cell it fans out from, when B is deleted: A's output transition = C's input transition; At this time, for C, the outputload remains unchanged, and the inputtransition changes, so the delay of C after the change is obtained; at this time, the delay of A and C is greater than the result when B is not deleted; therefore, for a timing path, if there is a slack x on the hold time path, and the unit delay of bufferB on the path is y, the unit delay change value of the previous and next stages A and C is z; if x+z>y, B is deleted. In the cell timing lib, there are tables of unit delays of different units regarding inputtransition and outputload, and the unit delay is obtained by looking up the table.
6. The method for optimizing buffer congestion removal based on a tree-shaped dynamic programming algorithm according to claim 5, characterized in that: When inputtransition and outputload are not on the preset grid points, the interpolation calculation method is applied. Let Z be the delay value, X be the inputtransition, and Y be the outputload. Z, X, and Y satisfy the following relationship: Z=A+B·X+C·Y+D·X·Y ABCD is set as a variable; for the non-grid inputtransition m, there is a smaller inputtransition node value α1, and a larger node value α2; for the non-grid outputload n, there is a smaller outputload node value β1, and a larger node value β2; assuming that α1, β1 correspond to delay values Z11, α2, β1 correspond to delay values Z21, α1, β2 correspond to delay values Z12, α2, β2 correspond to delay values Z22; then: Z11=A+B·α1+C·β1+D·α1·β1 Z12=A+B·α1+C·β2+D·α1·β2 Z21=A+B·α2+C·β1+D·α2·β1 Z22=A+B·α2+C·β2+D·α2·β2 Obtain the value of ABCD, and bring in the inputtransition and outputload at the current non-grid point to obtain the approximate solution Zmn at the non-grid point Zmn=A+B·m+C·n+D·m·n.