Key Path Delay Optimization Method and System Based on Logic Depth-Driven Graph Partitioning

Through the method based on logic-deep-driven diagram division, the critical path delay of the microprocessor is optimized, which solves the problem of difficulty in achieving optimal circuit configuration and back-end optimization in the prior art, and achieves more efficient critical path delay optimization and performance improvement.

CN119918471BActive Publication Date: 2025-07-01NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510413023.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-01
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

In the prior art, when optimizing the critical path delay of a microprocessor, it is difficult to achieve the optimal circuit configuration, and the back-end optimization is limited by the standard unit library, which cannot effectively reduce the critical path delay.

Method used

The critical path delay optimization method based on logic depth-driven diagram division is adopted. By constructing the critical path logic cone, sub-circuits with optimization value are identified, and these sub-circuits are optimized through the sub-graph logic reconstruction method to obtain a reconstruction logic unit with better delay.

Benefits of technology

Without changing the original design, it effectively reduces critical path delays, improves microprocessor performance, and improves the scalability and reusability of the design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119918471B_ABST
    Figure CN119918471B_ABST
Patent Text Reader

Abstract

The present invention discloses a critical path delay optimization method and system based on logic depth-driven graph partitioning. The method includes the steps of: 1) obtaining a logic netlist of an integrated circuit; 2) constructing a critical path logic cone based on the logic netlist; 3) screening out sub-circuits with optimization value from the logic cone by a logic depth-driven graph partitioning method; 4) simplifying the sub-circuits with optimization value by a sub-graph logic reconstruction method to obtain a reconstructed logic unit with better delay; 5) screening out the reconstructed logic units corresponding to sub-graphs with a reuse rate greater than a preset value and replacing them into the initial logic netlist to obtain an optimized logic netlist; 6) performing equivalence checking on the initial logic netlist and the optimized logic netlist; 7) performing delay evaluation on the logic netlist passing the equivalence checking, and if the evaluation passes, obtaining the final reconstructed logic unit and logic netlist. The present invention has the advantages of low cost, effectively reducing the critical path delay, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the technical field of microprocessor design, and particularly relates to a critical path delay optimization method and system based on logic depth-driven graph partitioning. Background Art

[0002] Vigorously promoting the independent research, design, and manufacturing of high-performance microprocessors is a key measure to enhance the competitiveness of the integrated circuit (IC) industry and effectively reduce external dependence, and is of great significance for aspects such as scientific and technological progress, economic development, and national defense security. Currently, large-scale complex microprocessor designs usually utilize electronic design automation (EDA) tools, based on commercial standard cell libraries, and are implemented through semi-custom design methods. This design method can significantly shorten the complex IC design process and accelerate the product time to market, but often cannot achieve the optimal circuit configuration. Especially under the constraint of a relatively high main frequency, there are still a large number of critical paths with timing violations in the logic netlist obtained under the harsh constraints of commercial EDA tools. When the critical path delay exceeds the set clock cycle constraint, it will reduce the design main frequency of the microprocessor, thereby restricting the further improvement of its performance. The critical path delay is the key factor restricting the performance and main frequency improvement of microprocessors, directly determining the maximum operating frequency and data processing rate of the processor. Therefore, optimizing the critical path delay is of great significance in the design of high-performance microprocessors.

[0003] Optimizing the critical path delay is a multi-faceted and multi-level complex process, involving multiple links in the microprocessor design. After applying methods such as pipeline insertion, retiming, logic replication, and logic optimization to complete the optimization in the design stage, there may still be room for further optimization in the design. Since the front-end design and optimization iterations require a large amount of resources, the cost of iteration is high and the number of times is limited. Therefore, it is an important means to improve performance at limited cost to effectively reduce the critical path delay through backend optimization techniques without changing the original design. Currently, the backend optimization is usually carried out for specific designs, which may not be applicable to subsequent design iterations or expansions, restricting the scalability and reusability of the design, and the optimization at this stage is limited by the standard cell library and often cannot achieve the optimal circuit configuration. Summary of the Invention

[0004] Aiming at the technical problems existing in the prior art, the present invention provides a critical path delay optimization method and system based on logic depth-driven graph partitioning that can effectively reduce the critical path delay.

[0005] To solve the above technical problems, the technical solution proposed by the present invention is as follows:

[0006] A critical path delay optimization method based on logic depth-driven graph partitioning, comprising the steps of:

[0007] 1) Obtain the logic netlist, the given standard cell library, and the timing constraint conditions in the semi-custom design process of the integrated circuit;

[0008] 2) Analyze the characteristics of the critical path in the logic netlist, and construct a critical path logic cone based on the given standard cell library and timing constraint conditions;

[0009] 3) Screen out the sub-circuits with optimization value from the critical path logic cone through the logic depth-driven graph partitioning method;

[0010] 4) Simplify the sub-circuits with optimization value through the sub-graph logic reconstruction method to obtain a reconstructed logic unit with better delay;

[0011] 5) Screen out the reconstructed logic units corresponding to the sub-graphs with a reuse rate greater than the preset value and replace them into the initial logic netlist to obtain an optimized logic netlist;

[0012] 6) Perform equivalence checking on the initial logic netlist and the optimized logic netlist; if the equivalence checking is passed, proceed to the next step;

[0013] 7) Perform delay evaluation on the logic netlist that passes the equivalence checking. If the evaluation is passed, obtain the final reconstructed logic unit and the optimized logic netlist.

[0014] Preferably, in step 2), the direction of constructing the critical path logic cone is exactly opposite to the signal transmission direction; first, obtain the end point of the critical path from the logic netlist; then, gradually backtrack from the end point to the output port of the timing unit and stop. In this process, the module hierarchy and the internal connections of the module are not retained, only the standard cells and their port information, as well as the logical relationships between them are retained; by identifying the driving nodes corresponding to all input ports of the node, all the standard cells and their connection relationships between the end point of the critical path and the timing unit can be obtained in sequence; in the constructed logic cone, each obtained standard cell is regarded as a node; where the timing unit is a flip-flop, a static random access memory, or an IP in the process library, and its output is defined as the starting point of the logic cone.

[0015] Preferably, in step 2), in the critical path logic cone, the logical depth of the standard cell corresponding to the node is used as its weight, stored in the form of a weighted directed acyclic graph, and visually displayed; where each node contains one or more logic gates inside, and the logical depth of the node is the number of logic gates on the longest path from the input to the output of the node.

[0016] Preferably, the specific steps of step 3) are:

[0017] Initialize the sub - graph structure: Consider each node in the logic cone as an independent minimum sub - graph, and calculate the logical depth of each sub - graph; the logical depth of a sub - graph is the logical depth on the longest path from the input port to the output port of the sub - graph.

[0018] Search for all pairs of sub - graphs that can be merged; the pairs of sub - graphs for which merging operations are performed must satisfy: the merged graph conforms to the port - quantity constraint, and the logical depth of the newly merged graph is greater than the logical depth of the sub - graphs before merging.

[0019] Calculate the logical - depth increment: For each pair of sub - graphs, calculate the logical - depth increment after merging; the greater the logical - depth increment, the greater the merging benefit; if the logical - depth increment is less than or equal to 0, it means that this merging method is not advisable; the logical - depth increment is the difference between the logical depth of the newly merged graph and the average logical depth of the pair of sub - graphs before merging.

[0020] Select the best merge: For all pairs of sub - graphs identified as possible for merging operations, find the pair of sub - graphs with the largest logical - depth increment after merging, merge to obtain a new sub - graph; replace the merged sub - graphs with the new sub - graph and update the sub - graph structure.

[0021] Iteratively perform the merging operation until there are no more pairs of sub - graphs to merge, and obtain the final partitioning result.

[0022] Preferably, the specific steps of step 4) are as follows:

[0023] Identify ports: Identify the input ports and output ports of the sub - graph to ensure that the port configuration of the reconstruction unit is consistent with the sub - graph.

[0024] Logical simplification: Implement two - level logical simplification for each output port of the sub - graph; specifically, find the set of minimum prime implicants by identifying and merging implicants, and then obtain the minimum SOP expression of the circuit to achieve two - level logical simplification to reduce the logical level of the circuit.

[0025] Extract inverted logic: On the basis of two - level logical simplification, extract and simplify all inverted logic upward, and finally obtain a new optimized reconstruction logic unit.

[0026] Preferably, in step 5), the specific steps for screening the reconstruction logic units corresponding to sub - graphs with a reuse rate greater than a preset value are as follows:

[0027] First, perform a topological sort on the nodes in the logic cone to form a linear sequence in logical pre - order.

[0028] Then, based on the sorting results, matching is performed sequentially in the form of a graph to match the part in the logic cone that has the same graph structure as the sub-graph; units with the same function are represented by nodes of the same form, the connections between the nodes represent the connections between the units, and arrows are used to indicate the data flow direction, while buffer units are directly represented in the form of connections rather than nodes;

[0029] The reuse rate of each sub-graph is statistically counted and sorted in sequence, and the sub-graphs with a reuse rate greater than a preset value are selected to obtain the corresponding reconstructed logic units.

[0030] Preferably, the specific process of the equivalence check in step 6) is as follows:

[0031] Taking the initial logic netlist as a reference model, on the premise that other irrelevant conditions remain the same, the equivalence between the initial logic netlist and the optimized logic netlist is verified. When the logical functions corresponding to the two netlists are equivalent, it is proved that the optimized logic cone is equivalent to the original logic cone, and thus the equivalence check is passed.

[0032] Preferably, in step 7), when evaluating the delay of the logic netlist that has passed the equivalence check, the change amount of the logical depth on the critical path before and after optimization is used as the key index for evaluation for quantitative analysis.

[0033] Preferably, the specific process of using the change amount of the logical depth on the critical path before and after optimization as the key index for evaluation for quantitative analysis is as follows:

[0034] Divide the logic cone G into k non-overlapping sub-graphs G1, G2, ..., Gk; where k is a positive integer;

[0035] Select the sub-graphs with a larger logical depth for logical reconstruction respectively, and use the reconstructed sub-graphs to replace the original sub-graphs in the logic cone G. Finally, the logic cone G is optimized into a new logic cone G′, and the logical depth corresponding to G′ is h(G′);

[0036] Evaluate the optimization effect by calculating the change amount of the graph logical depth before and after optimization; define the evaluation function as: ∆h = h(G) - h(G′);

[0037] When it satisfies that the logical depth of the optimized logic cone is less than the logical depth of the logic cone before optimization, that is, the ∆h value is greater than 0, the evaluation passes.

[0038] The present invention also discloses a critical path delay optimization system based on logic depth-driven graph partitioning, including a memory and a processor connected to each other. A computer program is stored on the memory, and when the computer program is run by the processor, it executes the steps of the above-mentioned method.

[0039] Compared with the prior art, the advantages of the present invention are:

[0040] The critical path delay optimization method based on logic depth-driven graph partitioning of the present invention deeply analyzes the characteristics of the critical path in the microprocessor logic netlist. Through the logic depth-driven graph partitioning method, sub-circuit structures with optimization value are accurately identified from the logic netlist of the IC semi-custom design process, and the identified complex circuit structures are optimized using logic reconstruction to obtain reconstructed units with better delay. By performing full-custom design on a small number of reconstructed units and replacing them in the logic netlist, the optimization of the logic depth and the number of cells on the critical path is achieved on the premise of consistent logic functions, and the critical path delay is effectively reduced while ensuring consistent logic functions, with the expectation of improving the performance of the microprocessor. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a flowchart of the critical path delay optimization method based on logic depth-driven graph partitioning provided by an embodiment of the present invention.

[0042] Figure 2 It is a flowchart of the critical path graph partitioning method driven by logic depth provided by an embodiment of the present invention.

[0043] Figure 3 It is a flowchart of the sub-graph logic reconstruction method provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The present invention will be further described below in conjunction with the accompanying drawings of the specification and specific embodiments.

[0045] As Figure 1 shown, the critical path delay optimization method based on logic depth-driven graph partitioning of an embodiment of the present invention includes the steps:

[0046] 1) Obtain the logic netlist, the given standard cell library, and the timing constraint conditions in the integrated circuit (IC) semi-custom design process;

[0047] 2) Analyze the characteristics of the critical path in the logic netlist, and construct a critical path logic cone based on the given standard cell library and timing constraint conditions;

[0048] 3) Screen out sub-circuits with optimization value from the critical path logic cone through the logic depth-driven graph partitioning method;

[0049] 4) Simplify the sub-circuits with optimization value through the sub-graph logic reconstruction method to obtain reconstructed logic units with better delay;

[0050] 5) Screen out the reconstructed logic units corresponding to sub-graphs with a reuse rate greater than a preset value and replace them in the initial logic netlist to obtain an optimized logic netlist;

[0051] 6) Perform equivalence checking on the initial logic netlist and the optimized logic netlist; if the equivalence checking is passed, proceed to the next step.

[0052] 7) Perform delay evaluation on the logic netlist that has passed the equivalence checking. If the evaluation is passed, the final reconstructed logic unit and the optimized logic netlist are obtained.

[0053] In step 1), the IC semi-custom design completes the circuit design by using pre-designed standard cells. In the logic synthesis step, an optimized gate-level netlist is generated according to the optimized Boolean description, the characteristics of the standard cell library, and the set constraint conditions. Under the constraint of a relatively high main frequency, there are still a large number of critical paths with timing violations in the logic netlist obtained under the harsh constraints of commercial EDA tools. Obtain the logic netlist in the IC semi-custom design process, as well as the corresponding commercial standard cell library and the given timing constraint conditions for subsequent optimization processing.

[0054] In step 2), extract the necessary design information from the library file corresponding to the selected standard cell library and the netlist file corresponding to the selected design respectively; analyze the characteristics of the critical paths in the logic netlist, abstract the critical paths and their associated parts in the netlist into critical path logic cones, without retaining the module hierarchy and the internal connections of the modules, only retaining the standard cells and their port information, as well as the logical relationships between them. In the constructed logic cone, each standard cell is regarded as a node; introduce a weight system, use the logical depth of the standard cell corresponding to the node as its weight, and store it in the form of a weighted directed acyclic graph (Directed Acyclic Graph, DAG); finally, visually display the constructed critical path logic cone.

[0055] In step 3), accurately analyze the critical path characteristics of the circuit based on the critical path logic cone, identify sub-circuits with optimization value by graph partitioning of the critical path, and model the critical path partitioning problem as a directed graph partitioning selection problem. Among them, the critical path graph partitioning method driven by logical depth uses logical depth as the partitioning basis, and iteratively selects the partitioning with the largest logical depth increment to accurately identify sub-circuits with greater optimization value.

[0056] In step 4), first identify and record the port information of the sub-circuits with optimization value to ensure that the port configuration of the reconstructed unit is consistent with the original circuit; then use two-level logic simplification to reduce the logical level of the circuit; at the same time, considering the complementary characteristics of CMOS circuits, extract the inverted logic upward; finally, obtain the reconstructed logic unit with optimized delay.

[0057] In step 5), the sub - circuits identified as having optimization value are matched in the netlist, and the reuse situation of the sub - circuits in the netlist is quantified. The reuse rate of each sub - graph is statistically calculated in sequence and sorted accordingly. According to different design requirements, some sub - graphs with higher reuse rates are selected, and the corresponding reconstructed logic units are obtained. The corresponding structures in the netlist are replaced with the reconstructed logic units.

[0058] In step 6), equivalence checking is performed using the initial logic netlist and the logic netlist after implementing logic optimization. Taking the initial logic netlist as the reference model, the equivalence of the logic before and after reconstruction is verified. When the logic functions corresponding to the two netlists are equivalent, it proves that the optimized design is equivalent to the original design, verifying the accuracy of the matching and replacement operations. If the logic functions corresponding to the two netlists are not equivalent, it proves that the matching and replacement operations are incorrect, and replacement and verification need to be performed again until the equivalence checking is passed.

[0059] In step 7), the standard cell delay and the wire delay between cells are not fixed values but an interval, making it difficult to directly perform quantitative analysis. The present invention uses the logic depth on the critical path as the key indicator for measuring the circuit delay, establishes an evaluation model, and evaluates the delay optimization effect by comparing the changes in the logic depth on the critical path before and after optimization.

[0060] The critical - path delay optimization method based on logic - depth - driven graph partitioning of the present invention first constructs a critical - path logic cone based on the logic netlist in the IC semi - custom design process, combined with a given standard cell library and timing constraint conditions, transforming the originally complex netlist critical - path problem into a more intuitive and easily - handled directed - graph problem. Secondly, the critical - path logic cone is graph - partitioned directly using the logic depth as the partitioning driver, modeling the critical - path partitioning problem as a directed - graph partitioning selection problem to identify sub - circuits with greater optimization value. Then, sub - graph logic reconstruction is performed on the identified complex logic circuits, and two - level logic simplification and inversion logic are used to reduce the logic depth of the design upward to obtain reconstructed units for delay optimization. A small number of general reconstructed units are screened and matched and replaced in the design netlist to achieve critical - path delay optimization. Finally, a model is established to evaluate the optimization effect, using the logic depth on the critical path to measure the circuit delay, and the change amount of the logic depth on the critical path before and after optimization as the key indicator for evaluating the delay optimization effect.

[0061] The present invention deeply analyzes the characteristics of critical paths in the microprocessor logic netlist. Through a logic-depth-driven graph partitioning method, it accurately identifies sub-circuit structures with optimization value from the logic netlist in the IC semi-custom design process, and uses logic reconstruction to optimize the identified complex circuit structures to obtain reconstructed units with better delay. By performing full-custom design on a small number of reconstructed units and replacing them into the logic netlist, under the premise of consistent logic functions, it realizes the optimization of logic depth and the number of units on the critical path, and effectively reduces the critical path delay under the premise of ensuring consistent logic functions, with a view to improving the performance of the microprocessor.

[0062] The logic-depth-driven critical path graph partitioning method of the present invention models the critical path partitioning problem as a directed acyclic graph partitioning selection problem. Using logic depth as the partitioning basis, it performs graph partitioning on the critical path logic cone to more accurately identify sub-circuits with greater optimization value, thereby providing options for circuit reconstruction optimization.

[0063] The method of sub-graph logic reconstruction of the present invention simplifies the identified complex logic circuits with optimization value by using two-level logic simplification and upward extraction of inverted logic, reduces the delay by reducing the logic depth and the number of logic gates in the circuit, and finally obtains reconstructed units with optimized delay.

[0064] The present invention is implemented based on the logic netlist in the IC semi-custom design process, giving full play to the respective advantages of semi-custom design and full-custom design, and can be conveniently embedded into the conventional IC design logic synthesis process, and provides an effective means for optimizing the critical path delay at limited cost.

[0065] To better understand the above technical solutions, the above technical solutions will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0066] As Figure 1 shown, the critical path delay optimization method based on logic-depth-driven graph partitioning provided by the embodiment of the present invention includes the following steps:

[0067] 1) Obtain the logic netlist, given standard cell library, and timing constraint conditions in the IC semi-custom design process;

[0068] 2) Deeply analyze the characteristics of critical paths in the logic netlist, combine the given standard cell library and timing constraint conditions, construct a critical path logic cone, and store it in the form of a weighted directed acyclic graph;

[0069] 3) Configure reasonable partitioning constraints, use logic depth as the partitioning driver, and use the graph partitioning method to screen out a batch of sub-circuits with optimization value from the logic cone;

[0070] 4) Reconstruct the logic of each sub-circuit generated by partitioning through a method of two-level logic simplification and extraction of the inversion logic to the outside, and obtain a reconstructed logic unit with better delay by reducing the logic depth of the sub-graph;

[0071] 5) Comprehensively consider various factors such as the reuse rate and the optimization value, select a small number of reconstructed logic units, and replace them into the initial logic netlist;

[0072] 6) Use the initial logic netlist as a reference model to verify the equivalence of the logic before and after reconstruction; if passing the equivalence test, proceed to the next step, otherwise return to step 5);

[0073] 7) Use the delay evaluation model, take the change amount of the logic depth of the critical path before and after optimization as the key index for evaluating the delay optimization effect, evaluate the delay of the reconstructed logic with equivalent logic, and judge whether the optimization target is met. If it is met, obtain a batch of reconstructed logic units with customization value and the optimized logic netlist; if not, jump to execute step 3).

[0074] In practical applications, when necessary, the method of the present invention can be used to perform cyclic optimization on the logic netlist multiple times to minimize the delay of the critical path as much as possible.

[0075] Specifically, step 1) includes: obtaining the logic netlist and the corresponding critical path information in the IC semi-custom design process; obtaining the standard cell library corresponding to the logic netlist, and extracting basic information such as the ports and logic functions of each standard cell; obtaining the given timing constraint conditions.

[0076] In step 2), the specific process of constructing the critical path logic cone based on the logic netlist includes:

[0077] In order to solve the problems of complex netlist information and difficult processing of cross-level modules, enhance the readability and optimization efficiency of the netlist, the present invention uses an abstract processing method to construct a critical path logic cone based on the given logic netlist. By introducing a new data structure, redundant information such as the hierarchical relationship and connection details between modules is ignored, and only the connection relationship between each standard cell is retained while maintaining the accuracy of the circuit design to improve the efficiency of subsequent optimization processing.

[0078] The present invention abstracts the critical path and its associated parts in the netlist into a weighted directed acyclic graph, transforming the originally complex netlist critical path problem into a more intuitive and easily processed directed graph problem, providing a new perspective and method for the optimization of circuit design. Use the weighted directed acyclic graph DAG to represent the circuit structure of the critical path and its associated parts in the netlist, where the vertices represent the instances of the standard cells in the circuit, the directed edges between the vertices represent the connection relationship between the cells, and the direction of the arrow represents the direction of signal transmission between the cells.

[0079] The direction of the critical path logic cone constructed in the present invention is exactly opposite to the signal transmission direction. First, obtain the endpoints of the critical path from the netlist; then, gradually trace back from these endpoints to the output ports of the timing units and stop. In this process, the module hierarchy and the internal connections of the modules are not retained, only the standard cells and their port information, as well as the logical relationships between them are retained. By identifying the driving nodes corresponding to all input ports of the nodes, all the standard cells and their connection relationships between the critical path endpoints and the timing units can be obtained in sequence. In the constructed logic cone, each obtained standard cell is regarded as a node. Among them, the timing unit can be a flip-flop, a static random access memory (SRAM) or an IP in the process library, and its output is defined as the starting point of the logic cone.

[0080] The present invention introduces a weight system in the abstracted logic cone, that is, the logical depth of the node is used as its weight. Each node contains one or more logic gates inside, and the logical depth of the node is the number of logic gates on the longest path from the input to the output of the node, which is the key factor determining the node delay.

[0081] In step 3), partitioning the critical path logic cone under certain constraints to identify subcircuits with optimization value is a dynamic process that requires multiple iterations, and each iteration may involve the reconstruction of different subcircuits. The method of critical path graph partitioning driven by logical depth provided by the present invention initially regards each node in the logic cone as a minimum subgraph, and then gradually merges those pairs of subgraphs that can generate the maximum delay optimization value after merging until the final partitioning result is obtained. In the present invention, the logical depth is used as the key index to evaluate the quality of subgraph partitioning, and the optimal partitioning is iteratively selected, aiming to make the logical depth inside the subgraph as large as possible after partitioning, which means that the optimization value of the subgraph is also greater.

[0082] Specifically, as Figure 2 shown, the content of the critical path logic cone graph partitioning driven by logical depth in step 3) is as follows:

[0083] Initialize the subgraph structure. Regard each node in the logic cone as an independent minimum subgraph, and calculate the logical depth of each subgraph. The logical depth of the subgraph is the logical depth on the longest path from the input port to the output port of the subgraph.

[0084] Search for all pairs of subgraphs that can be merged. In the present invention, considering the influence of the number of ports of a cell on its design cost, physical size, layout and wiring, cell delay and power consumption, etc., based on the empirical data of backend designers, the present invention restricts the number of input ports of the divided subgraphs to no more than 8, and the number of output ports to no more than 2. Therefore, the pairs of subgraphs that can be merged must satisfy: the graph obtained by merging conforms to the port number constraint, and the logical depth of the newly obtained graph after merging is greater than the logical depth of the subgraphs before merging.

[0085] Calculate the logical depth increment IncD (Incremental Depth). In the present invention, the logical depth increment is the difference between the logical depth of the newly obtained graph after merging and the average logical depth of the pair of subgraphs before merging, and the logical depth increment is used as an evaluation index for the merging operation. For each pair of subgraphs, calculate the logical depth increment after merging; the greater the logical depth increment, the greater the merging benefit; if the logical depth increment is less than or equal to 0, it means that this merging method is not advisable.

[0086] Select the best merge. For all pairs of subgraphs that can be merged identified, find the pair of subgraphs with the largest logical depth increment after merging, and merge to obtain a new subgraph. Use the new subgraph to replace the merged subgraphs and update the subgraph structure.

[0087] Iteratively perform the merging operation until there are no more pairs of subgraphs that can be further merged, and obtain the final partitioning result.

[0088] As Figure 3 shown, in step 4), for the sub-circuits with optimization value obtained by graph partitioning, perform logic reconstruction to obtain a reconstructed unit with optimized delay. The specific process is as follows:

[0089] Identify ports: Identify the input ports and output ports of the subgraph to ensure that the port configuration of the reconstructed unit is consistent with the subgraph.

[0090] Logic simplification: Perform two-level logic simplification on each output port of the subgraph. Specifically, the present invention finds the set of minimum prime implicants by identifying and merging implicants, and then obtains the minimum SOP (Sum of Products, a basic concept in Boolean algebra) expression of the circuit to achieve two-level logic simplification to reduce the logic level of the circuit.

[0091] Extract inverted logic: On the basis of two-level logic simplification, extract all inverted logic upward and simplify it. Finally, a new optimized reconstructed logic unit will be obtained.

[0092] In step 5), in practical applications, it is difficult to obtain new reconstructed logic units for all sub-circuits with optimization value through sub-graph logic reconstruction. For higher optimization benefits, the present invention carefully selects a small number of reconstructed logic units with higher optimization benefits, so as to achieve better optimization effects by reconstructing and replacing a small number of units, thereby improving the benefits of delay optimization.

[0093] Specifically, match the sub-circuits with optimization value in the netlist. Specifically, first perform topological sorting on the nodes in the logic cone, arrange them into a linear sequence in logical pre-order, ensuring that the driving unit of each unit is placed before it; then, based on the sorting result, perform matching in turn in the form of a graph, and match the part in the logic cone that has the same graph structure as the sub-graph. Among them, units with the same function are represented by nodes in the same form, the connections between nodes represent the connections between units, arrows indicate the data flow direction, and buffer units are directly represented in the form of connections rather than nodes. Count the reuse rate of each sub-graph in turn, and sort according to this, select some sub-graphs with higher reuse rate to obtain the corresponding reconstructed logic units, and finally use the reconstructed logic units to replace the corresponding structures in the netlist to obtain an optimized logic netlist.

[0094] In step 6), the specific process of equivalence checking is as follows:

[0095] Use the initial logic netlist and the optimized logic netlist to perform equivalence checking. Take the initial logic netlist as the reference model, and verify the equivalence of the logic before and after optimization on the premise that other irrelevant conditions remain the same. When the logical functions corresponding to the two netlists are equivalent, it proves that the optimized logic cone is equivalent to the original logic cone.

[0096] In step 7), combine an evaluation model to evaluate the critical path delay optimization effect. The specific evaluation method is as follows:

[0097] The delay between basic units is not a definite value, but an interval. Performance evaluation is completed using the evaluation model, and the change in logical depth on the critical path before and after optimization is used as the key evaluation index for quantitative analysis.

[0098] Specifically, divide the logic cone G into k (k is a positive integer) non-overlapping sub-graphs G1, G2,..., Gk;

[0099] Select the sub-graphs with larger logical depth for logical reconstruction respectively, and use the reconstructed sub-graphs to replace the original sub-graphs in the logic cone G. Finally, the logic cone G is optimized into a new logic cone G′, and the logical depth corresponding to G′ is h(G′).

[0100] The optimization effect is evaluated by calculating the change in the logical depth of the graph before and after optimization; the evaluation function is defined as: ∆h = h(G) - h(G’).

[0101] In the present invention, the critical path delay optimization method for logical depth-driven graph partitioning provided satisfies that the logical depth of the optimized logic cone is less than that of the logic cone before optimization, that is, the ∆h value is greater than 0.

[0102] The present invention re-optimizes the circuit designed by the EDA tool, can effectively reduce the logical depth on the critical path, and thus provides an effective means for optimizing the critical path delay at a limited cost. At the same time, it provides a choice for designing new logic units. By fully customizing the design of the logic reconstruction unit, a batch of new units can be obtained, and finally a new fully customized unit library is generated.

[0103] The present invention also provides a critical path delay optimization system based on logical depth-driven graph partitioning, including a memory and a processor connected to each other. A computer program is stored on the memory, and when the computer program is run by the processor, it executes the steps of the method described above. The optimization system of the present invention corresponds to the above optimization method and has the same advantages as those of the above optimization method.

[0104] The present invention can also implement all or part of the processes in the above method embodiments through hardware related to computer program instructions. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium includes: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. The memory is used to store the computer program and / or module, and the processor realizes various functions by running or executing the computer program and / or module stored in the memory and calling the data stored in the memory. The memory can include high-speed random access memory, and can also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices, etc.

[0105] The above has introduced in detail the method for optimizing the critical path delay of the logic depth-driven graph partitioning provided by the present invention. The present invention expounds the principle and implementation manner of the present invention through specific examples. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.

Claims

1. A critical path delay optimization method based on logic depth drive graph partitioning, characterized in that: Includes steps: 1) Obtain the logic netlist, given standard cell library and timing constraints in the integrated circuit semi-custom design process; 2) Analyze the characteristics of the critical path in the logic netlist and construct the critical path logic cone based on the given standard cell library and timing constraints; 3) Screen out sub-circuits with optimization value from the critical path logic cone through the logic depth driven graph partitioning method; 4) Simplify the sub-circuit with optimization value through the sub-graph logic reconstruction method to obtain the reconstructed logic unit with better delay; 5) Filter the reconstructed logic units corresponding to the subgraphs whose reuse rate is greater than the preset value and replace them into the initial logic netlist to obtain the optimized logic netlist; 6) Perform equivalence check on the initial logic netlist and the optimized logic netlist; If the equivalence test is passed, proceed to the next step; 7) Perform delay evaluation on the logic netlist that passes the equivalence test. If the evaluation passes, the final reconstructed logic unit and optimized logic netlist are obtained; The specific steps of step 3) are: Initialize the subgraph structure: treat each node in the logic cone as an independent minimum subgraph and calculate the logic depth of each subgraph; the logic depth of a subgraph is the logic depth of the longest path from the input port to the output port of the subgraph; Search for all possible subgraph pairs that can be merged; the subgraph pairs that can be merged must satisfy the following conditions: the merged graph meets the port number constraint, and the logical depth of the merged new graph is greater than the logical depth of the subgraph before merging. Calculate the logical depth increment: For each pair of subgraphs, calculate the logical depth increment after merging; The larger the logical depth increment, the greater the merging benefit; if the logical depth increment is less than or equal to 0, it means that this merging method is not desirable; the logical depth increment is the difference between the logical depth of the new graph after merging and the average logical depth of the subgraph before merging; Select the best merge: For all the identified subgraph pairs that may be merged, find the subgraph pair with the largest logical depth increment after merging, and merge them to obtain a new subgraph; Use the new subgraph to replace the merged subgraph and update the subgraph structure; The merging operation is iterated until there are no more merged subgraph pairs, and the final partitioning result is obtained.

2. The critical path delay optimization method based on logic depth drive graph partitioning according to claim 1 is characterized in that: In step 2), the direction of constructing the critical path logic cone is exactly opposite to the signal transmission direction; first, the end point of the critical path is obtained from the logic netlist; then, the end point is gradually traced back to the output port of the timing unit and stopped. In this process, the module hierarchy and the internal connections of the module are not retained, only the standard unit and its port information, as well as the logical relationship between them are retained; by identifying the driving nodes corresponding to all the input ports of the node, all the standard units from the end point of the critical path to the timing unit and their connection relationships can be obtained in turn; in the constructed logic cone, each obtained standard unit is regarded as a node; the timing unit is a trigger, static random access memory or IP in the process library, and its output is defined as the starting point of the logic cone.

3. The critical path delay optimization method based on logic depth drive graph partitioning according to claim 2 is characterized in that: In step 2), in the critical path logic cone, the logic depth of the standard cell corresponding to the node is used as its weight, stored in the form of a weighted directed acyclic graph, and visualized; each node contains one or more logic gates, and the logic depth of the node is the number of logic gates on the longest path from the input to the output of the node.

4. The critical path delay optimization method based on logic depth drive graph partitioning according to claim 1 is characterized in that: The specific steps of step 4) are: Identify ports: Identify the input ports and output ports of the subgraph to ensure that the port configuration of the reconstruction unit is consistent with the subgraph; Logic simplification: Implement two-level logic simplification for each output port of the subgraph; specifically, identify and merge the implied terms to find the minimum set of prime implied terms, and then obtain the minimum SOP expression of the circuit, implement two-level logic simplification to reduce the logic level of the circuit; Negation logic extraction: Based on the two-level logic simplification, all the negation logic is extracted and simplified upward, and finally a new optimized reconstructed logic unit is obtained.

5. The critical path delay optimization method based on logic depth drive graph partitioning according to claim 4 is characterized in that: In step 5), the specific steps of screening the reconstruction logic units corresponding to the subgraphs whose reuse rate is greater than the preset value are: First, the nodes in the logic cone are topologically sorted into a linear sequence of logical precedence; Then, based on the sorting results, matching is performed in sequence in the form of a graph, and the parts with the same graph structure as the subgraph are matched in the logic cone; units with the same functions are represented by nodes in the same form, and the lines between the nodes represent the connection between the units, and the arrows are used to indicate the data flow direction, while the buffer unit is directly represented in the form of a line instead of a node; The reuse rate of each sub-graph is counted and sorted in turn, and the sub-graph with a reuse rate greater than a preset value is selected to obtain the corresponding reconstruction logic unit.

6. The critical path delay optimization method based on logic depth drive graph partitioning according to claim 1, 2 or 3, characterized in that: The specific process of the equivalence test in step 6) is: Taking the initial logic netlist as the reference model, and under the premise that other irrelevant conditions remain the same, verify the equivalence of the initial logic netlist and the optimized logic netlist. When the logical functions corresponding to the two netlists are equivalent, it is proved that the optimized logic cone is equivalent to the original logic cone, and the equivalence test passes.

7. The critical path delay optimization method based on logic depth drive graph partitioning according to claim 1, 2 or 3, characterized in that: In step 7), when performing delay evaluation on the logic netlist that has passed the equivalence check, the change in logic depth on the critical path before and after optimization is used as a key indicator for evaluation and quantitative analysis is performed.

8. The critical path delay optimization method based on logic depth drive graph partitioning according to claim 7 is characterized in that: The specific process of quantitative analysis using the change in logic depth on the critical path before and after optimization as the key indicator for evaluation is as follows: Divide the logical cone G into k non-overlapping subgraphs G1, G2, . . . , Gk; where k is a positive integer; Select subgraphs with larger logic depths for logic reconstruction, and use the reconstructed subgraphs to replace the original subgraphs in the logic cone G. Finally, the logic cone G is optimized to a new logic cone G′, and the logic depth corresponding to G′ is h(G′); The optimization effect is evaluated by calculating the change in the graph logic depth before and after optimization; the evaluation function is defined as: ∆h=h(G)-h(G'); When the logic depth of the logic cone after optimization is less than the logic depth of the logic cone before optimization, that is, the ∆h value is greater than 0, the evaluation passes.

9. A critical path delay optimization system based on logic depth drive graph partitioning, comprising a memory and a processor connected to each other, wherein a computer program is stored on the memory, characterized in that: When the computer program is executed by a processor, the computer program performs the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Hybrid logic comprehensive optimization method and device of circuit and electronic equipment

    CN117521567A

  • Critical path delay optimization method based on logic netlist

    CN118171609A