GPU (Graphics Processing Unit) parallel acceleration global wiring method oriented to time sequence and congestion collaborative optimization

Through the Kruskal algorithm and Elmore delay model combined with GPU parallel core acceleration, the global wiring method of integrated circuits is optimized, the coordination problem between timing and congestion is solved, the wiring quality and timing performance are improved, and the requirements of high-performance circuit design are met.

CN120493854APending Publication Date: 2025-08-15SOUTHEAST UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510650400.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

When dealing with integrated circuit timing and congestion problems, existing global routing methods are difficult to effectively balance timing performance and congestion in the early stages, resulting in high subsequent optimization costs.

Method used

The Kruskal algorithm is used to combine and collect data structures to divide the network, calculate timing weights based on the Elmore delay model, accelerate hybrid mode wiring using GPU parallel cores, and optimize network topology through delay awareness and congestion-driven strategies.

Benefits of technology

In the ultra-large-scale integrated circuit design, coordinated optimization of timing and congestion is achieved, wiring quality and timing performance are improved, subsequent optimization burden is reduced, and design efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493854A_ABST
    Figure CN120493854A_ABST
Patent Text Reader

Abstract

The invention discloses a time sequence and congestion collaborative optimization-oriented GPU (Graphics Processing Unit) parallel acceleration global wiring method, which comprises the following steps of: according to a given netlist, dividing a super-large network by adopting a Kruskal algorithm in combination with a lookup set, and constructing a time sequence propagation path of the super-large network; performing network decomposition based on the timing margin estimation of the pins; calculating the time sequence weight of the two-pin network by adopting a time sequence weight calculation method based on an Elmore delay model; performing path cost calculation by adopting a cost function which comprehensively considers the time sequence and the congestion cost; executing two-level GPU parallel kernel acceleration mode wiring; optimizing the network topology by adopting a delay-aware pin connection improvement technology; the invention relates to a congestion-driven GPU (Graphics Processing Unit) accelerated routing and routing strategy for executing non-critical networks. According to the method, a high-quality wiring result with balanced time sequence congestion can be quickly obtained, the time sequence performance is effectively improved, and the requirement of a current super-large-scale high-performance circuit design wiring stage can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of FPGA physical design automation technology, and in particular to a GPU parallel acceleration global routing method for timing and congestion collaborative optimization. Background Art

[0002] Global routing plays a crucial role in the integrated circuit (IC) physical design process, significantly impacting timing performance and overall design quality. As circuit complexity continues to grow, ensuring timing closure becomes increasingly challenging, making global routing a crucial stage in early timing optimization. During the global routing phase, coarse-grained routing paths for signal nets are planned across multiple metal layers, laying the foundation for subsequent detailed routing. Critical nets have the greatest impact on worst negative slack (WNS) and total negative slack (TNS), so the global routing phase provides an opportunity to optimize timing-critical paths. Timing margins can be improved by reducing wire lengths, optimizing tree topologies, and using thicker metal layers. Effectively addressing these factors can significantly improve timing performance. A robust global routing solution not only improves routability but also lays a solid foundation for achieving timing closure in advanced IC designs.

[0003] Classic global routing methodologies primarily focus on wire length and congestion. While these congestion-driven approaches are widely adopted due to their effectiveness in improving routing feasibility, they often ignore timing constraints. Consequently, these methods can lead to timing violations or performance degradation in the subsequent physical design process, necessitating costly post-routing timing optimization steps. While wire length-driven global routing can improve timing to a certain extent, there are fundamental differences between wire length optimization and timing optimization. This highlights the importance of developing global routing techniques that directly address timing while also taking congestion into account.

[0004] Timing-driven global routing has the potential to mitigate critical path violations early in physical design and reduce reliance on later timing fixes. As an early evaluation stage, global routing requires repeated use, placing high demands on efficiency. As integrated circuits continue to increase in complexity, scalability of routing algorithms has become a key challenge.

[0005] Routing is a crucial and challenging optimization step in physical design, directly impacting the design's timing performance, power consumption, and manufacturability. Routing works closely with layout, fulfilling the critical task of achieving efficient interconnection within limited resources. With the continued expansion of integrated circuits and the evolution of process nodes, routing issues are characterized by high complexity, high congestion, and the coexistence of multiple constraints (such as timing, power consumption, and signal integrity). This places higher demands on the optimization capabilities and scalability of routing algorithms. Summary of the Invention

[0006] The present invention provides a GPU parallel accelerated global routing method for the coordinated optimization of timing and congestion to solve the problem of poor performance of global routing results in the prior art. This method can better balance timing performance and congestion in advanced integrated circuit design and improve overall performance.

[0007] An embodiment of the present invention provides a GPU parallel accelerated global routing method for timing and congestion coordinated optimization, comprising the following steps:

[0008] Step S1: Based on a given netlist, the Kruskal algorithm combined with the union-find data structure is used to partition the super-large network and construct the time-series propagation path of the super-large network;

[0009] Step S2, performing network decomposition based on pin-based timing slack estimation according to the timing propagation path, dividing the network into a critical network and a non-critical network, and decomposing spanning trees obtained from the critical network and the non-critical network into 2-pin networks;

[0010] Step S3, for the critical network, using a timing weight calculation method based on the Elmore delay model to calculate the timing weight of the 2-pin network, and determining the timing priority critical connection and the regular critical connection by sorting the timing weights;

[0011] Step S4, calculating the path cost of the timing-priority critical connection and the conventional critical connection using a cost function that comprehensively considers timing and congestion cost;

[0012] Step S5, executing two-level GPU parallel kernel accelerated mixed mode routing to obtain the original network topology;

[0013] Step S6, optimizing the original network topology using a delay-aware pin connection improvement technique;

[0014] Step S7: performing congestion-driven GPU accelerated routing and routing strategy on the non-critical network.

[0015] Optionally, in one embodiment of the present invention, in step S1, for an extremely large network with a number of pins greater than a preset number, a minimum spanning tree is first generated using the Kruskal algorithm in combination with a union-find data structure. During the minimum spanning tree construction process, when the size of a connected component exceeds a predefined threshold, the connected component is split into sub-networks.

[0016] After the partitioning is completed, each original network is divided into multiple sub-networks. The remaining minimum spanning tree edges connecting different sub-networks are then used to construct an adjacency table to represent the connection relationship between the sub-networks. Starting from the sub-network containing the source pin of the network, a breadth-first search traversal is performed. During the traversal process, a directed acyclic graph is constructed to record the path dependency relationship between the sub-networks.

[0017] After each sub-network is decomposed independently, the complete temporal propagation path across the sub-network is reconstructed using a directed acyclic graph.

[0018] Optionally, in one embodiment of the present invention, in step S2, the network is divided into a critical network and a non-critical network according to the estimated pin timing margin information. If a network contains any pin with a negative margin, it is considered a critical network; otherwise, it is considered a non-critical network.

[0019] For critical networks, the SALT algorithm is used to generate shallow lightweight trees to obtain shorter source-to-sink paths; for non-critical networks, the FLUTE algorithm is used to minimize bus length;

[0020] The resulting spanning trees are all decomposed into 2-pin networks, which are called a connection of this network.

[0021] Optionally, in one embodiment of the present invention, in step S3, the timing weight of the connection is calculated, and the specific calculation formula of the timing weight of the connection ab is:

[0022] Weight ab =-l ab LoadCap b SlackSum b (1)

[0023] Among them, l ab is the Manhattan distance between pin a and pin b, LoadCap b is the load capacitance of pin b, calculated similarly to the Elmore model:

[0024]

[0025] Among them, child is the collection of child nodes of the current node, l g is the Manhattan distance between the current node and the child node g, C0 is the unit capacitance value, here we take the median of the unit capacitance values of all metal layers, LoadCap g is the load capacitance of child node g, PinCap node When the node is a pin, it is the inherent capacitance of the pin; when it is not a pin, it is 0;

[0026] SlackSumb is the cumulative sum of the negative timing slack of pin b:

[0027]

[0028] Among them, SlackSum c Slack is the cumulative sum of the negative timing slack of child node g. node When the node is a pin, it is the timing margin of the pin; when it is not a pin, it is 0;

[0029] After calculating the timing weights of all connections, connections with non-zero timing weights are defined as critical connections, and connections with timing weights of 0 are defined as non-critical connections; the non-zero timing weights are further sorted in descending order, and critical connections are divided into two categories: timing-priority critical connections and regular critical connections; the classification is based on a predefined ratio parameter α (0<α<1), and the α×Nth value in the sorted timing weight vector is selected as the threshold, where N is the total number of connections; connections with timing weights greater than the threshold are identified as timing-priority critical connections, while the rest are classified as regular critical connections.

[0030] Optionally, in one embodiment of the present invention, for critical connections and non-critical connections, when routing, the critical connections of the network are first extracted and routed first, and the non-critical connections of the network are routed after the routing of all the critical connections of the network is completed.

[0031] Optionally, in one embodiment of the present invention, in step S4, pattern routing is performed using a differentiated cost function for the timing-priority critical connections and regular critical connections classified based on timing weights in the critical network, as well as the non-critical network. In the pattern routing, the cost of routing a line segment pq on the Lth metal layer is:

[0032] Cost = w c Cost c +w t Cost t (4)

[0033] Among them, w c and w t are congestion weight and timing weight, Cost c and Cost t are congestion cost and timing cost respectively, w t Defined as:

[0034]

[0035] Among them, when the current connection is a key connection, w t When the current connection is a non-critical connection, wt is 0;

[0036] Cost c Adjust based on whether the connection is timing-prioritized. For timing-prioritized critical connections:

[0037]

[0038] For regular critical and non-critical connections:

[0039]

[0040] Among them, w R and are the pre-set weights and the unit resistance of the Lth metal layer, Cost e is the congestion cost of a mesh edge, E wire is the set of mesh edges that the line segment passes through; for a mesh edge with capacity c and demand d, the congestion cost is calculated as:

[0041] Cost e =e s(d+1-c) -e s(d-c) (8)

[0042]

[0043] The timing cost of the line segment in formula (4) is t The calculation method is:

[0044] Cost t =RES Via (CAP wire +LoadCap q )+RES wire (0.5·CAP wire +LoadCap q ) (10)

[0045] Among them, RES Via Represents the total resistance of all vias connected to the upstream segment, LoadCap node is the cumulative capacitance of the node node, which is consistent with the node load capacitance used in the timing weight calculation; RES wire and CAP wire is the resistance and capacitance of the line segment, RES Via RES wire and CAP wire Expressed as:

[0046]

[0047] Among them, vias is the set of vias that the current segment passes through to connect to the upstream segment. represents the resistance of a single via between layer l and layer l+1, wl seg is the length of the line segment, is the unit resistance of the Lth layer, is the unit capacitance of the Lth layer.

[0048] Optionally, in one embodiment of the present invention, in step S5, the two-level GPU parallel core accelerated mixed mode routing specifically includes the following steps:

[0049] (1) Batch generation based on a novel overlap check: vertical and horizontal routing resources are distinguished during the check. L-shaped routing is used for connections with a length of less than 10 grid edges, while sparse Z-shaped routing is used otherwise. The semi-perimeter of the bounding box of the connection is divided into ten equal-length segments. Each segment point is used as a candidate position for the first inflection point of the Z-shape, thereby uniquely determining a two-dimensional Z-shaped candidate path.

[0050] (2) Flatten all candidate 2D connections: Flatten all candidate connections with L-shaped routing and Z-shaped routing;

[0051] (3) Parallel routing of candidate connections within a batch: A two-level parallel mechanism is used. At the level between connections, connections without resource conflicts are processed in parallel. At the same time, at the level within each connection, up to 10 candidate Z-shaped two-dimensional paths for each connection are parallelized. These two levels are combined to utilize the large number of threads on the GPU to perform routing and obtain the minimum routing cost for each candidate connection.

[0052] (4) Using GPU multithreading to find the best candidate connection for each connection and update the global grid;

[0053] (5) Repeat steps (3) and (4) until all batches are processed;

[0054] (6) Construct the original network topology: After completing the wiring of all connections, assemble the connection paths into a complete network wiring. For each network, the through-holes at the connection nodes are represented by an interval. After determining a connection path, update the intervals of the nodes at both ends of the connection to include the layer where the path is located. When all the connection wiring in the network is completed, a complete connected path is formed. Update the interval at each pin position to ensure that all pins in the network are correctly connected.

[0055] Optionally, in one embodiment of the present invention, in step S6, the delay ratio of the pin is calculated:

[0056]

[0057] Among them, delay actual is the actual delay of the pin, delay min is the delay assuming a direct shortest path connection from the source pin to this pin;

[0058] If the delay ratio is greater than the preset value, the connection of this pin is adjusted; the adjustment uses the A* algorithm to reconnect to the source pin. The reconnection process starts from this pin and uses breadth-first search to explore the wiring space. In each step of expansion, the G-cell with the lowest cost is expanded outward. The cost of each G-cell consists of two parts: the wiring cost cost p (u,s) and the guided cost cost g (u,t):

[0059]

[0060] cost g (u,t)=w g *dist(u,t) (16)

[0061] in, is the cost of G-cell u, u is traversed G-cell means u is the G-cell occupied by the network, u is the current G-cell in the expansion, s is the pin, t is the source pin of the network, cost g (u, t) represents the weighted two-dimensional distance to the target point, w g To guide the cost weight, dist(u,t) is the two-dimensional distance between u and t, cost p (u,s) represents the routing cost from G-cell u to this pin s, including the cost of grid edges and vias; the cost is calculated as a weighted combination of timing cost and congestion cost.

[0062] Optionally, in one embodiment of the present invention, in step S7, for the non-critical network, only the congestion cost is considered during routing, and congestion-driven GPU-accelerated routing is performed: first, a directed acyclic graph is constructed, and an L-shaped pattern is used for initial routing. For networks with overflow, the second stage is entered to create and move Steiner points to reorganize existing paths, thereby generating alternative paths.

[0063] The GPU-accelerated global routing method for co-optimizing timing and congestion, implemented in this embodiment of the present invention, can effectively improve routing quality in very large-scale integrated circuit (VLSI) designs, achieving better timing performance at the same congestion level and reducing the burden of subsequent optimization. This method can effectively improve routability and overall timing metrics, particularly in large-scale designs with stringent timing requirements, thereby enhancing chip performance and design closure efficiency.

[0064] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0066] Figure 1 A flowchart of a GPU parallel accelerated global routing method for timing and congestion collaborative optimization provided in accordance with an embodiment of the present invention;

[0067] Figure 2 This is a two-level GPU accelerated parallel algorithm according to an embodiment of the present invention. DETAILED DESCRIPTION

[0068] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.

[0069] The present invention provides a GPU parallel accelerated global routing method for timing and congestion collaborative optimization. The method first performs preprocessing to improve scalability and timing sensitivity, then adopts timing and congestion driven GPU accelerated hybrid three-dimensional routing based on Elmore delay weights for key networks, combines hybrid cost function optimization path, further accelerates the routing process through GPU parallel core, and finally performs delay-aware optimization on pin connections, and completes routing for non-critical networks using congestion driven GPU accelerated routing engine. Figure 1 As shown, the GPU parallel accelerated global routing method for timing and congestion coordinated optimization includes the following steps:

[0070] Step S1: Based on a given netlist, the Kruskal algorithm combined with the union-find data structure is used to partition the super-large network, and the time sequence propagation path of the super-large network is constructed.

[0071] Optionally, in one embodiment of the present invention, in step S1, for an extremely large network with a pin count greater than a preset number, such as a reset network with more than 10,000 pins, a minimum spanning tree (MST) is first generated using the Kruskal algorithm combined with a union-find data structure. During the minimum spanning tree construction process, when the size of a connected component exceeds a predefined threshold, the component is split to form a subnet, thereby ensuring that the size of each subnet remains within a manageable range.

[0072] After the partitioning is completed, each original network is divided into multiple sub-networks. The remaining minimum spanning tree edges connecting different sub-networks are then used to construct an adjacency table to represent the connection relationship between the sub-networks. Starting from the sub-network containing the source pin of the network, a breadth-first search traversal is performed. During the traversal process, a directed acyclic graph (DAG) is constructed to record the path dependency relationship between the sub-networks;

[0073] After each sub-network is decomposed independently, the complete timing propagation path across the sub-network is reconstructed using a directed acyclic graph, thereby achieving accurate timing analysis of large critical networks.

[0074] Step S2: performing network decomposition for pin-based timing margin estimation according to the timing propagation path, dividing the network into a critical network and a non-critical network, and decomposing the spanning trees obtained from the critical network and the non-critical network into 2-pin networks.

[0075] Optionally, in one embodiment of the present invention, in step S2, the network is divided into a critical network and a non-critical network according to the estimated pin timing margin information. If a network contains any pin with a negative margin, it is considered a critical network; otherwise, it is considered a non-critical network.

[0076] For critical networks, the SALT algorithm is used to generate shallow lightweight trees (SLT) to obtain shorter source-to-sink paths; for non-critical networks, the FLUTE algorithm is used to minimize bus length;

[0077] The resulting spanning trees are all decomposed into 2-pin networks, which are called a connection of this network.

[0078] Step S3: For the critical network, the timing weight calculation method based on the Elmore delay model is used to calculate the timing weight of the 2-pin network, and the timing priority critical connection and the regular critical connection are determined by sorting the timing weights.

[0079] Optionally, in one embodiment of the present invention, in step S3, the timing weight of the connection is calculated, and the connection The specific calculation formula for timing weight is:

[0080] Weightab =-l ab LoadCap b SlackSum b (1)

[0081] Among them, l ab is the Manhattan distance between pin a and pin b, LoadCap b is the load capacitance of pin b, calculated similarly to the Elmore model:

[0082]

[0083] Among them, child is the collection of child nodes of the current node, l g is the Manhattan distance between the current node and the child node g, C0 is the unit capacitance value, here we take the median of the unit capacitance values of all metal layers, LoadCap g is the load capacitance of child node g, PinCap node When the node is a pin, it is the inherent capacitance of the pin; when it is not a pin, it is 0;

[0084] SlackSum b is the cumulative sum of the negative timing slack of pin b:

[0085]

[0086] Among them, SlackSum c Slack is the cumulative sum of the negative timing slack of child node g. node When the node is a pin, it is the timing margin of the pin; when it is not a pin, it is 0;

[0087] After calculating the timing weights of all connections, connections with non-zero timing weights are defined as critical connections, while connections with a timing weight of 0 are defined as non-critical connections. The non-zero timing weights are further sorted in descending order, dividing critical connections into two categories: timing-first critical connections and regular critical connections. This classification is based on a predefined ratio parameter α (0 < α < 1), selecting the α×Nth value in the sorted timing weight vector as the threshold, where N is the total number of connections. Connections with timing weights greater than the threshold are identified as timing-first critical connections, while the rest are classified as regular critical connections. For timing-first connections, the focus is on minimizing latency and maximizing timing benefits.

[0088] For critical connections and non-critical connections, when wiring, the critical connections of the network are extracted first, and the critical connections are wired first. The non-critical connections of the network are wired after the wiring of all the critical connections of the network is completed.

[0089] Step S4 , calculating the path costs of the timing-priority critical connections and the conventional critical connections using a cost function that comprehensively considers timing and congestion costs.

[0090] Optionally, in one embodiment of the present invention, in step S4, pattern routing is performed using a differentiated cost function for the timing-priority critical connections and regular critical connections classified based on timing weights in the critical network, as well as the non-critical network. In the pattern routing, the cost of routing a line segment pq on the Lth metal layer is:

[0091] Cost = w c Cost c +w t Cost t (4)

[0092] Among them, w c and w t are congestion weight and timing weight, Cost c and Cost t are congestion cost and timing cost respectively, w t Defined as:

[0093]

[0094] Among them, when the current connection is a key connection, w t When the current connection is a non-critical connection, w t is 0;

[0095] Cost c Adjust based on whether the connection is timing-prioritized. For timing-prioritized critical connections:

[0096]

[0097] For regular critical and non-critical connections:

[0098]

[0099] Among them, w R and are the pre-set weights and the unit resistance of the Lth metal layer, Cost e is the congestion cost of a mesh edge, E wire is the set of mesh edges that the line segment passes through; for a mesh edge with capacity c and demand d, the congestion cost is calculated as:

[0100] Cost e =e s(d+1-c) -e s(d-c) (8)

[0101]

[0102] The timing cost of the line segment in formula (4) is t The calculation method is:

[0103] Cost t =RES Via (CAP wire +LoadCap q )+RES wire (0.5·CAP wire +LoadCap q ) (10)

[0104] Among them, RES Via Represents the total resistance of all vias connected to the upstream segment, LoadCap node is the cumulative capacitance of the node node, which is consistent with the node load capacitance used in the timing weight calculation; RES wire and CAP wire is the resistance and capacitance of the line segment, RES Via RES wire and CAP wire Expressed as:

[0105]

[0106] Among them, vias is the set of vias that the current segment passes through to connect to the upstream segment. represents the resistance of a single via between layer l and layer l+1, wl seg is the length of the line segment, is the unit resistance of the Lth layer, is the unit capacitance of the Lth layer.

[0107] Step S5: Execute two-level GPU parallel kernel accelerated mixed mode routing to obtain the original network topology.

[0108] Optionally, in one embodiment of the present invention, in step S5, the two-level GPU parallel core accelerated mixed mode routing specifically includes the following steps:

[0109] (1) Batch generation based on a novel overlap check: vertical and horizontal routing resources are distinguished during the check. L-shaped routing is used for connections with a length of less than 10 grid edges, while sparse Z-shaped routing is used otherwise. The semi-perimeter of the bounding box of the connection is divided into ten equal-length segments. Each segment point is used as a candidate position for the first inflection point of the Z-shape, thereby uniquely determining a two-dimensional Z-shaped candidate path.

[0110] (2) Flatten all candidate 2D connections: Flatten all candidate connections with L-shaped routing and Z-shaped routing;

[0111] (3) Parallel routing of candidate connections within a batch: A two-level parallel mechanism is used. At the level between connections, connections without resource conflicts are processed in parallel. At the same time, at the level within each connection, up to 10 candidate Z-shaped two-dimensional paths for each connection are parallelized. These two levels are combined to utilize the large number of threads on the GPU to perform routing and obtain the minimum routing cost for each candidate connection.

[0112] (4) Using GPU multithreading to find the best candidate connection for each connection and update the global grid;

[0113] (5) Repeat steps (3) and (4) until all batches are processed;

[0114] (6) Construct the original network topology: After completing the wiring of all connections, assemble the connection paths into a complete network wiring. For each network, the through-holes at the connection nodes are represented by an interval. After determining a connection path, update the intervals of the nodes at both ends of the connection to include the layer where the path is located. When all the connection wiring in the network is completed, a complete connected path is formed. Update the interval at each pin position to ensure that all pins in the network are correctly connected.

[0115] The algorithm is as follows:

[0116]

[0117] The first line of Algorithm 1 performs batch generation to ensure that the connections within the same batch do not conflict on routing resources, thereby achieving concurrent execution of routing tasks within each batch. Three-dimensional overlap detection is used to distinguish between metal layers in the horizontal and vertical directions, and a sparse Z-shaped routing strategy is applied to longer connections. In addition, if the connections within the same network are processed in parallel, the information of the upstream segment connected to the via may be missing, resulting in inaccurate timing cost calculation, especially for the via close to the source end, whose delay contribution cannot be ignored. To solve this problem, ensure that the first three levels of connections are routed strictly in sequence in different batches. For example, Figure 2 As shown, connection sa is at the first level, connection ab is at the second level, and connections bd and bc are at the third level. Lower-level connections must be scheduled for execution in subsequent batches relative to higher-level connections.

[0118] After the batch generation is completed, in line 2 of Algorithm 1, the purpose is to achieve parallel processing of multiple connections and multiple candidate paths for each connection. To this end, all candidate connections are flattened into an array so that a large number of GPU threads can process them simultaneously. Subsequently, the result array results is initialized in line 3. Different candidate paths for the same connection are distinguished based on the first turning point of each candidate path. For each candidate connection, its minimum routing cost and the corresponding metal layer index are calculated, and the results are stored in the result array. Since Z-type routing requires the storage of three layer information values and one cost value (a total of four values), while L-type routing and direct routing require fewer storage values (three and one), all entries are uniformly filled with four values for unified index processing in subsequent calculations.

[0119] Subsequently, as shown in lines 4 to 8 of Algorithm 1, after determining the minimum cost for all candidate connections, the optimal candidate path must be selected for each connection, its corresponding 3D routing path extracted, and the grid resource usage updated. To this end, a kernel is launched, with one thread assigned to each connection. Each thread scans its corresponding candidate connection result array to determine the path with the minimum cost. Because the candidate connections are stored in a non-uniform structure, a prefix sum array, connSum, is calculated to map connection indices to corresponding candidate connection indices. Once the optimal cost is determined, the first turning point of the selected path can be immediately determined because there is a one-to-one correspondence between them.

[0120] Step S6: Using delay-aware pin connection improvement technology to optimize the original network topology to further reduce path delay.

[0121] Optionally, in one embodiment of the present invention, in step S6, the delay ratio of the pin is calculated:

[0122]

[0123] Among them, delay actual is the actual delay of the pin, delay min is the delay assuming a direct shortest path connection from the source pin to this pin;

[0124] If the delay ratio is greater than the preset value, the connection of this pin is adjusted, and the A* algorithm is used to reconnect to the source pin. The reconnection process starts from this pin and uses breadth-first search (BFS) to explore the wiring space. In each step of expansion, the G-cell with the lowest cost is expanded outward. The cost of each G-cell consists of two parts: the wiring cost and the cost of the wiring. p (u,s) and the guided cost cost g (u,t); wiring cost cost p(u,s) represents the routing cost from G-cell u to this pin s, guiding cost cost g (u, t) represents the weighted two-dimensional distance from u to the target point t. To avoid the path overlapping with the existing routing, the routing cost of the occupied G-cell is given infinitely large:

[0125]

[0126] cost g (u,t)=w g *dist(u,t) (16)

[0127] in, is the cost of G-cell u, u is traversed G-cell means u is the G-cell occupied by the network, u is the current G-cell in the expansion, s is the pin, t is the source pin of the network, cost g (u, t) represents the weighted two-dimensional distance to the target point, w g To guide the cost weight, dist(u,t) is the two-dimensional distance between u and t, cost p (u,s) represents the routing cost from G-cell u to this pin s, including the cost of grid edges and vias; the cost is calculated as a weighted combination of timing cost and congestion cost.

[0128] When the search reaches the target (i.e. the source of the network), the found path will be integrated into the routing scheme of the network. Otherwise, the original routing path and via intervals will be restored to maintain the previous connectivity.

[0129] Step S7: Perform congestion-driven GPU accelerated routing and routing strategies on non-critical networks to effectively alleviate congestion.

[0130] Optionally, in one embodiment of the present invention, in step S7, for non-critical networks, only congestion costs are considered during routing, and congestion-driven GPU-accelerated routing is performed: first, a directed acyclic graph is constructed, and an L-shaped pattern is used for initial routing. For networks with overflow, the second stage is entered to create and move Steiner points to reorganize existing paths and generate alternative paths.

[0131] According to the GPU parallel accelerated global routing method for timing and congestion collaborative optimization proposed in an embodiment of the present invention, based on a given netlist, the ultra-large network is divided using the Kruskal algorithm combined with a search set to construct its timing propagation path; network decomposition based on pin-based timing margin estimation is performed; the timing weight calculation method based on the Elmore delay model is used to calculate the timing weight of a 2-pin network; a cost function that comprehensively considers timing and congestion costs is used to calculate path cost; two-level GPU parallel kernel acceleration mode routing is performed; network topology is optimized using delay-aware pin connection improvement technology; congestion-driven GPU accelerated routing of non-critical networks and its winding strategy are performed. The present invention can quickly obtain high-quality routing results with timing congestion balance, effectively improve timing performance, and meet the needs of the current ultra-large-scale high-performance circuit design routing stage.

[0132] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples without contradiction.

[0133] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "N" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0134] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or N executable instructions for implementing a custom logical function or step of a process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

Claims

1. A GPU parallel accelerated global routing method for timing and congestion coordinated optimization, characterized in that: The following steps are involved: Step S1: Based on a given netlist, the Kruskal algorithm combined with the union-find data structure is used to partition the super-large network and construct the time-series propagation path of the super-large network; Step S2, performing network decomposition based on pin-based timing slack estimation according to the timing propagation path, dividing the network into a critical network and a non-critical network, and decomposing spanning trees obtained from the critical network and the non-critical network into 2-pin networks; Step S3, for the critical network, using a timing weight calculation method based on the Elmore delay model to calculate the timing weight of the 2-pin network, and determining the timing priority critical connection and the regular critical connection by sorting the timing weights; Step S4, calculating the path cost of the timing-priority critical connection and the conventional critical connection using a cost function that comprehensively considers timing and congestion cost; Step S5, executing two-level GPU parallel kernel accelerated mixed mode routing to obtain the original network topology; Step S6, optimizing the original network topology using a delay-aware pin connection improvement technique; Step S7: performing congestion-driven GPU accelerated routing and routing strategy on the non-critical network.

2. The method according to claim 1, characterized in that In step S1, for a very large network with more pins than a preset number, the minimum spanning tree is first generated using the Kruskal algorithm combined with a union-find data structure. During the minimum spanning tree construction process, when the size of a connected component exceeds a predefined threshold, the connected component is split into subnetworks. After the partitioning is completed, each original network is divided into multiple sub-networks. The remaining minimum spanning tree edges connecting different sub-networks are then used to construct an adjacency table to represent the connection relationship between the sub-networks. Starting from the sub-network containing the source pin of the network, a breadth-first search traversal is performed. During the traversal process, a directed acyclic graph is constructed to record the path dependency relationship between the sub-networks. After each sub-network is decomposed independently, the complete temporal propagation path across the sub-network is reconstructed using a directed acyclic graph.

3. The method according to claim 1, characterized in that In step S2, the network is divided into critical networks and non-critical networks according to the estimated pin timing margin information. If a network contains any pin with negative margin, it is considered a critical network; otherwise, it is considered a non-critical network. For critical networks, the SALT algorithm is used to generate shallow lightweight trees to obtain shorter source-to-sink paths; for non-critical networks, the FLUTE algorithm is used to minimize bus length; The resulting spanning trees are all decomposed into 2-pin networks, which are called a connection of this network.

4. The method according to claim 1, wherein In step S3, the timing weight of the connection is calculated, and the connection The specific calculation formula for timing weight is: Weight ab =-l ab ·LoadCap b ·SlackSum b (1) Among them, l ab is the Manhattan distance between pin a and pin b, LoadCap b is the load capacitance of pin b, calculated similarly to the Elmore model: Among them, child is the collection of child nodes of the current node, l g is the Manhattan distance between the current node and the child node g, C0 is the unit capacitance value, here we take the median of the unit capacitance values of all metal layers, LoadCap g is the load capacitance of child node g, PinCap node When the node is a pin, it is the inherent capacitance of the pin; when it is not a pin, it is 0; SlackSum b is the cumulative sum of the negative timing slack of pin b: Among them, SlackSum c Slack is the cumulative sum of the negative timing slack of child node g. node When the node is a pin, it is the timing margin of the pin; when it is not a pin, it is 0; After calculating the timing weights of all connections, connections with non-zero timing weights are defined as critical connections, and connections with timing weights of 0 are defined as non-critical connections; the non-zero timing weights are further sorted in descending order, and critical connections are divided into two categories: timing-priority critical connections and regular critical connections; the classification is based on a predefined ratio parameter α (0<α<1), and the α×Nth value in the sorted timing weight vector is selected as the threshold, where N is the total number of connections; connections with timing weights greater than the threshold are identified as timing-priority critical connections, while the rest are classified as regular critical connections.

5. The method according to claim 4, characterized in that For critical connections and non-critical connections, when wiring, the critical connections of the network are extracted first, and the critical connections are wired first. The non-critical connections of the network are wired after the wiring of all the critical connections of the network is completed.

6. The method according to claim 1, characterized in that In step S4, pattern routing is performed using a differentiated cost function for the timing-priority critical connections and regular critical connections classified based on timing weights in the critical network, as well as the non-critical network. In pattern routing, the cost of routing a line segment pq on the Lth metal layer is: Cost=w c ·Cost c +w t ·Cost t (4) Among them, w c and w t are congestion weight and timing weight, Cost c and Cost t are congestion cost and timing cost respectively, w t Defined as: Among them, when the current connection is a key connection, w t When the current connection is a non-critical connection, w t is 0; Cost c Adjust based on whether the connection is timing-prioritized. For timing-prioritized critical connections: For regular critical and non-critical connections: Among them, w R and are the pre-set weights and the unit resistance of the Lth metal layer, Cost e is the congestion cost of a mesh edge, E wire is the set of mesh edges that the line segment passes through; for a mesh edge with capacity c and demand d, the congestion cost is calculated as: Cost e =e s(d+1-c) -e s(d-c) (8) The timing cost of the line segment in formula (4) is t The calculation method is: Cost t =RES Via (CAP wire +LoadCap q )+RES wire (0.5·CAP wire +LoadCap q ) (10) Among them, RES Via Represents the total resistance of all vias connected to the upstream segment, LoadCap node is the cumulative capacitance of the node node, which is consistent with the node load capacitance used in the timing weight calculation; RES wire and CAP wire is the resistance and capacitance of the line segment, RES Via RES wire and CAP wire Expressed as: Among them, vias is the set of vias that the current segment passes through to connect to the upstream segment. represents the resistance of a single via between layer l and layer l+1, wl seg is the length of the line segment, is the unit resistance of the Lth layer, is the unit capacitance of the Lth layer.

7. The method according to claim 1, characterized in that In step S5, the two-level GPU parallel kernel accelerated mixed mode routing specifically includes the following steps: (1) Batch generation based on a novel overlap check: vertical and horizontal routing resources are distinguished during the check. L-shaped routing is used for connections with a length of less than 10 grid edges, while sparse Z-shaped routing is used otherwise. The semi-perimeter of the bounding box of the connection is divided into ten equal-length segments. Each segment point is used as a candidate position for the first inflection point of the Z-shape, thereby uniquely determining a two-dimensional Z-shaped candidate path. (2) Flatten all candidate 2D connections: Flatten all candidate connections with L-shaped routing and Z-shaped routing; (3) Parallel routing of candidate connections within a batch: A two-level parallel mechanism is used. At the level between connections, connections without resource conflicts are processed in parallel. At the same time, at the level within each connection, up to 10 candidate Z-shaped two-dimensional paths for each connection are parallelized. These two levels are combined to utilize the large number of threads on the GPU to perform routing and obtain the minimum routing cost for each candidate connection. (4) Using GPU multithreading to find the best candidate connection for each connection and update the global grid; (5) Repeat steps (3) and (4) until all batches are processed; (6) Construct the original network topology: After completing the wiring of all connections, assemble the connection paths into a complete network wiring. For each network, the through-holes at the connection nodes are represented by an interval. After determining a connection path, update the intervals of the nodes at both ends of the connection to include the layer where the path is located. When all the connection wiring in the network is completed, a complete connected path is formed. Update the interval at each pin position to ensure that all pins in the network are correctly connected.

8. The method according to claim 1, characterized in that In step S6, the delay ratio of the pin is calculated: Among them, delay actual is the actual delay of the pin, delay min is the delay assuming a direct shortest path connection from the source pin to this pin; If the delay ratio is greater than the preset value, the connection of this pin is adjusted; the adjustment uses the A* algorithm to reconnect to the source pin. The reconnection process starts from this pin and uses breadth-first search to explore the wiring space. In each step of expansion, the G-cell with the lowest cost is expanded outward. The cost of each G-cell consists of two parts: the wiring cost cost p (u,s) and the guided cost cost g (u,t): cost g (u,t)=w g *dist(u,t) (16) in, is the cost of G-cell u, u is traversed G-cell means u is the G-cell occupied by the network, u is the current G-cell in the expansion, s is the pin, t is the source pin of the network, cost g (u, t) represents the weighted two-dimensional distance to the target point, w g To guide the cost weight, dist(u,t) is the two-dimensional distance between u and t, cost p (u,s) represents the routing cost from G-cell u to this pin s, including the cost of grid edges and vias; the cost is calculated as a weighted combination of timing cost and congestion cost.

9. The method according to claim 1, characterized in that In step S7, for the non-critical networks, only congestion costs are considered during routing, and congestion-driven GPU-accelerated routing is performed: first, a directed acyclic graph is constructed, and an L-shaped pattern is used for initial routing. For networks with overflow, the second stage is entered to reorganize existing paths by creating and moving Steiner points, thereby generating alternative paths.

Citation Information

Cited By

  • Method and device for merging and decomposing FPGA lookup table, equipment and medium

    CN121503381A