A GPU-accelerated chip routing method for constructing a minimum rectilinear steiner tree
By using a breadth-first search approach with GPU acceleration, the Steiner tree search process is optimized, solving the problem of low CPU efficiency in large-scale integrated circuit chip wiring. This achieves efficient construction of the minimum right-angle Steiner tree, improving chip wiring efficiency.
Patent Information
- Application Number
- CN202211285801.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-20
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2042-10-20
AI Technical Summary
In the existing technology of large-scale integrated circuit chip wiring, the efficiency of using CPU to construct the minimum right-angle Steiner tree is low, and it is difficult to efficiently complete the wiring estimation of millions of wire nets.
GPU-accelerated computation is employed, and the Steiner tree search process is optimized through a breadth-first search approach. By leveraging the massive data parallel computing capabilities of GPUs, parallel partitioning and merging of the nets are achieved to construct the minimum right-angle Steiner tree.
Without reducing the solution accuracy, the construction efficiency of the minimum right-angle Steiner tree is significantly improved, thereby increasing the computational efficiency of integrated circuit chip wiring.
Smart Images

Figure CN115563927B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of integrated circuit design automation, and relates to a Steiner tree construction technology in integrated circuit chip wiring technology, in particular to a GPU accelerated computing minimum rectilinear Steiner tree construction method for chip wiring. BACKGROUND
[0002] Wiring is a core step of chip design, in which a chip design automation system calculates a wiring scheme for each net in the chip. A net in the chip is a set of pins, each pin being a two-dimensional coordinate representing its position on the chip. The task of chip wiring is to connect the pins inside each net using metal wires. In the early steps of chip design, the wiring scheme is usually approximated using a minimum rectilinear Steiner tree (RSMT). For a given net, using only horizontal and vertical lines to connect all the pins of the net, with any intermediate points, a rectilinear Steiner tree is obtained. The minimum rectilinear Steiner tree is defined as the rectilinear Steiner tree with the shortest total wire length.
[0003] Since the minimum rectilinear Steiner tree construction problem is difficult to solve accurately, existing methods seek a balance between time and solution accuracy. Among them, the Fast Look-Up Table Estimation (FLUTE) is the most typical method, which can maintain a relatively optimal solution accuracy while achieving high computational efficiency. However, as the size of chip design continues to grow, tens of millions of nets need to be solved for the minimum rectilinear Steiner tree for one wiring, which takes a lot of time. Existing methods including FLUTE can only run on CPUs, limited by the parallel computing and memory scheduling capabilities of multi-core CPUs, and can only utilize up to 16 threads of parallel computing resources, making it difficult to efficiently complete fast wiring estimation for large-scale circuits. SUMMARY
[0004] The purpose of the present application is to provide a GPU accelerated computing minimum rectilinear Steiner tree construction method to efficiently complete fast wiring estimation for large-scale circuits, overcoming the deficiencies of existing minimum rectilinear Steiner tree construction methods.
[0005] The application designs the Steiner tree search process into a GPU-friendly breadth-first calculation process, uses the large-scale data parallel computing capability of the GPU to simultaneously process the minimum right-angle Steiner tree construction task of a large number of line nets, improves the construction efficiency of the minimum right-angle Steiner tree, and improves the calculation efficiency of the integrated circuit chip wiring without reducing the solving accuracy.
[0006] The technical scheme of the application is:
[0007] A chip wiring method for constructing a minimum right-angle Steiner tree by using a GPU to accelerate calculation, comprising the steps of lookup table initialization, line net data initialization, line net parallel segmentation, and line net parallel solution merging.
[0008] A. lookup table initialization;
[0009] The FLUTE is used to construct a minimum right-angle Steiner tree lookup table for all line nets with a degree less than a lookup table threshold in a chip, comprising mapping from line net pin relative position coding to a potential minimum right-angle Steiner tree, wherein the degree of the line net is defined as the number of pins in the line net, the potential minimum right-angle Steiner tree is obtained by using the FLUTE method, and the lookup table threshold is a constant.
[0010] The minimum right-angle Steiner tree lookup table is flattened to obtain a flattened Steiner tree branch list and a branch lookup table index. The specific process is to store all branches of the potential minimum right-angle Steiner tree in an array according to the pin relative position coding order, so that the last branch of the previous Steiner tree is adjacent to the first branch of the next Steiner tree, and the obtained branch array is the flattened Steiner tree branch list; the index of the first branch of each potential minimum right-angle Steiner tree in the array is calculated to obtain the branch lookup table index.
[0011] The flattened Steiner tree branch list and the branch lookup table index are copied from the CPU memory to the GPU memory.
[0012] B. Line net data initialization
[0013] All line nets in the chip (i.e. the input line net) are flattened to obtain a pin list and a pin start position index of the line net, wherein the pin list of the line net includes a horizontal coordinate list and a vertical coordinate list. The specific process is to represent all pins in the input line net as horizontal coordinates and vertical coordinates, and store the horizontal coordinates and vertical coordinates in two arrays according to the line net number order, respectively, to obtain the horizontal coordinate list and the vertical coordinate list, which is the pin list of the line net; the index of the first pin of each line net in the horizontal coordinate list and the vertical coordinate list is calculated to obtain the pin start position index.
[0014] The pin list and the start position index are copied from the CPU memory to the GPU memory.
[0015] C. Line net parallel segmentation
[0016] The pin list and the pin start position index of the line net are iteratively segmented on the GPU to establish a hierarchical line net segmentation forest, and the line net segmentation forest is composed of multiple line net segmentation trees. The node of each line net segmentation tree is a line net, and the tree edge connecting the parent node and the child node represents the line net segmented from the parent node. The root node is the input line net, and the leaf node is the line net that does not need to be further segmented, i.e. the line net with a pin number less than the lookup table threshold.
[0017] The specific process of iterative segmentation is to start segmentation operation from the input line net, which constitutes the first layer line net. For each line net in the current layer, the interval of the pin of the line net in the pin list is obtained through the pin start position index. For the line net whose degree is less than the threshold of the lookup table, no operation is performed; for the line net whose degree is greater than or equal to the threshold of the lookup table, one pin is selected as a segmentation point to divide the line net into two line nets with smaller degrees, which are the child nodes of the current line net on the line net segmentation tree. The specific way of segmentation is to divide the original line net into two parts above and below or left and right, try multiple segmentation schemes, select the segmentation pin and the segmentation direction, so that the line nets after segmentation are as uniform as possible, which is to enumerate all the segmentation pin and the segmentation direction, calculate the difference between the degrees of the two line nets obtained by segmentation, and select the segmentation scheme that can make the difference between the degrees the smallest. For each original line net, segmentation is performed to obtain multiple groups of child nodes. Each line net obtained by segmentation is connected to its original line net by a tree edge of the line net segmentation tree to obtain the next layer line net pin list. The segmentation operation is repeated on the obtained next layer line net pin list and pin start position index until the number of pins of all line nets is less than the threshold of the lookup table.
[0018] D. Parallel solving and merging of line nets
[0019] The segmented line net is solved on the GPU. The specific way of solving is:
[0020] 1. If the degree of the line net is less than the threshold of the lookup table, the minimum right-angle Steiner tree of the line net is obtained through the flattened Steiner tree branch list and the branch lookup table index. The specific method is to calculate the relative position encoding of the pin, obtain the interval of all branches of the minimum right-angle Steiner tree in the branch list through the branch lookup table index, and these branches constitute the minimum right-angle Steiner tree of the line net.
[0021] 2. If the degree is greater than the threshold of the lookup table, the minimum right-angle Steiner tree of the line net is obtained by merging the solving results of the lower layer line nets, and the specific method is to connect the minimum right-angle Steiner trees of the lower layer line nets at the segmentation point. For multiple segmentation schemes, the scheme with the minimum total length of the right-angle Steiner tree is selected as the segmentation scheme of the line net.
[0022] The order of solving the line net is from the lower layer to the upper layer on the hierarchical line net segmentation forest. The merging result of the top layer line net is obtained, which is the solving result of the minimum right-angle Steiner tree of the input line net.
[0023] Compared with the prior art, the beneficial effects of the present application are:
[0024] The present application restructures the minimum rectilinear Steiner tree search process into a GPU-friendly breadth-first computation flow by flattening the Steiner tree lookup table and the net data (i.e. steps A, B), and uses the large-scale data parallel computing capability of the GPU to simultaneously process the minimum rectilinear Steiner tree construction task of a large number of nets. Without reducing the solving accuracy, the construction efficiency of the minimum rectilinear Steiner tree is improved, and the integrated circuit chip wiring estimation is efficiently realized. BRIEF DESCRIPTION OF DRAWINGS
[0025] Figure 1 A flowchart of a chip wiring method for constructing a minimum rectilinear Steiner tree by GPU accelerated computation provided by the present application.
[0026] Figure 2 A schematic diagram of a minimum rectilinear Steiner tree. In the diagram, the solid circle points represent the four pins of the net, and the solid lines represent the metal wires on the Steiner tree.
[0027] Figure 3 A schematic diagram of net pin coordinates and relative position encodings. In the diagram, Pin 0-3 represent the four pins of the net, xs0-x s3 represent the horizontal coordinates of the four pins sorted in ascending order, ys0-ys3 represent the vertical coordinates of the four pins sorted in ascending order, and s0-s3 represent the relative position encodings of the pins of the net.
[0028] Figure 4 A schematic diagram of a hierarchical net partitioning forest. Layer represents the hierarchy of nodes, m represents the number of the last layer, and the numbering starts from 0, with layer 0 being the first layer. Net 0-net n represent the 0th to nth nets, subnetXX represents a child node net with the number XX, and the arrows represent the tree edges of the net partitioning tree.
[0029] Figure 5 A schematic diagram of the process of parallel net partitioning and parallel net solving and merging. In the diagram, the meanings of net, subnet are the same as those defined in Figure 4 , the right side "partitioning" label marks the position of the parallel net partitioning step, the "merging" label marks the position of the parallel net solving and merging step, the left side axis labeled "time" is the time axis of the processing process of the method, and RSMTXX represents the minimum rectilinear Steiner tree of the XX net calculated. DETAILED DESCRIPTION
[0030] The present application will be further described below by examples in conjunction with the drawings, but the scope of the present application is not limited in any way by the examples.
[0031] The present application provides a method for constructing a minimum rectilinear Steiner tree for chip wiring by GPU accelerated computation, which can efficiently construct a minimum rectilinear Steiner treeFigure 2 The minimum rectilinear Steiner tree is shown to realize the chip routing. The method includes the steps of lookup table initialization, net data initialization, net parallel partitioning, net parallel solving and merging (as shown in Figure 1 ). Specifically, the method includes the following steps:
[0032] A. Lookup table initialization
[0033] A minimum rectilinear Steiner tree lookup table is constructed for all nets whose degrees are less than the lookup table threshold using FLUTE. The lookup table contains the mapping from the pin relative position encoding of a net to the potential minimum rectilinear Steiner tree. As shown in Figure 3 , the horizontal and vertical coordinates of all pins are sorted respectively to obtain the xs and ys arrays, which represent the sorted horizontal and vertical coordinates. Then, the s array is calculated, and the kth position of the s array represents the vertical coordinate rank of the pin ranked kth in the horizontal coordinate. The s array is the pin relative position encoding of the net. The specific construction process of the minimum rectilinear Steiner tree lookup table adopts the FLUTE method. The lookup table threshold is a constant, and in this embodiment, the default setting of FLUTE is adopted, which is 9.
[0034] The minimum rectilinear Steiner tree lookup table is flattened to obtain a flattened Steiner tree branch list and a branch lookup table index. The specific process is to store all branches of the potential minimum rectilinear Steiner tree in an array according to the pin relative position encoding, so that the last branch of the previous Steiner tree is adjacent to the first branch of the next Steiner tree. The obtained branch array is the flattened Steiner tree branch list. The index of the first branch of each potential minimum rectilinear Steiner tree in the array is calculated to obtain the branch lookup table index. The pseudo code for calculating the flattened Steiner tree branch list and the branch lookup table index is as follows:
[0035]
[0036] where i represents the loop variable, L represents the size of the minimum rectilinear Steiner tree lookup table, the square brackets [] represent array addressing, TreeSize[i] represents the size of the ith potential minimum rectilinear Steiner tree, TableOffset array represents the branch lookup table index, Memcopy represents the memory copy operation, FlattenedBranch array represents the flattened Steiner tree branch list, and Branch[i] represents the branch list of the ith potential minimum Steiner tree.
[0037] The flattened Steiner tree branch list and the branch lookup table index are copied from the CPU memory to the GPU memory by calling the cudaMemcpyAsync function on the CUDA platform.
[0038] B. Net data initialization
[0039] The input line net is flattened to obtain the pin list and pin start position index of the line net, wherein the pin list of the line net comprises a horizontal coordinate list and a vertical coordinate list. The specific process is to represent all pins in the input line net as horizontal coordinates and vertical coordinates, and store the horizontal coordinates and vertical coordinates in two arrays in the order of line net numbering respectively to obtain the horizontal coordinate list and the vertical coordinate list, i.e. the pin list of the line net; and to calculate the subscript of the first pin of each line net in the horizontal coordinate list and the vertical coordinate list to obtain the pin start position index. The pseudo code for calculating the horizontal coordinate list, the vertical coordinate list and the pin start position index is as follows:
[0040]
[0041] wherein NetDegree[i] represents the degree of the i-th line net, NetOffset represents the pin start position index, X[i], Y[i] respectively represent the horizontal coordinates and vertical coordinates of all pins of the current i-th line net, and FlattenedX, FlattenedY respectively represent the horizontal coordinate list and the vertical coordinate list.
[0042] The pin list and the start position index are copied from the CPU memory to the GPU memory by calling the cudaMemcpyAsync function on the CUDA platform.
[0043] C. Parallel segmentation of line nets
[0044] The pin list and the pin start position index of the line net are iteratively segmented on the GPU to establish a hierarchical line net segmentation forest. As shown in Figure 4 , the line net segmentation forest is composed of multiple line net segmentation trees, and the nodes of each line net segmentation tree are line nets. The tree edges connecting the parent nodes and the child nodes represent the line nets obtained by segmenting the parent node line net, the root node is the input line net, and the leaf node is the line net that does not need to be further segmented, i.e. the line net with the number of pins less than the threshold of the lookup table. As shown in Figure 4 , net0 is a root node, which is the first input line net; subnet00 is a line net with the number of pins less than the threshold of the lookup table, which does not need to be further segmented and has no child node; and subnet01 is a line net with the number of pins greater than or equal to the threshold of the lookup table, which has child nodes.
[0045] The specific process of iterative segmentation is to start the segmentation operation from the input line net, and the input line net constitutes the first layer of line nets, i.e. layer 0 in Figure 4 . For each line net in the current layer, the interval in which the pins of the line net are located in the pin list is obtained through the pin start position index. The degree of the line net is defined as the number of pins in the line net, and the line net with a degree less than the threshold of the lookup table, such as Figure 4subnet00, do nothing; for the line net whose degree is greater than or equal to the threshold of the lookup table, such as subnet01 in Figure 4 subnet01, select one pin as the split point to split the line net into two smaller line nets, i.e., two child nodes. The specific way of selecting the split point is to select the pin and the split direction that can make the split line nets as uniform as possible, and split the line net into two parts, up and down or left and right. Connect the smaller line nets obtained by splitting to the original line net with the tree edge of the line net splitting tree, and obtain the pin list of the next layer line net, such as Figure 4 layer 1 is obtained from layer 0. Repeat the splitting operation on the obtained pin list of the next layer line net and the pin start position index to obtain layer 2 from layer 1, until the number of pins of all line nets is less than the threshold of the lookup table.
[0046] In this embodiment, we record which line net each child node comes from the previous layer, such as Figure 5 shown in the Break stage of
[0047] D. Parallel solution merging of line nets
[0048] Solve the split line net on the GPU. The specific way of solving is:
[0049] 1. If the degree of the line net is less than the threshold of the lookup table, obtain the minimum rectilinear Steiner tree of the line net through the flattened Steiner tree branch list and the branch lookup table index, such as Figure 5 subnet03 in Figure 3 the first step of Merge obtains RSMT011. The specific method is to calculate the relative position code of the pin according to the method described in the branch lookup table index to obtain the interval of all branches of the minimum rectilinear Steiner tree in the branch list. The specific implementation method is to use the TableOffset array in step A. If the relative position code of the pin corresponds to the i-th position of the minimum rectilinear Steiner tree lookup table, then the interval in the branch list is TableOffset[i] to TableOffset[i+1]-1. These branches constitute the minimum rectilinear Steiner tree of this line net.
[0050] Figure 5 2. If the degree is greater than the threshold of the lookup table, obtain the minimum rectilinear Steiner tree of the line net from the solution results of the lower layer line nets, and the specific method is to connect the minimum rectilinear Steiner trees of the lower layer line nets at the split point. For multiple split schemes, take the scheme with the minimum total rectilinear Steiner tree length as the split scheme of the line net. For example, RSMT000 and RSMT011 in merge to obtain RSMT00.
[0051] The order of solving the line net is from the lower layer to the upper layer on the hierarchical line net partition forest, which is contrary to the order of partitioning the line net in parallel, and in the Figure 5 The Break part and the Merge part are mirror-symmetric in the Figure 5 After RSMT000 and RSMT011 are merged to obtain RSMT00, RSMT01 is continued to be merged, RSMT02 and RSMT03 are merged to obtain two right-angle Steiner trees of net0, and the one with the minimum total length is taken as RSMT0. The merging results RSMT0 and RSMT1 of the top layer line net are obtained, which are the solving results of the minimum right-angle Steiner tree of the input line net.
[0052] It should be noted that the purpose of the disclosed embodiments is to help further understand the present application, but those skilled in the art can understand that various replacements and modifications are possible without departing from the scope of the present application and the appended claims. Therefore, the present application should not be limited to the disclosed embodiments, and the scope of the present application is defined by the scope of the claims.
Claims
1. A GPU-accelerated method of chip routing to construct a minimum rectilinear Steiner tree, The method is characterized by comprising the steps of: A. Look-up table initialization, obtaining the flattened Steiner tree branch list and branch look-up table index, and copying from the CPU memory to the GPU memory; comprising: A1. For all the wire nets in the chip whose degree is less than the look-up table threshold, constructing the minimum rectilinear Steiner tree look-up table; the minimum rectilinear Steiner tree look-up table contains the mapping from the pin relative position coding to the potential minimum rectilinear Steiner tree; the look-up table threshold is a constant; A2. Flattening the minimum rectilinear Steiner tree look-up table to obtain the flattened Steiner tree branch list and branch look-up table index; the specific process is to store all the branches of the potential minimum rectilinear Steiner tree in an array according to the pin relative position coding order, so that the last branch of the previous Steiner tree is adjacent to the first branch of the next Steiner tree, and the obtained branch array is the flattened Steiner tree branch list; the index of the first branch of each potential minimum rectilinear Steiner tree in the array is calculated to obtain the branch look-up table index; A3. Copying the flattened Steiner tree branch list and branch look-up table index from the CPU memory to the GPU memory; B. Wire net data initialization, obtaining the pin list and pin start position index of the wire net, and copying from the CPU memory to the GPU memory; B1. All the wire nets in the chip are input wire nets, and the input wire nets are flattened to obtain the pin list and pin start position index of the wire net, wherein the pin list of the wire net comprises a horizontal coordinate list and a vertical coordinate list; B2. Copying the pin list and start position index from the CPU memory to the GPU memory; C. Parallel wire net segmentation, establishing a hierarchical wire net segmentation forest; Iterative segmentation is performed on the pin list and pin start position index of the wire net on the GPU to establish a hierarchical wire net segmentation forest, and the wire net segmentation forest is composed of multiple wire net segmentation trees, the nodes of each wire net segmentation tree are wire nets, and the tree edges connecting the parent nodes and the child nodes represent the wire nets obtained by segmenting the parent wire net, the root node is the input wire net, and the leaf node is the wire net that does not need to be further segmented, i.e. the wire net whose pin number is less than the look-up table threshold; The specific process of iterative segmentation comprises: C1. Starting segmentation from the input wire net, and the input wire net constitutes the first layer of wire nets; C2. For each wire net in the layer, the interval in which the pin of the wire net is located in the pin list is obtained through the pin start position index; C3. Defining the degree of the wire net as the number of pins in the wire net, and not performing any operation on the wire net whose degree is less than the look-up table threshold; for the wire net whose degree is greater than or equal to the look-up table threshold, selecting one pin as a segmentation point to segment the wire net into two wire nets, i.e. the child nodes of the current wire net on the wire net segmentation tree; D. Parallel solving and merging of wire nets: solving the segmented wire nets on the GPU, and the specific way of solving is: D1. If the degree of the line net is less than the threshold value of the lookup table, the minimum rectilinear Steiner tree of the line net is obtained by the flattened Steiner tree branch list and branch lookup table index, and the specific method is to calculate the relative position code of the pin, obtain the interval of all branches of the minimum rectilinear Steiner tree in the branch list through the branch lookup table index, and these branches constitute the minimum rectilinear Steiner tree of the line net; D2. If the degree of the line net is greater than the threshold value of the lookup table, the minimum rectilinear Steiner tree of the line net is obtained by merging the solving results of the lower layer line net, and the specific method is to connect the minimum rectilinear Steiner trees of the lower layer line net at the split point; for multiple split schemes, the scheme with the minimum total length of the rectilinear Steiner tree is selected as the split scheme of the line net; D3. The order of solving the line net is from the lower layer to the upper layer in the hierarchical line net split forest, and the merging result of the top layer line net is obtained, that is, the solving result of the minimum rectilinear Steiner tree of the input line net; E. According to the obtained solving result of the minimum rectilinear Steiner tree of the input line net, the chip routing of the minimum rectilinear Steiner tree is realized by GPU acceleration.
2. The GPU accelerated chip routing method for constructing a minimum rectilinear Steiner tree of claim 1, wherein, In step A3, the cudaMemcpyAsync function is called in the CUDA platform to copy the flattened Steiner tree branch list and branch lookup table index from the CPU memory to the GPU memory.
3. The GPU accelerated chip routing method for constructing a minimum rectilinear Steiner tree of claim 1, wherein, In step B1, the input line net is flattened, and the specific process is to represent all pins in the input line net as horizontal coordinates and vertical coordinates, and store the horizontal coordinates and vertical coordinates in two arrays according to the line net number sequence to obtain the horizontal coordinate list and vertical coordinate list, that is, the pin list of the line net; the index of the first pin of each line net in the horizontal coordinate list and the vertical coordinate list is calculated to obtain the pin start position index.
4. The GPU accelerated chip routing method for constructing a minimum rectilinear Steiner tree of claim 1, wherein, In step B2, the cudaMemcpyAsync function is called in the CUDA platform to copy the pin list and start position index from the CPU memory to the GPU memory.
5. The GPU accelerated chip routing method for constructing a minimum rectilinear Steiner tree of claim 1, wherein, In step C3, the specific way of splitting is: The original line net is divided into two parts above and below or left and right, and multiple split schemes are tried, and the split point pin and the split direction are selected to make the split line net uniform; Each original line net is divided to obtain a plurality of groups of child nodes; Each line net obtained by splitting is connected to the corresponding original line net by the tree edge of the line net split tree to obtain the pin list of the next layer line net; The pin list and the pin start position index of the next layer line net are obtained, and the splitting operation is repeated until the number of pins of all line nets is less than the threshold value of the lookup table.
6. The GPU accelerated chip routing method for constructing a minimum rectilinear Steiner tree of claim 5, wherein, The method of making the split line net uniform is to enumerate all split point pins and split directions, calculate the degree difference of the two line nets obtained by splitting, and select the split scheme with the minimum degree difference.
Citation Information
Patent Citations
Complete optimal Steiner tree construction method based on lookup table
CN113947057A
Steiner tree handling equipment, steiner tree handling method, and steiner tree handling program
JP2005275780A