Chip partitioning and placement simulation method, device, and computer equipment in 3D scene
By constructing a hypergraph and optimizing node revocation using gain function, the problem of too tight component distances and weak correlation between hybrid bond terminals in chip layout design in the prior art is solved, and more efficient chip utilization and shorter half-circumference length are achieved, reducing chip power consumption.
Patent Information
- Application Number
- CN202310024383.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-09
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2043-01-09
AI Technical Summary
During the division and placement process of existing chip layout design methods, too much attention is paid to the number of wires between component modules, resulting in too close components distances, affecting the minimum half-circumference length and chip utilization, and the position correlation of hybrid bond terminals is weak, affecting the final placement effect.
By constructing a hypergraph, node hierarchical shrinkage and division are performed based on node weights, node undo and tuning are performed in combination with the gain function for 3D placement scenarios, the placement of standardized units is optimized, the final partition netlist is generated, and the placement model is input for simulation.
It improves the chip utilization rate and the half-circumference length of the standardized unit, reduces the overall power consumption of the chip, and improves the effect of placing simulation results.
Smart Images

Figure CN116245066B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of electronic design automation technology, and in particular to a chip partitioning and placement simulation method, apparatus, computer equipment, storage medium, and computer program product in a 3D scene. Background Art
[0002] With the increasing application of chips, chip layout design is an important component of chip design. Minimizing the chip's semi-perimeter is a crucial step in large-scale integrated chip layout design. During analog design of integrated chip layout, all component modules are placed on the target chip, connected by wires. Sometimes, these components need to be connected across planes. Therefore, it is necessary to study the placement of component modules with cross-plane connections to keep the wires between the components as short as possible.
[0003] Existing methods often use placement simulation methods, which divide and place the planes to which component modules belong according to certain rules. During the division and placement process, this method often pays too much attention to the number of wires between component modules, resulting in the distance between different components being too close. In turn, the division result cannot meet the chip placement scenario, affecting the final minimum half-circle line length. Summary of the Invention
[0004] Based on this, it is necessary to provide a chip partitioning and placement simulation method, device, computer equipment, computer-readable storage medium and computer program product in a 3D scene that can improve the chip placement effect in order to address the above technical problems.
[0005] In a first aspect, the present application provides a chip partitioning and placement simulation method in a 3D scene, the method comprising:
[0006] Preprocessing the parameters of the netlist to be placed and the standardized units in the netlist to obtain a hypergraph with the standardized units as nodes and the connecting lines between the standardized units as hyperedges;
[0007] Based on the node weights of the hypergraph, performing node hierarchical contraction and division on the nodes in the hypergraph to obtain an initial division result;
[0008] Deleting the contracted nodes in the hypergraph according to the initial partitioning result, and during the deleting process, estimating the change in the minimum semi-circle length after each movement using a gain function designed for a 3D placement scenario, and optimizing the standardized units during the node deleting process based on the change in the minimum semi-circle length to obtain a final partitioned netlist;
[0009] The partitioned netlist is input into a placement model to obtain a 3D placement simulation result.
[0010] In a second aspect, the present application provides a device for simulating chip partitioning and placement in a 3D scene, the device comprising:
[0011] A preprocessing module, configured to preprocess the parameters of the netlist to be placed and the standardized units in the netlist to obtain a hypergraph having the standardized units as nodes and the connecting lines between the standardized units as hyperedges;
[0012] A processing module, configured to perform node hierarchical contraction and partitioning on the nodes in the hypergraph based on the node weights of the hypergraph to obtain an initial partitioning result;
[0013] a partitioning module configured to cancel the contracted nodes in the hypergraph according to the initial partitioning result, and during the cancellation process, estimate the change in the minimum semi-circle length after each movement based on a gain function designed for a 3D placement scenario in an optimization algorithm, and optimize the standardized units during the node cancellation based on the change in the minimum semi-circle length to obtain a final partitioned netlist with a gain function designed for a 3D placement scenario;
[0014] A placement module is used to input the partitioned netlist into a placement model to obtain a 3D placement simulation result.
[0015] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0016] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when executed by a processor.
[0017] In a fifth aspect, a computer program product is provided, comprising a computer program, which implements the above method steps when executed by a processor.
[0018] The aforementioned chip partitioning and placement simulation method, apparatus, computer device, storage medium, and computer program product for 3D scenarios utilizes hypergraph-based node weights to perform hierarchical node contraction and partitioning of nodes within the hypergraph. By considering the order of standardized unit area sizes, chip utilization is fully improved, resulting in an initial partitioning result. A gain function designed for 3D placement scenarios is then used to optimize the initial partitioning result during the node removal process, resulting in a final partitioned netlist. This partitioned netlist is then input into the placement model to obtain a 3D placement simulation result. This method, while taking into account the chip utilization of the standardized units, minimizes the semi-circle length of the standardized units, resulting in a reasonable partitioned netlist, improving the placement simulation results and reducing overall chip power consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 A diagram illustrating an application environment of a chip partitioning and placement simulation method in a 3D scene according to an embodiment;
[0020] Figure 2 A schematic flow chart of a chip partitioning and placement simulation method in a 3D scene according to one embodiment;
[0021] Figure 3 is a schematic diagram of a netlist in one embodiment;
[0022] Figure 4 is a schematic diagram of an initial partitioning result in one embodiment;
[0023] Figure 5 1 is a flow chart of a method for dividing standardized units in one embodiment;
[0024] Figure 6 A schematic flow chart of a method for generating a partitioned netlist in one embodiment;
[0025] Figure 7 A schematic diagram of a simulation result is provided in one embodiment;
[0026] Figure 8 A schematic flow chart of a method for generating placement simulation results in one embodiment;
[0027] Figure 9 1 is a flow chart of a placement method with D2D vertical chain connection based on Dreamplace in one embodiment;
[0028] Figure 10 A structural block diagram of a device for simulating chip partitioning and placement in a 3D scene in one embodiment;
[0029] Figure 11 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0030] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0031] 3D placement is a critical but time-consuming step in the very large scale integration (VLSI) design process. It is designed to reduce on-chip wiring, thereby reducing power consumption and improving performance, and achieving the effect of reducing costs. Placement determines the location of standardized cells in the physical layout, and the quality of the layout has a significant impact on the later stages of the process, such as the post-layout optimization stage. Therefore, how to complete high-quality placement quickly and at low cost has become an important hotspot in the field of placement problem research. In a 3D placement scenario with D2D vertical connection, in addition to considering the layout of the upper and lower DIEs, it is necessary to additionally consider the layout of the hybrid bonding terminals connecting the upper and lower DIEs. At the same time, in order to reduce costs, the technologies used by the upper and lower DIEs may also be different, which makes the entire placement process more complicated. The current 3D placement simulation method is roughly divided into two steps:
[0032] 1. To achieve 3D placement from two separate 2D placements, first in the partitioning stage, the designed netlist is represented by a hypergraph, which is then partitioned into two halves. Separate 2D placement methods are then used for the two partitioned dies, and finally 3D via assignments are used to constrain the placement results for the first and second layers.
[0033] Alternatively, 3D layout can be formulated and solved as a nonlinear programming (NLP) problem. The NLP objective is to find the weighted sum of the half-perimeter wire length (HPWL) and the number of 3D vias, with the constraint that the normalized bin-wise area of the cells is less than the maximum bin area within the 3D layout region. A gain function is then used to determine the placement results for the top and bottom layers.
[0034] However, existing 3D partitioning and placement methods still have some significant flaws. First, netlist partitioning before placement has a significant impact on placement. The rationality of the netlist partitioning results is directly related to the placement results. However, in the current partitioning process, the netlist partitioning calculation mostly only considers minimizing the number of cut edges, which often ignores the maximum utilization of different dies and the differences in the adopted technologies. As a result, the partitioning results are not well adapted to the placement scenario, affecting the final minimum half-perimeter line length (HPWL).
[0035] Second, the position of the hybrid bond terminals directly affects the calculation of the final half-circle length. However, in most current placement technologies, the placement results of the upper and lower dies are weakly correlated with the hybrid bond terminals. Hybrid bond terminals are often used to tune the standardized units of the upper and lower dies during the final optimization phase of placement. This ignores the critical role of hybrid bond terminals in the placement process, which in turn affects the gain function.
[0036] In view of this, the chip partitioning and placement simulation method in a 3D scene provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the data storage system can store data that the server 102 needs to process. The data storage system can be integrated on the server 102, or placed on the cloud or other network servers.
[0037] The server 102 obtains the parameters of the netlist and the standardized units in the netlist from the data storage system, and preprocesses the parameters of the netlist to be placed and the standardized units in the netlist to obtain a hypergraph with standardized units as nodes and connecting lines between standardized units as hyperedges; the server 102 performs node hierarchical contraction and division on the nodes in the hypergraph based on the node weights of the hypergraph to obtain an initial division result; the server 102 cancels the contracted nodes in the hypergraph according to the initial division result, and during the cancellation process, the standardized units at the time of node cancellation are tuned through the gain function designed for the 3D placement scenario to obtain the final divided netlist with the gain function designed for the 3D placement scenario; the server 102 inputs the divided netlist into the placement model to obtain a 3D placement simulation result.
[0038] The server 102 may be implemented as an independent server or a server cluster consisting of multiple servers.
[0039] In one embodiment, Figure 2 As shown, a chip partitioning and placement simulation method in a 3D scene is provided, and the method is applied to Figure 1 The following steps are used as an example to illustrate the server in the example:
[0040] S202 , preprocessing the parameters of the netlist to be placed and the standardized units in the netlist to obtain a hypergraph with the standardized units as nodes and the connecting lines between the standardized units as hyperedges.
[0041] The netlist can be the input format of the chip, and includes the chip layout, standardized cells, and the connection lines of the standardized cells. The netlist can be obtained from the corresponding technology library. The netlist in the technology library contains standardized cell parameters, that is, the size of the standardized cells and the location information of the pins of each standardized cell.
[0042] Specifically, a netlist to be placed is obtained from a technology library, and the netlist is preprocessed according to the chip layout and parameters of standardized units contained in the netlist.
[0043] The preprocessing of the netlist may be to convert the netlist into a hypergraph. A hypergraph is a generalized graph structure including at least one set of nodes and hyperedges. Specifically, the hypergraph can be constructed by using standardized units as nodes and connecting lines between standardized units as hyperedges.
[0044] The location information of the pins in the standardized unit may also be used as the node.
[0045] Specifically, if Figure 3 The schematic diagram of the netlist shown includes: standardized cells (C1, C2, C3, C4, and C5), pins of the standardized cells (Pin P1, Pin P2, and Pin P3), connecting lines of the standardized cells, etc. The pins of each standardized cell are connected to form a network net through the connecting lines of the standardized cells.
[0046] S204: Based on the node weights of the hypergraph, the nodes in the hypergraph are shrunk and divided in a hierarchical manner to obtain an initial division result.
[0047] The node weight of the hypergraph may be the area of the standardized unit. The larger the area of the standardized unit, the greater the node weight of the standardized unit.
[0048] Specifically, a hypergraph is constructed based on the number of standardized units in the integrated netlist, the number of connection lines between standardized units, the area of standardized units, and the location information of the pins of each standardized unit. The expression of the hypergraph is as follows:
[0049] H=(V,E,c,ω);
[0050] Where H represents a hypergraph, V represents a group of n nodes, and E represents a group of m hyperedges. Nodes can correspond to standardized units, and hyperedges can correspond to the connecting lines of standardized units. c represents the weight of the node. Considering the utilization limit of each chip, the node weight is designed to be equal to the area of each standardized unit. ω represents the weight of the hyperedge, and the hyperedge weight can be set to 1.
[0051] Among them, shrinking the nodes in the hypergraph refers to shrinking the highly connected nodes in the hypergraph, thereby reducing the number of connection lines and the length of connection lines of standardized units in the hypergraph, obtaining a hypergraph with a smaller structure, and facilitating subsequent initial partitioning steps and local searches.
[0052] Specifically, hierarchical contraction is adopted for the nodes of the hypergraph, including: first contracting the hypergraph nodes with smaller node weights in the hypergraph nodes, and then contracting the hypergraph nodes with larger node weights in the hypergraph nodes to obtain the contraction results of the hypergraph nodes.
[0053] Among them, initial division can be performed according to the contraction result of the hypergraph nodes, and the standardized units corresponding to each hypergraph node can be divided into different chips to obtain the initial division result.
[0054] Specifically, if Figure 4 The schematic diagram of the initial partitioning result shown includes: standardized cells (C1, C2, C3, C4, and C5), the bottom chip, and the top chip. Among them, standardized cells C1, C3, and C5 are divided into the bottom chip, and standardized cells C2 and C4 are divided into the top chip. It should be noted that after the partitioning, the position of the standardized cells needs to be fine-tuned so that the position of the standardized cells is within the row height of each chip and is placed across the rows on the surface, meeting the chip placement specifications.
[0055] S206, cancel the contracted nodes in the hypergraph according to the initial partitioning result, and during the cancellation process, estimate the change value of the minimum half-circle length after each movement through the gain function designed for the 3D placement scenario, and based on the change value of the minimum half-circle length, adjust the standardized units when canceling the nodes to obtain the final partitioned netlist.
[0056] Among them, all the shrunk nodes of the initial partitioning result can be canceled level by level. In the process of each level of cancellation, the gain function designed for the 3D placement scenario can be used to tune the standardized unit to obtain the final partitioned netlist.
[0057] It can be understood that the estimated change value of the minimum half-circle length after each movement is not equivalent to the actual value of the minimum half-circle length. Among them, adjustment can be performed based on the estimated change value of the minimum half-circle length after each movement. Specifically, if the change value is less than 0, it means that the minimum half-circle length decreases after the movement. If the change value is greater than 0, it means that the minimum half-circle length increases after the movement. The movement with a change value less than 0 can be selected for adjustment.
[0058] Specifically, during the process of retracting the shrunk nodes, optimization can also be performed. The basic idea behind this optimization is to move the nodes randomly or according to the optimization algorithm. Based on the semi-circle length of the network of standardized units corresponding to the moved nodes, the semi-circle length of the standardized units calculated by the gain function designed for 3D placement scenarios in the optimization algorithm is obtained. It should be noted that a smaller semi-circle length of the standardized units calculated by the gain function designed for 3D placement scenarios in the optimization algorithm indicates a greater gain effect from the movement of the optimized standardized units.
[0059] Specifically, the FM algorithm can be used to tune the partitioning results, repeatedly executing the node movement with the highest gain calculated by the gain function designed for the 3D placement scenario. Before each movement, it is necessary to determine whether the total area of the nodes in the chip after the movement exceeds the maximum usable area of the corresponding die. If so, the movement is canceled until all the shrunk nodes are revoked to obtain the final partitioned netlist.
[0060] S208 , inputting the partitioned netlist into the placement model to obtain a 3D placement simulation result.
[0061] The placement model is a chip placement model that processes the partitioned netlist and hybrid bond terminals to obtain placement simulation results for standardized cells and hybrid bond terminals. Specifically, the placement model can be a 3D placement model, which is a proprietary chip placement model for 3D placement scenarios.
[0062] Hybrid bonding terminals are the connecting components that connect the wires of standardized units in different chips. Hybrid bonding terminals are generally located on the top layer of the chip. The 3D structure of the chip after placement is hybrid bonding terminals, top chip, and bottom chip from top to bottom.
[0063] When the placement model is used to process the partitioned netlist and the hybrid bonding terminals, the hybrid bonding terminals and the standardized units are uniformly placed according to the placement model to obtain 3D placement simulation results of the hybrid bonding terminals and the standardized units.
[0064] It should be noted that the hybrid bonding terminals may be generated by generating the hybrid bonding terminals according to the partitioned netlist.
[0065] Specifically, after placement, it is necessary to check whether the standardized cells meet the placement specifications. The placement specifications include: the standardized cells are within the row height range, the standardized cells do not overlap with each other, and the standardized cells meet the chip utilization requirements after placement.
[0066] In the chip partitioning and placement simulation method described above for 3D scenarios, nodes in the hypergraph are hierarchically contracted and partitioned based on hypergraph node weights. By considering the order of standardized unit area sizes, chip utilization is fully improved, resulting in an initial partitioning result. A gain function designed for 3D placement scenarios is then used to optimize the initial partitioning result during the node removal process, resulting in a final partitioned netlist. This partitioned netlist is then input into the placement model to produce a 3D placement simulation result. This method, while taking into account the chip utilization of the standardized units, also minimizes the half-circle length of the standardized units, resulting in a reasonable partitioned netlist, improving the placement simulation results and reducing overall chip power consumption.
[0067] In one embodiment, based on the node weights of the hypergraph, the nodes in the hypergraph are hierarchically contracted and divided to obtain an initial division result, such as Figure 5 As shown, a flow chart of the method for dividing the standardized units includes:
[0068] S502: Obtain the contraction score of each node and neighboring node according to the weight of each node and neighboring node in the hypergraph.
[0069] Among them, neighbor nodes represent nodes adjacent to the node in the hypergraph.
[0070] Specifically, the contraction score of the node and its neighbor nodes can be obtained based on the product of the weights of node u and neighbor node v. The calculation formula for the contraction score of the node and its neighbor nodes is as follows:
[0071]
[0072] Among them, c(u) represents the node weight of node u, c(v) represents the node weight of node v, and size() represents the total number of nodes on the hypergraph network.
[0073] It should be noted that c(v)×c(u) represents the product of the node weights of node u and the adjacent node v, and also represents the product of the areas of the normalized units.
[0074] S504: Determine the node with the smallest area according to the contraction score of the node and its neighboring nodes, and contract the node with the smallest area.
[0075] The larger the shrinkage score of the node and its neighboring nodes, the smaller the area of the standardized unit. Therefore, the node with the smallest area is shrunk according to the shrinkage score of the node and its neighboring nodes.
[0076] Among them, after each contraction, the weight of the contracted node becomes the sum of the weights participating in the contraction.
[0077] S506, starting from a randomly selected vertex, the shrunken hypergraph is traversed using BFS to partition it and obtain the partition result.
[0078] BFS, short for Breadth-First-Search, is a graph-based search algorithm that is relatively easy to understand.
[0079] The BFS algorithm starts from the initial state of the problem (the starting point) and, according to the state transition rules (edges in the graph structure), traverses all possible states (other nodes) until it finds the final state (the end point). Therefore, the complexity of the BFS algorithm is closely related to the total number of states.
[0080] Specifically, BFS traversal is used to initialize and divide the shrunken hypergraph. Starting from a random vertex, BFS traversal is performed on the hypergraph until half of the hypergraph is found. The vertices visited during the traversal constitute block V0, and all remaining vertices constitute V1. In this way, the hypergraph is initialized and divided into two parts, that is, the current partition result is obtained.
[0081] V0 may represent the bottom chip, and V1 may represent the top chip. Specifically, a BFS traversal algorithm is used to divide each node (standardized unit) into the bottom chip or the top chip.
[0082] S508 , updating the shrunken node weights according to the shrunken normalized units.
[0083] The weight of the node after contraction can be obtained by calculating the sum of the weights of the node and its neighboring nodes before contraction, or by converting the area of the normalized unit after contraction.
[0084] S510, record the number of divisions.
[0085] S512, returning to the step of obtaining the contraction points of the nodes and neighboring nodes according to the weights of each node and neighboring nodes in the hypergraph, and obtaining multiple division results until the number of divisions reaches a preset number.
[0086] The preset number of times is at least two times, and the preset number of times can be set to 10 times.
[0087] S514: taking the division result that meets the preset requirements among the multiple division results as the initial division result.
[0088] Among them, the preset requirement can be the chip utilization requirement. For example, when the chip utilization in the division result exceeds the chip utilization requirement, it is considered that the division result does not meet the preset requirement. For example, when the chip utilization in multiple division results does not exceed the chip utilization requirement, the division result with the largest utilization is selected as the initial division result.
[0089] Among them, the preset requirement can also be a requirement for dividing the number of standardized units. For example, when the standardized unit in the first division result is divided into two equal parts, and the standardized unit in the second division result is divided into two unequal parts, the first division result is selected as the initial division result.
[0090] In this embodiment, by using node weight as the basis for shrinkage, nodes with smaller standardized unit areas corresponding to their corresponding node weights are prioritized for shrinkage. This prevents the subsequent removal step from preventing the highest-gain node movement from being executed due to the removed node area being too large and not meeting the chip's maximum utilization area. This method and steps achieve a reasonable division of standardized units while fully considering chip utilization.
[0091] In one embodiment, a division result that meets preset requirements among multiple division results is used as an initial division result, including: a division result that is determined to meet the optimal cutting requirement and the minimum imbalance requirement among multiple division results is used as the initial division result; if multiple division results cannot simultaneously meet the optimal cutting requirement and the minimum imbalance requirement, then the division result that meets the minimum imbalance requirement is selected as the initial division result.
[0092] Among them, the minimum imbalance requirement can be that the number of standardized units meets the minimum imbalance requirement, that is, in the division result, the difference in the number of standardized units on different chips is the smallest, for example, the difference in the number of standardized units on the top chip and the bottom chip is the smallest.
[0093] The minimum imbalance requirement may be that the standardized unit meets the minimum imbalance requirement in chip utilization, that is, in the partitioning result, under the premise that the chip utilization is lower than the preset chip utilization, the difference between the utilizations of different chips is minimum.
[0094] In this embodiment, by selecting a division result that meets preset requirements from multiple division results, the impact of the number of standardized units and chip utilization on the chip division result can be fully considered, providing a basis for subsequent placement of standardized units.
[0095] In one embodiment, the contracted nodes in the hypergraph are deactivated according to the initial partitioning results. During the deactivation process, the gain function designed for the 3D placement scenario is used to estimate the change in the minimum half-circle length after each move. Based on the change in the minimum half-circle length, the standardized units during node deactivation are optimized to obtain the final partitioned netlist, as shown in FIG. Figure 6 The flowchart of the method for generating a partitioned netlist shown includes:
[0096] S602: Perform multiple node cancellations on the contracted nodes in the hypergraph according to the initial partitioning result.
[0097] Among them, according to the initial partitioning result, multiple node cancellations are performed on the contracted nodes in the hypergraph to obtain a partitioned netlist after multiple node cancellations.
[0098] Among them, multiple node cancellations are performed on the shrunk nodes to obtain a partitioned netlist after multiple node cancellations. At this time, the nodes of the partitioned netlist need to be moved to obtain the final partitioned netlist. However, during the movement process, the total area of the nodes after the movement may be greater than the preset maximum utilization area of the chip. Therefore, it is necessary to compare the total area of the nodes of the partitioned netlist after the nodes are moved with the preset maximum utilization area of the chip.
[0099] S604: In the process of canceling the contracted nodes, the nodes of the netlist in the initial partitioning result are tuned and exchanged using a gain function designed for the 3D placement scenario. During each tuning process, it is necessary to ensure that the total area of the corresponding netlist nodes after tuning is less than the preset maximum chip utilization area.
[0100] The nodes of the netlist in the initial partitioning result are optimized and swapped using a gain function designed for 3D placement scenarios. If the total node area of the partitioned netlist after node optimization and swapping is greater than the preset maximum chip utilization area, the partitioned netlist generated after the node removal is filtered out. If the total node area of the partitioned netlist after node optimization and swapping is less than or equal to the preset maximum chip utilization area, the move is retained to obtain the partitioned netlist.
[0101] In the process of undoing the contraction node, the normalized unit is tuned by a gain function designed for the 3D placement scenario.
[0102] Specifically, the gain function is calculated as follows:
[0103] First, the expression of the gain function designed for the 3D placement scene is:
[0104] hpwl_gain=∑ e∈I(c)(hpwl_simulate(e∩v)-hpwl_simulate(e∩vc)+hpwl_simulate(e∩b)-hpwl_simulate(e∩b+c));
[0105] Wherein, hpwl_gain represents the gain value of the gain function, hpwl_simulate represents the calculation formula of the half-circle length, v and b represent different chips, and e represents the total number of standardized units connected to the standardized unit c.
[0106] Specifically, this function represents the gain function of moving normalization unit c from chip v to chip b. Considering that the movement of normalization unit c will affect the connection line of each normalization unit containing normalization unit c, the calculation of the gain function designed for the 3D placement scenario needs to take into account the connection line of each normalization unit related to normalization unit c. In the formula, I(c) represents the set of all associated nets of normalization unit c.
[0107] It can be understood that the standardized unit c is located on chip v before the movement and on chip b after the movement. Among them, hpwl_simulate(e∩v) represents the half-circle length of the network net of chip v before the movement of the standardized unit c, hpwl_simulate(e∩vc) represents the half-circle length of the network net of chip v after the movement of the standardized unit c, hpwl_simulate(e∩b) represents the half-circle length of the network net of chip b before the movement of the standardized unit c, and hpwl_simulate(e∩b+c) represents the half-circle length of the network net of chip b after the movement of the standardized unit c.
[0108] The hpwl_simulate(V) function represents the calculation of the gain function of the standardized unit set V.
[0109] Specifically, hpwl_simulate(V) is obtained in the following way:
[0110] To calculate hpwl, we first need to calculate the height of the normalized unit set V:
[0111]
[0112] In the formula It means that, ideally, all standardized units in set V are considered a square, and then the square root is taken to get the side length of this ideal square. The reason why set V is considered a square here is that under the same area, the side length of the circle is the smallest, followed by the square. In the placement model of the D2D vertical link scenario, all standardized units are arranged in rows, so it is impossible to form a circle with the standardized unit set, so it is formed into a square. Then divide it by the row height and round up to get the number of occupied rows. This is because in the actual placement process, standardized units are placed in rows, and the number of occupied rows needs to be an integer, and then multiply it by the row height to get the total height of set V. Then the length of set V can be calculated by the total height of set V, and the expression is as follows:
[0113]
[0114] Then, the estimated value of the minimum half-circle length of the standardized unit set V under ideal conditions can be calculated through the calculation formula of hpw1:
[0115] hpwl_simulate(V)=hpwl_width_simulate(V)+hpwl_length_simulate(V)
[0116] In summary, if the gain value hpwl_gain of the gain function is less than zero, it means that moving the standardized unit reduces the overall semi-circular length of the chip. If the gain value of the gain function is less than zero, and the absolute value of the gain value of the gain function is larger, it means that moving the standardized unit reduces the overall semi-circular length of the chip and the chip's power consumption is lower. Similarly, if the gain value hpwl_gain of the gain function is greater than zero, it means that moving the standardized unit increases the overall semi-circular length of the chip. If the gain value of the gain function is greater than zero, and the absolute value of the gain value of the gain function is larger, it means that moving the standardized unit increases the overall semi-circular length of the chip and the chip's power consumption is higher.
[0117] S606: After all the contracted nodes are cancelled and optimized, the final partitioned netlist is obtained.
[0118] The smallest half-circle length of the standardized unit in the calculation result, that is, the movement with the largest gain value of the gain function, is obtained according to the tuning process of the standardized unit to obtain the final partitioned netlist.
[0119] In this embodiment, by calculating the gain function designed for the 3D placement scenario of the standardized unit, the half-circle lengths of the standardized unit before and after the movement are compared, and the movement with the best gain value of the gain function is used as the tuning result of the standardized unit to obtain the most reasonable final partitioned netlist, which can provide a reasonable partitioned netlist for subsequent placement and improve the subsequent placement effect.
[0120] In one embodiment, the partitioned netlist is input into the placement model to obtain a 3D placement simulation result, including: placing standardized cells and hybrid bonding terminals according to the partitioned netlist based on the constraints of the placement model and the standardized cells of the placement model to obtain a 3D placement simulation result.
[0121] The constraints for placing the standardized cells of the model include: the standardized cells are within the row height range, the standardized cells do not overlap with each other, and the standardized cells meet the chip utilization requirements after placement.
[0122] Specifically, the partitioned netlist determines whether the connection lines between the standardized cells are on the same plane. If so, the standardized cells are placed according to the partitioned netlist to obtain a placement simulation result. If the connection lines between the standardized cells are not on the same plane, the standardized cells and hybrid bonding terminals are placed according to the partitioned netlist to obtain a placement simulation result.
[0123] In this embodiment, when the connection lines of the standardized units are not on the same plane, the hybrid bonding terminals and the standardized units are placed to obtain 3D simulation results, thereby enhancing the correlation between the hybrid bonding terminals and the standardized units.
[0124] In one embodiment, based on the placement model and the constraints of the standardized cells of the placement model, the standardized cells and the hybrid bonding terminals are placed according to the partitioned netlist to obtain a 3D placement simulation result, including: dividing multiple standardized cells into upper-level standardized cells and lower-level standardized cells according to the partitioned netlist; based on the placement model and the constraints of the standardized cells of the placement model, the upper-level standardized cells, the lower-level standardized cells and the hybrid bonding terminals are placed to obtain a 3D placement simulation result.
[0125] Among them, Figure 7 The schematic diagram of the placement simulation results shown includes: the bottom chip and the lower standardized cells in the bottom chip, the top chip and the upper standardized cells in the top chip, and the hybrid bonding terminals running through the bottom chip and the top chip.
[0126] The upper standardized unit and the lower standardized unit also have connecting wires, which connect the standardized units to the hybrid bonding terminals. The connecting wires are usually rectangular, and the perimeter of the rectangle is generally the half perimeter length HPWL.
[0127] In this embodiment, by placing the upper-layer standardized units, the lower-layer standardized units and the hybrid bonding terminals, a placement simulation result is obtained, and the position of the hybrid bonding terminals is taken into consideration. This can enhance the correlation between the placement of the upper and lower layers, thereby reducing the minimum half-circle length of the connection line across the upper and lower layers, and further reducing the total minimum half-circle length, thereby reducing the power consumption of the chip.
[0128] In one embodiment, based on the placement model and the constraints of the standardized cells of the placement model, the upper standardized cells, the lower standardized cells and the hybrid bonding terminals are placed to obtain a placement simulation result, such as Figure 8 The flowchart of the method for generating placement simulation results shown in FIG. 1 further includes:
[0129] S802: Acquire the initial position of the hybrid bonding terminal.
[0130] The random placement is performed based on the chip center, and the random placement position is the initial position of the hybrid bonding terminal.
[0131] Specifically, the process of randomly initializing the hybrid bonding terminal is divided into initialization of the x-axis coordinate and initialization of the y-axis. Both initialization parts use the Gaussian distribution method. The distribution formula of the hybrid bonding terminal coordinate on the x-axis is as follows:
[0132] X~N(μ,σ 2 );
[0133] Where X represents the random variable of the hybrid bond terminal on the x-axis, μ represents the mean of the probability distribution, which corresponds to the center of the distribution in this case, and σ represents the standard deviation of the probability distribution, which corresponds to the width of the distribution.
[0134] The mean μ and standard deviation σ of the probability distribution are expressed as follows:
[0135]
[0136] σ=(xh-xl)×0.001;
[0137] Where xl represents the coordinate of the left edge of the chip on the x-axis, and xh represents the coordinate of the right edge of the chip on the x-axis.
[0138] Similarly, the distribution of the hybrid bonding terminal coordinates on the y-axis follows the Gaussian distribution Y~N(μ,σ 2 ) performs an initialization operation, thereby obtaining the initial distribution of the hybrid bonding terminals on the x-axis and y-axis, that is, the initial position of the hybrid bonding terminals on the chip.
[0139] S804 , placing the upper layer standardized cells and the lower layer standardized cells based on the placement model and the standardized cell constraints of the placement model to obtain a first placement result.
[0140] The final partitioned netlist is converted into bookshelf file format and used as Dreamplace input to generate upper-level netlist files and lower-level netlist files.
[0141] A first placement result is generated according to the upper-layer netlist file and the lower-layer netlist file, based on the placement model and the standardized cell constraints of the placement model.
[0142] Specifically, the placement of the upper and lower chips is treated as two sub-problems, and then two sub-processes are created for each sub-problem. Each sub-process uses half of the system's computing power to calculate the placement position. In each sub-process, the hybrid bonding terminals and standardized units are input into the Dreamplace placement framework, and the corresponding placement results are output. Finally, the results are returned to the main process to obtain the first placement result.
[0143] S806 : Re-place the hybrid bonding terminal based on the first placement result to obtain an updated hybrid bonding terminal position.
[0144] The hybrid bonding terminals are also taken into consideration and the first placement result is re-placed. Specifically, the hybrid bonding terminals are treated as standardized units and placed in the placement model to obtain updated hybrid bonding terminal positions.
[0145] S808 : Replace the initial position of the hybrid bonding terminal with the updated hybrid bonding terminal position.
[0146] S810 , repeatedly performing the step of placing the upper standardized unit, the lower standardized unit, and the hybrid bonding terminal a preset number of times to obtain a plurality of second placement results.
[0147] S812 , selecting the second placement result in which the half-circle length of the upper standardized unit, the lower standardized unit, and the hybrid bonding terminal is the smallest among the second placement results as the 3D placement simulation result.
[0148] The second placement result may contain overlapping and cross-row placement.
[0149] Specifically, when overlap and cross-row placement occur in the second placement result, a greedy strategy is used to move the positions of these standardized units to make the result conform to the specification. First, the x-axis coordinates of the lower left corners of all standardized units on the chip after global placement are recorded through array X. Then, array X is sorted from small to large. Then, the standardized units are moved according to the sorting order, with vertical movement being prioritized to place the standardized units on the row. If overlap occurs, horizontal movement is performed. Each movement adopts the following strategy:
[0150] min[(x(i)-x ′ (i)) 2 +(y(i)-y ′ (i)) 2 ]
[0151] Among them, x(i) represents the x-axis coordinate before moving, x ′ (i) represents the x-axis coordinate after the move. This strategy describes the distance between the global position and the legal position. Each move is performed with the shortest straight-line distance in the current state. The move needs to meet the following constraints:
[0152] x(i)-x(i-1)≥w(i-1), i≥2
[0153] The w in the above formula represents the width of the normalized unit. This constraint is used to determine whether there is any overlap in the results after the movement.
[0154] The standardized cells are continuously moved through the above greedy strategy until all standardized cells meet the legalization requirements and the positions of the standardized cells and hybrid bonding terminals are obtained.
[0155] Among them, the half-circumference lengths of the upper standardized unit, the lower standardized unit and the hybrid bonding terminal can be calculated according to the positions of the standardized unit and the hybrid bonding terminal.
[0156] Specifically, the normalized unit position ( i ,y i ), calculate the minimum half-circle length of the connection line of a standardized unit on this chip as:
[0157] hpwlL e =max i∈e {x i}-min i∈e {x i}+max i∈e {y i}-min i∈e {y i}
[0158] Next, we can calculate the minimum half-circle length of this chip:
[0159]
[0160] Therefore, the minimum semi-circular length of the upper and lower dies is added together to obtain the minimum semi-circular length of the 3D model. The specific formula is as follows:
[0161] hpwl 3D =hpwl top +hpwl bottom
[0162] Among them, h top represents the connection line length of all standardized cells and hybrid bonding terminals in the top chip, h bottom Represents the connection wire length of all standardized cells and hybrid bond terminals in the bottom chip.
[0163] In this embodiment, the placement of the upper and lower chips is continuously optimized by adjusting and updating the positions of the hybrid bonding terminals after each placement round. Furthermore, a parallel placement method using shared memory is used during placement, saving significant time. This ensures accurate chip placement while also reducing simulation time.
[0164] In one embodiment, Figure 9 A placement method with D2D vertical chain connection based on Dreamplace is provided, including:
[0165] S902 , preprocessing the parameters of the netlist to be placed and the standardized units in the netlist to obtain a hypergraph with the standardized units as nodes and the connecting lines between the standardized units as hyperedges.
[0166] S904: Obtain the contraction score of each node and neighboring node according to the weight of each node and neighboring node in the hypergraph.
[0167] S906 , determining the node with the smallest area according to the contraction score of the node and its neighboring nodes, and contracting the node with the smallest area.
[0168] S908, starting from a randomly selected vertex, the shrunken hypergraph is traversed using BFS to partition it and obtain the partition result.
[0169] S910: Update the shrunken node weights according to the shrunken normalized units.
[0170] S912, record the number of divisions.
[0171] S914, determine whether the number of divisions reaches the preset number, if so, execute S916, if not, execute S904.
[0172] S916: The partitioning result that satisfies both the optimal cutting requirement and the minimum imbalance requirement is selected as the initial partitioning result. If none of the multiple partitioning results satisfies both the optimal cutting requirement and the minimum imbalance requirement, the partitioning result that satisfies the minimum imbalance requirement is selected as the initial partitioning result.
[0173] S918: Perform multiple node cancellations on the contracted nodes in the hypergraph according to the initial partitioning result.
[0174] S920: In the process of canceling the contracted nodes, the nodes of the netlist in the initial partitioning result are tuned and exchanged using a gain function designed for the 3D placement scenario. During each tuning process, it is necessary to ensure that the total area of the corresponding netlist nodes after tuning is less than the preset maximum chip utilization area.
[0175] S922: After all the contracted nodes are cancelled and optimized, the final partitioned netlist is obtained.
[0176] S924 , dividing the plurality of standardized units into upper-layer standardized units and lower-layer standardized units according to the partitioned netlist.
[0177] S926 , obtaining the initial position of the hybrid bonding terminal.
[0178] Among them, the initial positions of the hybrid bonding terminals can be generated according to the partitioned netlist.
[0179] S928 , placing the upper layer standardized cells and the lower layer standardized cells based on the placement model and the standardized cell constraints of the placement model to obtain a first placement result.
[0180] S930 : Re-place the hybrid bonding terminal based on the first placement result to obtain an updated hybrid bonding terminal position.
[0181] S932: Replace the initial position of the hybrid bonding terminal with the updated hybrid bonding terminal position.
[0182] S934 , repeating the step of placing the upper standardized unit, the lower standardized unit, and the hybrid bonding terminal for a preset number of times to obtain a plurality of second placement results.
[0183] S936 , selecting the second placement result in which the half-circle length of the upper standardized unit, the lower standardized unit, and the hybrid bonding terminal is the smallest among the second placement results as the 3D placement simulation result.
[0184] Specifically, the half-circle lengths calculated by multiple tests were compared with those calculated by other division and placement methods. The comparison results are shown in Table 1:
[0185] Table 1 Minimum half-circle length calculated by three division and placement methods
[0186] Partition placement method Case 1 Case 2 Case 3 Hmeits division + placement 2576930 43432394 373748369 Kahapar division + placement 2589842 43337481 372575321 This method 2543763 42786578 365142135
[0187] Among them, Case 1 is the division and placement of 2,735 standardized units, Case 2 is the division and placement of 44,764 standardized units, and Case 3 is the division and placement of 220,845 standardized units.
[0188] Hmeits division + placement can be Hmeits division + Dreamplace placement, and Kahapar division + placement can be Kahapar division + Dreamplace placement.
[0189] Compared with the Hmeits partitioning and placement method and the Kahapar partitioning and placement method, this method achieves the lowest final HPWL (minimum half-perimeter length) for standardized units of different orders of magnitude in all three cases. It should be noted that the results in the table may be the average of multiple test results, for example, the average of 20 tests is used as the specific value in the table.
[0190] Compared with the Hmeits partition + placement method, this method reduces HPWL by 1.29%, 1.24%, and 2.21% on the datasets of cases 1, 2, and 3, respectively. Compared with the Kahapar partition + placement method, this method reduces HPWL by 1.78%, 1.49%, and 1.99% on the datasets of cases 1, 2, and 3, respectively.
[0191] It can be seen that the gain function h_ designed for the 3D placement scenario proposed by this method can better optimize the division and placement results and better adapt to the 3D division and placement scenario.
[0192] In this embodiment, the nodes in the hypergraph are hierarchically contracted and divided based on the node weight of the hypergraph. By considering the order of the area size of the standardized units, the chip utilization rate is fully improved to obtain the initial division result. The gain function designed for the 3D placement scenario is used to adjust the node cancellation process of the initial division result to obtain the final division netlist. The division netlist is input into the placement model to obtain the 3D placement simulation result. On the one hand, this method considers the utilization rate of the chip where the standardized unit is located. On the other hand, it makes the half-circle length of the standardized unit as short as possible, obtains a reasonable division netlist, improves the placement effect of the placement simulation result, and reduces the overall power consumption of the chip. This method, (1) takes the area of the standardized unit into consideration in the calculation of the gain function for the placement scenario, which can effectively prevent the total area of the standardized units divided on a single chip from being too large, thereby exceeding the effective utilization area of the chip. At the same time, the standardized unit areas on the two chips are balanced, so that it can well adapt to the D2D vertical connection 3D placement model, thereby affecting the minimum half-circle length of the final placement. (2) The position of the hybrid interlayer terminals is taken into account when placing the upper and lower layers. This can enhance the correlation between the upper and lower layers, thereby reducing the minimum half-circle length of the connecting line across the upper and lower layers, and thus reducing the total minimum half-circle length. At the same time, the position of the hybrid interlayer terminals is adjusted and updated after each round of placement, so that the placement results of the upper and lower layers can be continuously optimized. In addition, a parallel placement method with shared memory is used during placement, which saves a lot of time. (3) It is universal and can be applied to 3D placement tasks in various situations.
[0193] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0194] Based on the same inventive concept, embodiments of the present application also provide a device for simulating chip partitioning and placement in a 3D scene, for implementing the aforementioned method for simulating chip partitioning and placement in a 3D scene. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more embodiments of the device for simulating chip partitioning and placement in a 3D scene provided below can be found in the aforementioned limitations of the method for simulating chip partitioning and placement in a 3D scene, and will not be further elaborated here.
[0195] In one embodiment, Figure 10 As shown, a chip partitioning and placement simulation device in a 3D scene is provided, comprising: a pre-processing module 1002, a processing module 1004, a partitioning module 1006 and a placement module 1008, wherein:
[0196] A preprocessing module 1002 is used to preprocess the parameters of the netlist to be placed and the standardized units in the netlist to obtain a hypergraph with the standardized units as nodes and the connecting lines between the standardized units as hyperedges;
[0197] The processing module 1004 is used to perform node hierarchical contraction and partitioning on the nodes in the hypergraph based on the node weights of the hypergraph to obtain an initial partitioning result;
[0198] A partitioning module 1006 is configured to cancel the contracted nodes in the hypergraph according to the initial partitioning result. During the cancellation process, the gain function designed for the 3D placement scenario is used to estimate the change in the minimum semi-circle length after each movement. Based on the change in the minimum semi-circle length, the standardized units used during the node cancellation are optimized to obtain a final partitioned netlist.
[0199] The placement module 1008 is used to input the partitioned netlist into the placement model to obtain a 3D placement simulation result.
[0200] In one embodiment, the processing module 1004 is further used to obtain the shrinkage score of the node and the neighboring nodes according to the weight of each node and the neighboring nodes in the hypergraph; determine the node with the smallest shrinkage area according to the shrinkage score of the node and the neighboring nodes, and shrink the node with the smallest shrinkage area; start from a randomly selected vertex of the shrunken hypergraph, use BFS traversal to divide it, and obtain the current division result; update the weight of the shrunken node according to the standardized unit after shrinkage; record the number of divisions; return to the step of obtaining the shrinkage score of the node and the neighboring nodes according to the weight of each node and the neighboring nodes in the hypergraph, and obtain multiple division results until the number of divisions reaches a preset number; and use the division result that meets the preset requirements among the multiple division results as the initial division result.
[0201] In one embodiment, the processing module 1004 is further used to use the division result that satisfies the optimal cutting requirement and the minimum imbalance requirement among multiple division results as the initial division result; if multiple division results cannot simultaneously meet the optimal cutting requirement and the minimum imbalance requirement, the division result that meets the minimum imbalance requirement is selected as the initial division result.
[0202] In one embodiment, the partitioning module 1006 is also used to perform multiple node cancellations on the contraction nodes in the hypergraph according to the initial partitioning results; in the process of canceling the contraction nodes, the nodes of the netlist in the initial partitioning result are tuned and exchanged through the gain function designed for the 3D placement scenario, and in each tuning process, it is necessary to ensure that the total area of the corresponding netlist nodes after tuning is less than the preset maximum chip utilization area; after all the contraction nodes are cancelled and tuned, the final partitioned netlist is obtained.
[0203] In one embodiment, the placement module 1008 is further configured to place the standardized cells and hybrid bonding terminals based on the placement model and the constraints of the standardized cells of the placement model and according to the partitioned netlist to obtain a 3D placement simulation result.
[0204] In one embodiment, the placement module 1008 is further used to divide multiple standardized cells into upper-level standardized cells and lower-level standardized cells according to the partitioned netlist; based on the placement model and the standardized cell constraints of the placement model, the upper-level standardized cells, the lower-level standardized cells and the hybrid bonding terminals are placed to obtain 3D placement simulation results.
[0205] In one embodiment, the placement module 1008 is also used to obtain the initial position of the hybrid bonding terminal; based on the placement model and the standardized unit constraints of the placement model, the upper standardized unit and the lower standardized unit are placed to obtain a first placement result; according to the first placement result and the initial position of the hybrid bonding terminal, the hybrid bonding terminal is re-placed to obtain an updated hybrid bonding terminal position; the updated hybrid bonding terminal position replaces the initial position of the hybrid bonding terminal; repeat the steps of placing the upper standardized unit, the lower standardized unit and the hybrid bonding terminal a preset number of times to obtain multiple second placement results; select the second placement result with the smallest semi-circumference length of the upper standardized unit, the lower standardized unit and the hybrid bonding terminal in the second placement result as the 3D placement simulation result.
[0206] Each module in the aforementioned chip partitioning and placement simulation device in a 3D scene can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0207] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 11 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store netlist data to be placed. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a chip partitioning and placement simulation method in a 3D scene is implemented.
[0208] Those skilled in the art will understand that Figure 11 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0209] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0210] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0211] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0212] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0213] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.
[0214] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0215] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A chip partitioning and placement simulation method in a 3D scene, characterized in that: The method comprises: Preprocessing the parameters of the netlist to be placed and the standardized units in the netlist to obtain a hypergraph with the standardized units as nodes and the connecting lines between the standardized units as hyperedges; Based on the node weights of the hypergraph, performing node hierarchical contraction and division on the nodes in the hypergraph to obtain an initial division result; Deleting the contracted nodes in the hypergraph according to the initial partitioning result, and during the deleting process, estimating the change in the minimum semi-circle length after each movement using a gain function designed for a 3D placement scenario, and optimizing the standardized units during the node deleting process based on the change in the minimum semi-circle length to obtain a final partitioned netlist; The partitioned netlist is input into a placement model to obtain a 3D placement simulation result.
2. The method according to claim 1, characterized in that The step of performing node hierarchical contraction and partitioning on the nodes in the hypergraph based on the node weights of the hypergraph to obtain an initial partitioning result includes: Obtaining a node and neighbor node contraction score according to the weights of each node and neighbor node in the hypergraph; Determine the node with the smallest area according to the contraction score of the node and its neighboring nodes, and contract the node with the smallest area; Starting from a randomly selected vertex, the shrunk hypergraph is traversed using BFS to obtain the partitioning result. updating the shrunk node weights according to the shrunk normalized unit; Record the number of divisions; Returning to the step of obtaining a contraction score of the node and its neighboring nodes according to the weights of each node and its neighboring nodes in the hypergraph, obtaining multiple division results until the number of divisions reaches a preset number; The division result that meets the preset requirements among the multiple division results is used as the initial division result.
3. The method according to claim 2, characterized in that The division result that meets the preset requirements among the multiple division results is used as the initial division result, including: Determining a division result that meets the optimal cutting requirement and the minimum imbalance requirement from the multiple division results as an initial division result; If the plurality of division results cannot simultaneously meet the optimal cutting requirement and the minimum imbalance requirement, the division result that meets the minimum imbalance requirement is selected as the initial division result.
4. The method according to claim 1, wherein The contracted nodes in the hypergraph are cancelled according to the initial partitioning result, and during the cancellation process, a gain function designed for a 3D placement scenario is used to estimate the change in the minimum semi-circle length after each movement, and based on the change in the minimum semi-circle length, the standardized units during the node cancellation are tuned to obtain a final partitioned netlist, including: Performing multiple node cancellations on the contracted nodes in the hypergraph according to the initial partitioning result; During the process of undoing the shrinking nodes, the nodes in the netlist of the initial partitioning result are tuned and exchanged using a gain function designed for 3D placement scenarios. During each tuning process, it is necessary to ensure that the total area of the corresponding netlist nodes after tuning is less than the preset maximum chip utilization area. After all the contraction nodes are cancelled and optimized, the final partitioned netlist is obtained.
5. The method according to claim 1, wherein Inputting the partitioned netlist into a placement model to obtain a 3D placement simulation result includes: Based on the placement model and the constraints of the standardized cells of the placement model, and according to the partitioned netlist, the standardized cells and hybrid bonding terminals are placed to obtain a 3D placement simulation result.
6. The method according to claim 5, characterized in that The placing of the standardized cells and the hybrid bonding terminals based on the placement model and the constraints of the standardized cells of the placement model according to the partitioned netlist to obtain a 3D placement simulation result includes: Dividing the plurality of standardized units into upper-layer standardized units and lower-layer standardized units according to the partitioned netlist; Based on the placement model and the standardized cell constraints of the placement model, upper-layer standardized cells, lower-layer standardized cells, and hybrid bonding terminals are placed to obtain a 3D placement simulation result.
7. The method according to claim 6, characterized in that The method further comprises placing upper standardized cells, lower standardized cells, and hybrid bonding terminals based on the placement model and the constraints of the standardized cells of the placement model to obtain a 3D placement simulation result. Obtaining an initial position of the hybrid bonding terminal; Based on the placement model and the standardized unit constraints of the placement model, placing the upper layer standardized units and the lower layer standardized units to obtain a first placement result; Re-place the hybrid bonding terminal based on the first placement result to obtain an updated hybrid bonding terminal position; Replacing the initial position of the hybrid bonding terminal with the updated hybrid bonding terminal position; Repeating the steps of placing the upper standardized unit, the lower standardized unit, and the hybrid bonding terminal a preset number of times to obtain a plurality of second placement results; The second placement result in which the half-circle length of the upper standardized unit, the lower standardized unit, and the hybrid bonding terminal is the smallest among the second placement results is selected as the 3D placement simulation result.
8. A chip partitioning and placement simulation device in a 3D scene, characterized in that: The device comprises: A preprocessing module, configured to preprocess the parameters of the netlist to be placed and the standardized units in the netlist to obtain a hypergraph having the standardized units as nodes and the connecting lines between the standardized units as hyperedges; A processing module, configured to perform node hierarchical contraction and partitioning on the nodes in the hypergraph based on the node weights of the hypergraph to obtain an initial partitioning result; a partitioning module configured to cancel the contracted nodes in the hypergraph according to the initial partitioning result, and during the cancellation process, estimate the change in the minimum semi-circle length after each movement based on a gain function designed for a 3D placement scenario in an optimization algorithm, and optimize the standardized units during the node cancellation based on the change in the minimum semi-circle length to obtain a final partitioned netlist with a gain function designed for a 3D placement scenario; A placement module is used to input the partitioned netlist into a placement model to obtain a 3D placement simulation result.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
A layer allocation method applied to a three-dimensional integrated circuit
CN109033580A
Method for balancing interconnection number between different partitions of circuit and readable storage medium
CN113221501A