Line length and line time delay collaborative prediction method after digital integrated circuit layout
By adopting the line length and line delay collaborative prediction method based on graph neural network and multi-task learning in digital integrated circuit design, the problem of difficulty in accurately predicting line length and delay after layout in the logical synthesis stage is solved, and a more efficient design process and better design quality are achieved.
Patent Information
- Application Number
- CN202510426423.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The prior art is difficult to accurately predict the line length and delay after layout in the logical synthesis stage of digital integrated circuits, resulting in increased resource consumption and affecting design efficiency.
The collaborative prediction method of line length and line delay based on graph neural network and multi-task learning is adopted. By constructing heterogeneous graphs, node and edge features are extracted, and feature interaction and task sharing are used for feature interaction and task sharing by using TransformerConv layer and cross-stitch units, a collaborative prediction model is built.
This method can accurately predict the line length and delay after layout in the logical synthesis stage, reduce the number of design iterations, improve design efficiency and quality, and reduce the time cost of training and reasoning.
Smart Images

Figure CN120012669A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital integrated circuit electronic design automation, and in particular to a method for collaboratively predicting line length and line delay after digital integrated circuit layout. Background Art
[0002] The design process of digital integrated circuits includes multiple links, each of which has an important impact on the performance of the final design. Among them, logic synthesis and layout are indispensable and important stages in the design process. Logic synthesis is to convert hardware description language (such as Verilog or VHDL) into a gate-level netlist, laying the foundation for subsequent physical design. Layout focuses on placing circuit components reasonably on the chip to ensure that wiring can proceed smoothly.
[0003] In both stages, timing analysis is an important step, and its purpose is to verify whether the design meets the established timing requirements. The core task of timing analysis is to calculate the propagation time of the signal in the circuit, that is, the delay. Delay refers to the time required for a signal to be transmitted from one point to another, which directly affects the operating frequency and response speed of the circuit. The influence of line length cannot be ignored in the calculation of delay. Line length refers to the length of the wire connecting two or more circuit elements. Since the transmission of the signal on the line is not instantaneous, but takes a certain amount of time to cross these distances, this generates additional delay caused by the wire. Longer wires increase resistance and capacitance effects, resulting in greater delays and destroying the originally expected timing characteristics.
[0004] In the logic synthesis stage, due to the lack of detailed physical location information, timing analysis usually predicts line length based on idealized assumptions or some estimation models, but this approach makes it difficult to accurately estimate the additional delay caused by the actual line length. When the design enters the layout stage, as the specific location of the circuit elements is determined, the impact of the line length becomes clear, which will change the originally estimated delay value and thus affect the overall timing performance. This deviation often leads to multiple rounds of iterations to adjust the design constraints and rearrange the layout to meet strict timing requirements. This process not only prolongs the design cycle, but also increases resource consumption, greatly affecting design efficiency. Therefore, how to accurately predict the line length and even the delay of the subsequent stages (such as layout) in the early stages of the design (such as logic synthesis) has become the key to improving design efficiency and quality.
[0005] The document (Xie Z, Liang R, Xu X, et al. Preplacement net length and timing estimation by customized graph neural network [J]. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, 2022, 41 (11): 4667-4680.) is the most similar existing solution to the present invention. In view of the difficulty in predicting the delay after layout due to the lack of wire length information in the logic synthesis stage, this study innovatively proposed a two-step prediction process: first, the wire length is predicted by a graph neural network, and then these prediction results are input as features into the random forest model to achieve the prediction of unit delay and line delay. Specifically, this study models each wire net in the design as a node in the graph, combines pre-defined node features and uses graph neural networks to complete the wire length prediction task, and then constructs two independent prediction models based on random forests for predicting unit delay and line delay. The prediction model uses the wire length prediction value as an input feature. The experimental results show that this method can significantly improve the accuracy of delay prediction. Although this paper has some innovations in methods and has improved prediction accuracy, it still has some limitations and shortcomings:
[0006] Limitations of single-task models: This paper uses a single-task model to handle the problems of line length and delay prediction separately. The single-task model is relatively limited in capturing the correlation between tasks and cannot effectively share features and information. Although the two-step method in the paper attempts to partially consider the correlation between the two by first predicting the line length and then predicting the delay, and to a certain extent considers the impact of line length on delay, this approach is slightly insufficient because the models used for the two tasks are still processed independently and fail to fully capture the complex interactive relationship between line length and delay. This separate processing method will limit the model's comprehensive use of related features, which may miss some important information and affect the accuracy of the final prediction.
[0007] Speed issue: The two-step prediction process in this paper may lead to longer training and reasoning time. It is necessary to train two independent models in sequence - one for line length prediction and the other for delay prediction. In addition, before performing delay prediction, each feature used for delay prediction in the design needs to be added with the corresponding line length prediction value one by one, whether it is training or reasoning, before performing delay prediction. This two-step method is not only cumbersome, but may also cause significant time overhead in practical applications, especially when facing large-scale circuit design, the efficiency problem is particularly prominent. Summary of the invention
[0008] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for collaboratively predicting line length and line delay after digital integrated circuit layout.
[0009] To achieve the above purpose, the technical solution provided by the present invention is:
[0010] A method for collaboratively predicting line length and line delay after digital integrated circuit layout, comprising:
[0011] Construct a preliminary heterogeneous graph based on the connection information of the circuit netlist;
[0012] Extract features for line length and line delay prediction of line edges from netlist files and process files;
[0013] Add the extracted features to the corresponding nodes and edges of the initial heterogeneous graph to obtain the final heterogeneous graph with connection relationships and features;
[0014] Construct a line length and line delay collaborative prediction model based on graph neural network and multi-task learning;
[0015] Train the line length and line delay collaborative prediction model based on graph neural network and multi-task learning;
[0016] The final version of the heterogeneous graph is input into the trained line length and line delay collaborative prediction model based on graph neural network and multi-task learning, and finally the line length prediction results and line delay prediction results are obtained.
[0017] Further, the features extracted from the netlist file and the process file include node features and edge features;
[0018] Among them, the edge features used for line length prediction of line edges include the driving strength of the source pin unit, the source pin fan-out number, the source pin unit area, the source pin unit pin number, the target pin unit area, and the target pin unit fan-out number;
[0019] The node characteristics used for line delay prediction include transition time, margin, capacitance of input pins, and maximum capacitance of output pins;
[0020] The edge features used for line delay prediction include line network capacitance and line network resistance.
[0021] Furthermore, the line length and line delay collaborative prediction model based on graph neural network and multi-task learning includes three TransformerConv layers, each TransformerConv layer is followed by a cross-stitch unit, and the last cross-stitch unit connects two task branches, both of which include a wired edge embedding generation module and a fully connected layer.
[0022] Furthermore, the working mechanism of the TransformerConv layer includes:
[0023] For each node i in the graph, its input feature is expressed as Where d is the feature dimension and l is the index of the current network layer. First, through the learnable weight matrix Perform linear transformation on node features to obtain query vectors Key Vector Sum value vector
[0024]
[0025] The query vector Used to measure the degree of attention of node i to other nodes; key vector Used to match query vectors of other nodes to calculate attention weights; value vector Contains the actual information of the node, which is used for weighted summation to update the node features;
[0026] In order to utilize the information of edge features in the graph, the TransformerConv layer uses a learnable weight matrix Opposite edge features Perform a linear transformation and map it to the same dimensional space as the node features:
[0027]
[0028] Then, the attention weight coefficient α of node i and its neighbor node j is calculated according to the above transformation results ij :
[0029]
[0030] in Dot product operation q T k measures the similarity between the query vector and the key vector, while the normalization in the denominator ensures that the sum of the attention weights is 1. The introduction of enhances the model's ability to utilize edge feature information;
[0031] Based on the attention weight coefficient α ij , new features of nodes The weighted sum of the value vectors of neighbor nodes is obtained:
[0032]
[0033] in is a learnable weight matrix used to transform the features of the node itself; The node's own characteristic information is retained, while Fusion of neighbor node information, through attention weights Implementing weighted aggregation; edge features Adding it to the aggregation process enhances the model's ability to express the relationship between nodes.
[0034] Furthermore, a cross-stitch unit is set after each TransformerConv layer to perform dynamic node feature interaction during the feature extraction process of the line length prediction task T1 and the line delay prediction task T2. In the first layer of TransformerConv, the node features of the two tasks T1 and T2 are represented as and Where i represents the node index and d is the feature dimension. The cross-stitch unit performs weighted fusion of feature representations of different tasks by learning the linear combination weights between tasks, thereby generating a new feature representation. and
[0035]
[0036] Among them, α 11 , α 12 , α 21 , α 22 are all attention weight coefficients.
[0037] Furthermore, in the edge embedding generation module of each task branch, the embedding representations of the source node and the target node are fused into a new embedding representation, namely the edge embedding, through a splicing operation. This splicing operation retains the complete information of the source node and the target node and captures the relationship between the two.
[0038] Each task branch processes the embedded representation through a fully connected layer, which consists of three consecutive hidden layers. Each layer uses a nonlinear activation function ReLU to enhance the expressiveness of the model and finally outputs the prediction result.
[0039] Furthermore, for the line length and line delay collaborative prediction model based on graph neural network and multi-task learning, the line length prediction task T1 and the line delay prediction task T2 use the mean square error as the single task loss function, which are defined as follows:
[0040]
[0041] Where N is the total number of edges, y wirelength,i and Represent the true length and predicted value of the i-th line edge, y netdelay,i and They represent the actual line delay value and predicted value of the i-th line edge respectively;
[0042] In order to effectively balance the optimization objectives of the two tasks in the multi-task learning framework, the geometric mean loss is used as the joint loss function. Its definition is as follows:
[0043]
[0044] Compared with the prior art, the principles and advantages of this technical solution are as follows:
[0045] This technical solution creates a collaborative prediction model for line length and line delay based on graph neural networks and multi-task learning. The model introduces a multi-task learning framework and considers the collaborative relationship between line length and line delay. In this case, predicting line length and line delay through this model can not only reduce the time cost of training and reasoning and improve prediction efficiency, but also ensure the deep integration and full utilization of information related to line length and line delay, overcoming the problem of relatively independent processing of these two indicators in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the services required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0047] Figure 1 This is a principle flow chart of a method for collaboratively predicting line length and line delay after layout of a digital integrated circuit according to an embodiment of the present invention;
[0048] Figure 2 This is a multi-task learning framework diagram;
[0049] Figure 3 An example diagram for converting a circuit diagram into a heterogeneous graph;
[0050] Figure 4 This is an architecture diagram of a line length and line delay collaborative prediction model based on graph neural network and multi-task learning created in a method for collaboratively predicting line length and line delay after digital integrated circuit layout in an embodiment of the present invention. DETAILED DESCRIPTION
[0051] The present invention will be further described below in conjunction with specific embodiments:
[0052] The present embodiment describes a method for collaboratively predicting line length and line delay after digital integrated circuit layout, which defines the collaborative prediction problem to be solved as a multi-task learning problem, as follows:
[0053] Given two supervised learning tasks TI , I = 1, 2, where task T1 is a graph edge-level regression task for predicting the line length of connections between pins of different units. The line length prediction target of this embodiment is the length of the connection between two pins belonging to different units, because it is directly related to the propagation path of the signal inside the chip, and can finely model the signal propagation characteristics, and can more accurately locate potential problem connections, providing guidance for subsequent optimization. More importantly, this fine-grained selection maintains a consistent granularity with the line delay between pins, thus providing a natural fit for feature sharing in the multi-task learning framework. Task T2 is a graph edge-level regression task for line delay prediction. For these two tasks, the input is a heterogeneous graph with feature information from the logic synthesis stage, and the construction of the heterogeneous graph will be described in detail in the data representation. There is a significant correlation between Task T1 and Task T2, because longer connecting lines generally result in higher resistance and capacitance effects, thereby increasing the signal propagation time. In addition, the relationship between line length and line delay is not a simple linear relationship, but is affected by multiple factors (such as load capacitance and driving capability). Therefore, through a multi-task learning framework (such as Figure 2 The proposed method (as shown in Figure 2) helps to capture this complex nonlinear relationship and helps to improve the prediction accuracy of both line length and line delay.
[0054] For each task T I , assuming that the training data set D I Contains I training samples (i.e. n I Figure), then:
[0055]
[0056] Among them G I,J is a task T I The constructed heterogeneous graph contains the feature information of nodes and edges; Y I,J It is a vector containing the true values of the line length or line delay of all the edges related to the line in the graph.
[0057] Assume φ mtl is a line length and line delay collaborative prediction model based on graph neural network and multi-task learning, θ mtl is the total trainable parameter of the model, then the multi-task regression task is expressed as follows:
[0058] Y=φ mtl (G,θ mtl )
[0059] Where Y=(Y 1 ,Y 2 ), I=1,2 represents the true value set, G=(G 1 ,G 2 ), I=1,2 represents a graph set. Further, let φ1 and φ2 be φ mtl The task-specific branch of θ1 and θ2 are the trainable parameters of the corresponding task, then the above formula is expanded to:
[0060]
[0061] Among them, θ mtl =θ1∪θ2. Let the loss function of multi-task regression be L mtl , calculated as follows:
[0062]
[0063] in and They represent the predicted values of tasks T1 and T2 respectively, L1 and L2 represent the loss functions of the two tasks, Represents a joint function that is used to balance the loss weights of the two tasks. By minimizing the total loss function L mtl The method described in this embodiment can simultaneously and efficiently learn and accurately predict line length and line delay.
[0064] like Figure 1 As shown, the method for collaboratively predicting line length and line delay after digital integrated circuit layout described in this embodiment includes the following steps:
[0065] S1. Construct the initial heterogeneous graph based on the connection information of the circuit netlist;
[0066] A heterogeneous graph is a special graph structure, which contains multiple types of nodes and edges, which enables it to model the diverse relationships in complex systems more finely. Therefore, we choose to build a heterogeneous directed graph G = (V, E c ,E n ) to effectively represent the circuit design. Where V represents the node set consisting of all pins; E c Represents the unit edge set, which is used to describe the connection relationship between the pins inside the same unit; E n Represents a line edge set, which is used to describe the connection relationship between pins of different units. Figure 3 As shown, the pins in the circuit are modeled as nodes in the graph (the terms "node" and "pin" are used interchangeably in the present invention), and different types of edges are used to represent their relationships according to the type of connection between the pins: unit edges (dashed arrows) represent the connection between pins within the same component, while line edges (solid arrows) represent the connection between pins of different components.
[0067] Figure 3(a) shows a circuit consisting of three units U1, U2 and U3 and a net Net1. Each unit contains input and output pins. Taking unit U1 as an example, U1 / A1 and U1 / A2 are input pins, and U1 / ZN is an output pin. Therefore, the total number of pins of unit U1 is 3. Figure 3 In (b), for the line edge U1 / ZN-U2 / A2, U1 / ZN is the source pin, that is, the starting point of the signal, U2 / A2 is the target pin, receiving the signal from the source pin, and U1, as the unit where the source pin U1 / ZN is located, is called the driver unit. Another line edge U1 / ZN-U3 / A shares the same source pin U1 / ZN with the line edge U1 / ZN-U2 / A2, but the target pin is different. In this case, the fan-out number of the source pin U1 / ZN is 2, that is, it is subsequently connected to two pins, and these two line edges belong to the same line net Net1.
[0068] S2, extracting features for line length and line delay prediction of line edges from netlist files and process files;
[0069] In this embodiment, the features extracted for predicting the line length and line delay of line edges include node features and edge features;
[0070] Among them, the edge features used for line length prediction of line edges include the driving strength of the source pin unit, the source pin fan-out number, the source pin unit area, the number of pins of the source pin unit, the target pin unit area, and the target pin unit fan-out number; the node features used for line delay prediction include conversion time, margin, capacitance of the input pin, and maximum capacitance of the output pin; the edge features used for line delay prediction include wire mesh capacitance and wire mesh resistance.
[0071] The following is a detailed description of each feature:
[0072] Drive strength of source pin unit: Drive strength is an important indicator to measure the ability of a unit to effectively drive subsequent loads. In digital circuit design, drive strength directly affects the transmission speed and stability of signals. For high fan-out or long wire nets, it is particularly important to select units with higher drive strength, because high drive strength can reduce the delay caused by excessive load during signal transmission, thereby ensuring that the signal can be reliably transmitted to multiple target pins without affecting performance.
[0073] Source pin fan-out number: The fan-out number directly reflects the number of other pins connected to a pin, and is an important parameter for evaluating circuit complexity and load capacity. In actual circuits, if a pin has a higher fan-out number, it often means that it needs to drive more loads, which not only increases the connection length, but also introduces additional line delay.
[0074] Source pin unit area, source pin unit pin count, target pin unit area, target pin unit pin count: A larger unit area or a larger number of pins may mean that more physical space needs to be allocated to it during layout. This increase in space requirements not only affects the arrangement and layout of the unit, but may also indirectly lead to changes in the length of the connection.
[0075] Target pin unit fan-out number: Similar to the fan-out number of the source pin, the fan-out number of the unit to which the target pin belongs will also affect the length of its subsequent driving network. When it has a larger fan-out number, the subsequent network it drives may be more complex, and this complexity may affect the length of the current predicted network. By providing information about neighboring networks, the model can capture the interactions and potential impacts from these networks.
[0076] Conversion time: refers to the time required for a signal to switch from one logical state (such as low level) to another logical state (such as high level). A longer conversion time means that the signal changes more slowly, which will cause a longer delay in the rising or falling edge of the signal, thereby directly increasing the total delay of signal propagation.
[0077] Margin: Margin is a key indicator in timing analysis. It is used to measure the time margin of a signal under the conditions of meeting the design constraints. Margin represents the difference between the actual signal arrival time and the expected arrival time. If the margin is positive, it means that the signal can arrive on time within the specified time and meet the design requirements; if the margin is negative, it means that there is a risk of timing violation and further optimization is required. This feature can reflect the timing pressure of the path and provide global timing information for the model.
[0078] Capacitance (input pins): The capacitance of an input pin significantly increases the load on the unit driving the pin, which has a significant impact on signal propagation delay. Larger input capacitance will increase the charging and discharging time, making the signal change slower, which in turn causes higher line delay.
[0079] Maximum capacitance (output pin): The maximum capacitance of the output pin limits the maximum load capacitance it can drive and is an important parameter for evaluating the drive capability of the unit. During the design process, if a unit is expected to have a large fan-out number or needs to drive a longer wire net, a unit with a higher maximum capacitance of the output pin is generally preferred to ensure that the signal can be reliably transmitted to all target pins.
[0080] Mesh capacitance and mesh resistance: Mesh capacitance and mesh resistance are the main factors affecting line delay. Mesh capacitance mainly comes from the parasitic effects of metal interconnection and the input capacitance of the pin, while mesh resistance is determined by the material properties and geometric dimensions of the metal conductor. Higher mesh capacitance and resistance will cause the signal to decay faster during transmission, thereby significantly increasing line delay.
[0081] S3. Build a line length and line delay collaborative prediction model based on graph neural network and multi-task learning;
[0082] like Figure 4 As shown, the line length and line delay collaborative prediction model based on graph neural network and multi-task learning includes three TransformerConv layers. A cross-stitch unit is provided after each TransformerConv layer. The last cross-stitch unit connects two task branches. Both task branches include a wired edge embedding generation module and a fully connected layer.
[0083] Among them, the working mechanism of the TransformerConv layer includes:
[0084] For each node i in the graph, its input feature is expressed as Where d is the feature dimension and l is the index of the current network layer. First, through the learnable weight matrix Perform linear transformation on node features to obtain query vectors Key Vector Sum value vector
[0085]
[0086] The query vector Used to measure the degree of attention of node i to other nodes; key vector Used to match query vectors of other nodes to calculate attention weights; value vector Contains the actual information of the node, which is used for weighted summation to update the node features;
[0087] In order to utilize the information of edge features in the graph, the TransformerConv layer uses a learnable weight matrix Opposite edge features Perform a linear transformation and map it to the same dimensional space as the node features:
[0088]
[0089] Then, the attention weight coefficient α of node i and its neighbor node j is calculated according to the above transformation results ij :
[0090]
[0091] in Dot product operation q Tk measures the similarity between the query vector and the key vector, while the normalization in the denominator ensures that the sum of the attention weights is 1. The introduction of enhances the model's ability to utilize edge feature information;
[0092] Based on the attention weight coefficient α ij , new features of nodes The weighted sum of the value vectors of neighbor nodes is obtained:
[0093]
[0094] in is a learnable weight matrix used to transform the features of the node itself; The node's own characteristic information is retained, while Fusion of neighbor node information, through attention weights Implementing weighted aggregation; edge features Adding it to the aggregation process enhances the model's ability to express the relationship between nodes.
[0095] Each TransformerConv layer is followed by a cross-stitch unit, which performs dynamic node feature interaction during the feature extraction process of the line length prediction task T1 and the line delay prediction task T2. In the first layer of TransformerConv, the node features of the two tasks T1 and T2 are represented as and Where i represents the node index and d is the feature dimension. The cross-stitch unit performs weighted fusion of feature representations of different tasks by learning the linear combination weights between tasks, thereby generating a new feature representation. and
[0096]
[0097] Among them, α 11 , α 12 , α 21 , α 22 are all attention weight coefficients.
[0098] In the edge embedding generation module of each task branch, the embedding representations of the source node and the target node are fused into a new embedding representation, namely edge embedding, through a splicing operation. This splicing operation preserves the complete information of the source node and the target node and captures the relationship between the two.
[0099] Each task branch processes the embedded representation through a fully connected layer, which consists of three consecutive hidden layers. Each layer uses a nonlinear activation function ReLU to enhance the expressiveness of the model and finally outputs the prediction result.
[0100] For the line length and line delay collaborative prediction model based on graph neural network and multi-task learning, the line length prediction task T1 and the line delay prediction task T2 use mean square error as the single task loss function, which are defined as follows:
[0101]
[0102] Where N is the total number of edges, y wirelength,i and Represent the true length and predicted value of the i-th line edge, y netdelay,i and They represent the actual line delay value and predicted value of the i-th line edge respectively;
[0103] In order to effectively balance the optimization objectives of the two tasks in the multi-task learning framework, the geometric mean loss is used as the joint loss function. Its definition is as follows:
[0104] S4. Train the line length and line delay collaborative prediction model based on graph neural network and multi-task learning;
[0105] S5. Input the extracted feature data into the trained line length and line delay collaborative prediction model based on graph neural network and multi-task learning, and finally obtain the prediction results of line length and line delay.
[0106] In order to fully verify the effectiveness of the method proposed in this invention, multiple benchmark data sets were selected as experimental data sources, including ISCAS'89, ITC'99, ANUBIS and IWLS'05. On this basis, the original data was screened, and the smaller designs were eliminated, leaving only those more representative and challenging circuit designs. The purpose of this strategy is to ensure that the experimental data can better reflect the true characteristics of complex circuits, thereby enhancing the practical significance and reference value of the experimental results. After screening, 20 circuit designs of different sizes were finally obtained for experimental analysis. The detailed information of these designs is shown in Table 1, of which the 20 designs were randomly divided into 12 for training and 8 for testing. All designs are based on the 45nm NanGate process library, and the logic synthesis and layout process are completed by EDA tools. After the layout is completed, the line length and line delay between pins in each design are extracted by EDA tools as training labels. Among them, the line delay is the maximum value of the signal rising edge delay (Rise Delay) and the falling edge delay (Fall Delay). Both features and labels are normalized to accelerate model convergence.
[0107] Table 1 Detailed information on data set division and design
[0108]
[0109] In the model parameter setting, the number of hidden channels of the TransformerConv layer is set to 64. The design of the fully connected layer is as follows: the input dimension of the first fully connected layer is 64×2 (i.e., the concatenation result of the source node and the target node embedding), and the output dimension is 64; the input dimension of the second fully connected layer is 64, and the output dimension is 32; the input dimension of the third fully connected layer is 32, and the output dimension is 1, which is finally used to predict the regression value of the edge. Considering the large difference in the scale of graph data, each batch contains only one graph during the training process, and the training data will be randomly shuffled during the training process to avoid the model from overfitting to data in a specific order. In addition, an early stopping mechanism is introduced during the training process. Since multi-task learning requires repeated trade-offs between different tasks, the model may require more training rounds to achieve full convergence. For this reason, the patience value is set to 30, that is, when the performance of the validation set does not improve within 30 consecutive rounds, the training process will be terminated early. In terms of the choice of optimizer, the Adam optimizer is used, with an initial learning rate set to 0.005 and a weight decay set to 0.0001. This configuration aims to balance the convergence speed and regularization effect of the model, thereby effectively preventing overfitting and improving generalization ability.
[0110] In the experiment, the methods proposed in two papers (Xie Z, Liang R, Xu X, et al. Preplacement netlength and timing estimation by customized graph neural network [J]. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, 2022, 41 (11): 4667-4680.) and (Guo Z, Liu M, Gu J, et al. A timing engine inspired graph neural network model for pre-routing slack prediction [C] / / Proceedings of the 59th ACM / IEEE Design Automation Conference. 2022: 1207-1212.) are used as baseline models to compare their performance with the model proposed in this invention.
[0111] Xie et al. proposed two models Net 2f and Time fThey are used to predict line length and line delay respectively. 2f It is a model based on Graph Attention Network (GAT). Its core idea is to regard lines as nodes to construct an undirected graph and use GAT to extract high-level features of nodes to achieve line length prediction. 2f A three-layer GAT was used for feature extraction, each layer contained two attention heads, and the hidden layer dimension was 64. In addition, batch normalization was used after each layer of GAT to stabilize the training process. Finally, the node features of all GAT layers were concatenated to form the final node embedding representation, and the line length prediction value was output through two fully connected layers. The input dimension of the first fully connected layer was 3×64×2, corresponding to the dimension of the node embedding, and the output dimension was 64. The input dimension of the second fully connected layer was 64 and the output dimension was 1. The optimizer uses Stochastic Gradient Descent (SGD), the learning rate is set to 0.005, the momentum factor is 0.9, and the weight decay is 0.0001. Time f It is a model based on random forest, whose input features include not only Net 2f The predicted line length also includes the predicted line length features and other line delay related features. The specific configuration of the random forest model is that the number of trees is 80 and the maximum tree depth is 12.
[0112] On the other hand, the model implementation of Guo et al. is based on its open source code, and the core part is its customized graph convolution layer (NetConv). The NetConv layer captures the complex relationship in the wire network structure through two customized operations (graph broadcasting and graph reduction). Graph broadcasting allows information to flow from the source node to the target node. Its operation is to splice the features of the source node, target node and edge, and generate new features of the target node through a fully connected layer; graph reduction allows information to flow back from the target node to the source node. This operation is implemented through two channels, namely the sum channel (summing the features of all target nodes) and the maximum channel (taking the maximum value of the features of all target nodes). In the experiment, a multi-task learning model for line length prediction and line delay prediction was implemented based on NetConv. Three layers of NetConv layers were stacked for each task, and cross-stitch units were inserted between each layer to achieve multi-task learning. Finally, each task outputs the line length prediction results and line delay prediction results respectively through a fully connected layer. The Adam optimizer was used to train the model, with the learning rate set to 0.005 and the weight decay to 0.0001.
[0113] The experimental running environment is a high-performance workstation with hardware configuration including Intel(R) Xeon(R) Platinum 8368 CPU@2.40GHz processor and NVIDIA A100 80GB PCIe GPU. The software environment is based on Ubuntu 20.04LTS operating system and Python 3.9.21 is used as the programming language. In the experiment, the deep learning model is constructed using the PyTorch Geometric 2.6.1 framework, and the traditional machine learning algorithm is implemented based on Scikit-learn 1.6.1.
[0114] 1. Evaluation Metrics
[0115] Select the Pearson correlation coefficient R and the determination coefficient R, which are commonly used in many research works. 2 To evaluate the performance of the model.
[0116] The Pearson correlation coefficient R is used to measure the linear correlation between the predicted value and the true value, and its value range is [-1,1]. When R is close to 1, it means that the predicted value and the true value have a high positive correlation; when R is close to -1, it means that the two are highly negatively correlated; when R is close to 0, it indicates that there is almost no linear relationship between the two. The formula is defined as:
[0117]
[0118] Among them, y i and Represent the true value and predicted value of the i-th sample respectively, and are the means of the true value and the predicted value, respectively.
[0119] Coefficient of determination R 2 It is used to measure the model's ability to explain the changes in the target variable. Its value range is (-∞,1]. When R 2 When it is close to 1, it means that the model can capture the changing trend of the target variable well; when R 2 When it is close to 0 or negative, it indicates that the model fit is poor. The formula is defined as:
[0120] Among them, the molecule The sum of squared errors between the model prediction and the true value reflects the part that the model cannot explain; the denominator It represents the sum of squared deviations between the true value and its mean, reflecting the overall fluctuation of the target variable.
[0121] 2. Experimental results
[0122] Tables 2 and 3 show the prediction performance comparison of different methods, except for the two baseline models (MTL-Guo and Net 2f -Time f ), STL-Ours represents that the proposed method performs two prediction tasks independently, i.e., the shared strategy of multi-task learning is removed.
[0123] In the comparison between single-task learning (STL-Ours) and multi-task learning (MTL-Ours), it can be found that the performance of the two methods in online time delay prediction is almost the same. However, in terms of online time delay prediction, the multi-task learning model (MTL-Ours) shows a higher correlation coefficient R and determination coefficient R 2 This result shows that the multi-task learning framework can significantly improve the prediction ability of line delay without affecting the accuracy of line length prediction, which fully demonstrates the effectiveness and advantages of multi-task learning.
[0124] Table 2 Comparison of prediction performance of different methods on the test set (R)
[0125]
[0126]
[0127] Table 3 Comparison of prediction performance of different methods on the test set (R 2 )
[0128]
[0129] When the proposed method (MTL-Ours) is compared with the baseline models (MTL-Guo and Net 2f -Time f ), it can be found that the overall performance of the proposed method in all test designs is significantly better than the two baseline models, which shows that the method has stronger generalization ability and better adaptability. In the online long prediction task, the correlation coefficient R of MTL-Ours is respectively higher than that of the baseline model Net 2f and MTL-Guo were 12% and 1% higher, respectively, and the coefficient of determination R 2 In the online delay prediction task, the correlation coefficient R of MTL-Ours is 26.3% and 21.2% higher than that of the baseline model Time f and MTL-Guo were 4.7% and 4.8% higher, respectively, and the coefficient of determination R 2 10.4% and 15.1% higher. In addition, it should be noted that the baseline model Time fThe performance on the training set is significantly better than that on the test set, showing obvious overfitting, which limits its generalization ability. In contrast, the stability and superiority of the method on the test set further highlight its potential in practical applications.
[0130] In addition, Table 4 also records the time taken for training and reasoning of each method. The results show that by introducing a multi-task learning framework to integrate the line length and line delay prediction process, the time cost of training and reasoning can be reduced and the prediction efficiency can be improved.
[0131] Table 4 Training time and inference time of different methods
[0132]
[0133] The embodiments described above are only preferred embodiments of the present invention and are not intended to limit the scope of implementation of the present invention. Therefore, all changes made according to the shape and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for collaboratively predicting line length and line delay after digital integrated circuit layout, characterized in that: include: Construct a preliminary heterogeneous graph based on the connection information of the circuit netlist; Extract features for line length and line delay prediction of line edges from netlist files and process files; Add the extracted features to the corresponding nodes and edges of the initial heterogeneous graph to obtain the final heterogeneous graph with connection relationships and features; Construct a line length and line delay collaborative prediction model based on graph neural network and multi-task learning; Train the line length and line delay collaborative prediction model based on graph neural network and multi-task learning; The final version of the heterogeneous graph is input into the trained line length and line delay collaborative prediction model based on graph neural network and multi-task learning, and finally the line length prediction results and line delay prediction results are obtained.
2. The method for co-predicting line length and line delay after digital integrated circuit layout according to claim 1, characterized in that: The features extracted from the netlist file and the process file include node features and edge features; Among them, the edge features used for line length prediction of line edges include the driving strength of the source pin unit, the source pin fan-out number, the source pin unit area, the source pin unit pin number, the target pin unit area, and the target pin unit fan-out number; The node characteristics used for line delay prediction include transition time, margin, capacitance of input pins, and maximum capacitance of output pins; The edge features used for line delay prediction include line network capacitance and line network resistance.
3. The method for co-predicting line length and line delay after digital integrated circuit layout according to claim 1, characterized in that: The line length and line delay collaborative prediction model constructed based on graph neural network and multi-task learning includes three TransformerConv layers. There is a cross-stitch unit after each TransformerConv layer. The last cross-stitch unit connects two task branches. Both task branches include a wired edge embedding generation module and a fully connected layer.
4. The method for collaboratively predicting line length and line delay after digital integrated circuit layout according to claim 3, characterized in that: The working mechanism of the TransformerConv layer includes: For each node i in the graph, its input feature is expressed as Where d is the feature dimension and l is the index of the current network layer. First, through the learnable weight matrix Perform linear transformation on node features to obtain query vectors Key Vector Sum value vector The query vector Used to measure the degree of attention of node i to other nodes; key vector Used to match query vectors of other nodes to calculate attention weights; value vector Contains the actual information of the node, which is used for weighted summation to update the node features; In order to utilize the information of edge features in the graph, the TransformerConv layer uses a learnable weight matrix Opposite edge features Perform a linear transformation and map it to the same dimensional space as the node features: Then, the attention weight coefficient α of node i and its neighbor node j is calculated according to the above transformation results ij : in Dot product operation q T k measures the similarity between the query vector and the key vector, while the normalization in the denominator ensures that the sum of the attention weights is 1. The introduction of enhances the model's ability to utilize edge feature information; Based on the attention weight coefficient α ij , new features of nodes The weighted sum of the value vectors of neighbor nodes is obtained: in is a learnable weight matrix used to transform the features of the node itself; The node's own characteristic information is retained, while Fusion of neighbor node information, through attention weights Implementing weighted aggregation; edge features Adding it to the aggregation process enhances the model's ability to express the relationship between nodes.
5. The method for co-predicting line length and line delay after digital integrated circuit layout according to claim 4, characterized in that: Each TransformerConv layer is followed by a cross-stitch unit, which performs dynamic node feature interaction during the feature extraction process of the line length prediction task T1 and the line delay prediction task T2. In the first layer of TransformerConv, the node features of the two tasks T1 and T2 are represented as and Where i represents the node index and d is the feature dimension; The cross-stitch unit generates a new feature representation by weighted fusion of feature representations of different tasks by learning the linear combination weights between tasks. and Among them, α 11 , α 12 , α 21 , α 22 are all attention weight coefficients.
6. A method for collaboratively predicting line length and line delay after digital integrated circuit layout according to any one of claims 3 to 5, characterized in that: In the edge embedding generation module of each task branch, the embedding representations of the source node and the target node are fused into a new embedding representation, namely edge embedding, through a splicing operation. This splicing operation preserves the complete information of the source node and the target node and captures the relationship between the two. Each task branch processes the embedded representation through a fully connected layer, which consists of three consecutive hidden layers. Each layer uses a nonlinear activation function ReLU to enhance the expressiveness of the model and finally outputs the prediction result.
7. The method for co-predicting line length and line delay after digital integrated circuit layout according to claim 6, characterized in that: For the line length and line delay collaborative prediction model based on graph neural network and multi-task learning, the line length prediction task T1 and the line delay prediction task T2 use mean square error as the single task loss function, which are defined as follows: Where N is the total number of edges, y wirelength,i and Represent the true length and predicted value of the i-th line edge, y netdelay,i and They represent the actual line delay value and predicted value of the i-th line edge respectively; In order to effectively balance the optimization objectives of the two tasks in the multi-task learning framework, the geometric mean loss is used as the joint loss function. Its definition is as follows:
Citation Information
Patent Citations
Netlist-level line delay prediction method and device based on LightGBM, and medium
CN113609812A
Pre-wiring time delay prediction method and device, equipment, storage medium and program product
CN118350339A
Time sequence analysis method and related device
CN119067065A
Generating integrated circuit placements using neural networks
US20210334445A1
Generative self-supervised learning to transform circuit netlists
US20230334215A1
Cited By
Business process prediction method based on multi-scale feature fusion and graph time sequence modeling
CN121258671A