FPGA multi-performance index prediction method based on multi-task learning
Through the multi-task learning method, combined with class-graph image data and directed acyclic graph structure data, and using CNN-GCN combined model to extract and fusion features, the problem that single-task learning is difficult to predict multiple FPGA design performance indicators is solved, achieving higher prediction accuracy and more reliable design optimization support.
Patent Information
- Application Number
- CN202510172522.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-10
AI Technical Summary
Existing single-task learning algorithms can only predict one specific FPGA design performance metric and cannot capture the complex trade-offs and interdependencies between different design goals, making it difficult to predict multiple performance metrics at the same time.
The multi-performance index prediction method based on multi-task learning is adopted. By extracting the layout of the class diagram image data and directed acyclic graph structure data, the CNN-GCN combined model is used for feature extraction and fusion, and different decoders are used to predict different prediction targets, and congestion, line length, power consumption and critical path delay are predicted at the same time.
The ability to predict multiple FPGA design performance indicators simultaneously is realized. Through information sharing and multi-task collaborative training, the prediction accuracy is improved and more reliable decision support is provided for FPGA design optimization.
Smart Images

Figure CN120124583A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital integrated circuit electronic design automation (EDA), and particularly relates to an FPGA multi-performance index prediction method based on multi-task learning. Background Art
[0002] With the continuous expansion of the scale of FPGA chips and the improvement of the complexity of on-chip design, the process of FPGA circuit design has gradually become more standardized and complex. In this context, the role of electronic design automation (EDA) technology has become increasingly important. EDA combines software, hardware, and related services, aiming to provide automated solutions for the design and manufacture of very large scale integration (VLSI) circuits. Its core objectives usually involve system design, physical design, and customization services. The process of physical design can be subdivided into links such as technology mapping, placement, and routing.
[0003] Among these links, placement and routing are considered the most time-consuming steps. The placement stage is responsible for arranging all design components to legal positions on the FPGA chip, while the routing stage determines the interconnection method between different logic units, aiming to optimize the design objectives. The traditional placement and routing process usually needs to obtain the performance indexes of the design after completing the routing, and then readjust the placement or routing according to these quality results (such as wire length, congestion, power consumption, and timing, etc.) until the best trade-off between various objectives is achieved. However, the placement scheme that overly relies on routing resources often has an adverse impact on the subsequent routing process.
[0004] In recent years, the rapid development of machine learning has almost affected all disciplines. Machine learning surpasses traditional methods through a large amount of data and optimization-based algorithms and achieves results at almost the human level. The integration of supervised learning technology has brought significant changes to the placement and routing design process. By providing training data and labels composed of placement results and the quality of results after routing (QoR), the supervised learning model can learn the patterns and relationships between placement features and routing performance, use this learned information to predict the QoR after routing based on the new placement results without physical implementation, and then readjust the placement strategy according to the predicted performance indexes, which avoids unnecessary routing attempts and greatly improves the overall efficiency and effectiveness of the design process. Although the prediction model has shown good results in accelerating the design speed and improving the design quality, most existing methods adopt a single-task learning algorithm (STL) model, which can only predict a specific design objective and cannot capture the complex trade-offs and interdependencies between different design objectives. Therefore, a machine learning model with the ability to jointly predict multiple design indexes can be used as an alternative and improvement to the single-task learning model.
[0005] In summary, existing single-task prediction works usually focus only on a single specific performance metric, and few multi-task prediction works can well cover multiple metrics sufficient to represent the performance of FPGA designs.
[0006] In view of this, there is an urgent need to develop a method that can fully utilize the rich information obtained in the placement stage, enabling the model to predict multiple performance metrics simultaneously. Summary of the Invention
[0007] The purpose of the present invention is to provide an FPGA multi-performance metric prediction method based on multi-task learning to solve at least one technical problem in the prior art.
[0008] The technical solution of the present invention is as follows:
[0009] An FPGA multi-performance metric prediction method based on multi-task learning, comprising:
[0010] Based on the extracted class diagram image data after placement, obtaining directed acyclic graph structure data of the connection relationship between pins;
[0011] Using a CNN-GCN combined model with an encoder-decoder structure to extract features from the class image data and the directed acyclic graph data, obtaining class image data features and directed acyclic graph data features;
[0012] Fusing the class image data features and the directed acyclic graph data features;
[0013] Using different decoders for prediction according to different prediction targets to simultaneously predict multiple performance metrics;
[0014] The performance metrics include: congestion, wire length, power consumption, and critical path delay.
[0015] The obtaining of the directed acyclic graph structure data of the connection relationship between pins includes:
[0016] Taking pins as nodes and the connection relationship between pins as edges to form a directed acyclic graph structure;
[0017] The pins include one or several of the input pins IPIN of combinational logic components, output pins OPIN, input SINK of sequential logic components, output SOURCE, and clock input pins CPIN;
[0018] The edges include: clock_capture when CPIN points to SINK; or clock_launch when CPIN points to SOURCE; or other combinations.
[0019] Any of the nodes and edges contains delay information obtained from a relatively pessimistic analysis.
[0020] The CNN-GCN combined model of the decoder-decoder structure includes: a CNN encoder, a GCN encoder, and a decoder;
[0021] Among them, the CNN encoder is composed of three convolutional layers and two max pooling layers, and each convolutional layer contains two regular convolutional kernels, and applies the LeakyReLU activation function and the InstanceNorm2d normalization layer; among them, the max pooling layer downsamples the feature map to improve the abstraction degree of the features; the last convolutional layer combines batch normalization and the hyperbolic tangent activation function;
[0022] The GCN encoder consists of three convolutional layers and an average pooling layer, and each convolutional layer consists of a graph convolutional layer, a residual residual connection, a BatchNorm1d batch normalization layer, and a Relu activation function;
[0023] The decoder is composed of 4 parallel network structures, and any one of the network structures is used to process one task. For different prediction tasks, different structures are used for decoding.
[0024] The features of the nodes in the graph are captured through the three convolutional layers in the GCN encoder, the end nodes and their features are extracted, and all nodes are averaged pooled and the end nodes are averaged pooled to obtain global features and end features;
[0025] Among them, the end features are used for the critical path delay prediction task, and the global features are upsampled for feature fusion and participate in the prediction tasks of congestion, wire length, and power consumption.
[0026] The decoder is composed of 4 parallel network structures, and any one of the network structures is used to process one task. For different prediction tasks, different structures are used for decoding, including:
[0027] For the congestion prediction task, the decoder adopts the FCN decoder architecture, including three transposed convolutional layers and two convolutional layers; moreover, in the second convolutional layer, the skip connection architecture is used to connect the same-shaped features of the encoder with the decoder, so as to fuse features of different scales and output a congestion heat map.
[0028] The decoder is composed of 4 parallel network structures, and any one of the network structures is used to process one task. For different prediction tasks, different structures are used for decoding, including:
[0029] For the wire length and power consumption prediction tasks, the decoder adopts a CNN decoder, including an AdaptiveAvgPool2d layer and four fully-connected layers, and the output is a continuous scalar; or,
[0030] For the critical path delay prediction task, the decoder adopts two fully-connected layers, and its output is a continuous scalar.
[0031] The described FPGA multi-performance index prediction method based on multi-task learning further includes a parameter optimization process for the CNN-GCN combined model, including:
[0032] Optimizing the trainable parameters in the CNN-GCN combined model through gradient descent.
[0033] The parameter optimization process of the CNN-GCN combined model includes:
[0034] Represent the total parameters of the WCPTNet to be trained as θ total , and the equation of the total loss function is as follows:
[0035]
[0036] Among them, L w (θ w ), L p (θ p ), L c (θ c ), L t (θ t ) are the loss functions of wire length, power consumption, congestion, and critical path delay respectively, and at the same time θ w , θ p , θ c , θ t are parameters that can be trained;
[0037] For wire length prediction, the real data is represented as a single continuous value, denoted as The loss function is defined as the predicted wire length value, denoted as And the mean square error between it and the real label wire length value is:
[0038]
[0039] The loss function of power consumption prediction is defined as the predicted power consumption value And the mean square error between the real label power consumption value , and the loss function of the critical path delay is defined as the predicted critical path delay value And the mean square error between the real label value :
[0040]
[0041] For congestion prediction, the ground truth is a heatmap of continuous values, denoted as which contains the congestion value of each pixel; the loss function is defined as the mean squared error between the predicted value and the ground truth value of the k-th pixel in the j-th graph :
[0042]
[0043] where n pixel represents the number of pixels in the congestion heatmap; n 1 = n 2 = n 3 = n 4 , and the relationship between different trainable parameters is described as follows:
[0044] θ total = θ w ∪ θ p ∪ θ c ∪ θ t ;
[0045] To minimize L total , the trainable parameters are optimized by gradient descent.
[0046] The beneficial effects of the present invention at least include:
[0047] The method of the present invention, firstly, on the basis of extracting the class diagram image data after layout, obtains the directed acyclic graph structure data of the connection relationship between pins; then uses the CNN-GCN combined model with an encoder-decoder structure to extract features from the class image data and the directed acyclic graph data, obtaining class image data features and directed acyclic graph data features; and fuses the class image data features and the directed acyclic graph data features; finally, uses different decoders for prediction according to different prediction targets to simultaneously predict 4 performance indicators; the method of the present invention, by making full use of the rich information obtained in the layout stage, combines supervised learning models for different data types to extract features, and realizes information sharing through multi-task collaborative training, so that the model can simultaneously predict multiple performance indicators, and in this way, can more comprehensively capture the complex relationships and mutual influences in the design, improve the prediction accuracy, and provide more reliable decision support for FPGA design optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is the construction flow chart of WCPTNet;
[0049] Figure 2(a) shows the utilized logic resource blocks in the image_place as a dark pixel map;
[0050] Figure 2(b) shows the image_connect as a black fly-line map;
[0051] Figure 2(c) shows the color scheme map of the image_pin_util;
[0052] Figure 2(d) shows a directed acyclic timing diagram;
[0053] Figure 3 is the network architecture diagram of WCPTNet;
[0054] Figure 4 is the training loss diagram of the WCPTNet network. Detailed implementation manner
[0055] The present application will be further described below with reference to the accompanying drawings.
[0056] In the prior art, in order to fully utilize the advantages of this innovative design, it is crucial to establish an accurate and efficient prediction model, and some related studies have made remarkable progress. In these studies, the congestion task is usually modeled as a computer vision task, and the correlation between layout features is learned through the model. Since layout features can be naturally represented as data similar to images - for example, standard cells, routing channels, and other structural elements are usually visually arranged on a grid, computer vision techniques can be used to process this information. For example, a conditional generative adversarial network (pix2pix) is used to predict the heat map of congestion. Or, a CNN-based framework is used to predict the routability of the design. In the field related to wire length, linear regression and gradient boosting regression methods are adopted to study the relationship between the FPGA layout scheme and the internal wiring wire length. In the chip design process, timing optimization is crucial as it directly affects the performance, stability, and reliability of the final product. To this end, there is a framework based on LGBM (Light Gradient Boosting Machine) that accurately and effectively estimates the delay of the gate-level circuit by predicting the depth of the corresponding LUT logic after prediction technology mapping.
[0057] Although prediction models have shown good results in accelerating design speed and improving design quality, most existing methods adopt single-task learning algorithm (STL) models, which can only predict a specific design goal and cannot capture the complex trade-offs and interdependencies between different design goals. Therefore, a machine learning model with the ability to jointly predict multiple design metrics can be used as an alternative and improvement to the single-task learning model. However, there is still relatively little work on multi-task in the EDA field. In the FPGA field, there are multi-task prediction models that can simultaneously predict wire length, congestion, and power consumption using the features in the FPGA placement stage. In the ASIC field, there are also MTL models used to simultaneously predict IRdrop and routing congestion for each design based on the placement scheme, and then the predicted results provide guidance for the placement scheme. Although there has gradually been work on applying multi-task learning to the EDA field, these works still cannot cover more targets well. Therefore, there is still great potential for exploration of multi-task learning in EDA.
[0058] It can be seen that in past research, most researchers have been dedicated to prediction work for a single target, and prediction work for multiple targets has gradually emerged. However, these single-task prediction works usually only focus on a single specific performance metric, and a few multi-task prediction works have not been able to cover multiple metrics that are sufficient to represent FPGA design performance well. Specific Example 1:
[0060] The present invention provides an example:
[0061] In this embodiment, the FPGA multi-performance metric prediction method based on multi-task learning is named WCPTNet, and its overall structure is as Figure 1 shown. Based on the extraction of the class diagram image data after placement, the directed acyclic graph structure data with nodes as component pins and edges as the connection relationships between pins is extracted at the same time. The CNN-GCN combined model with an encoder-decoder structure is used to extract features from the class image data and the directed acyclic graph data, and then the features of the two types are fused so that different types of information can be jointly used to predict the performance metrics of FPGA, namely congestion, wire length, power consumption, and critical path delay. Different decoders are used to obtain results for different targets, and multiple tasks are co-trained to share information with each other so that it can simultaneously predict multiple performance metrics.
[0062] The specific steps are as follows:
[0063] S1: Feature extraction: Define the input features for training and inference, mainly including two parts: the layout feature image data and the directed acyclic graph data for timing analysis. The former is a feature proposed with reference to existing literature, and the latter is a novel feature proposed in this paper.
[0064] S101: Image data:
[0065] Define three feature images named image_place, image_connect, and image_pin_util respectively, which are used to describe the layout of logic resource blocks, the connection relationship between logic resource blocks, and the pin utilization rate of logic resource blocks. As shown in Figures 2(a), (b), and (c), the utilized logic resource blocks in image_place are represented by dark pixels, image_connect is represented by black fly lines, and the color scheme of image_pin_util is Plasma. Each logic block will be mapped to a corresponding color according to the pin utilization rate. Capturing the number of utilized logic resource blocks, layout positions, their topological connection information, and the actual wiring requirements of the surrounding area of the block provided by the pin utilization rate from the three feature maps makes an important contribution to the accuracy of MTL model prediction.
[0066] S102: Directed acyclic timing diagram data:
[0067] Whether in the FPGA field or the ASIC field, the current mainstream timing prediction methods mainly use graph neural networks (GNNs). Therefore, it is very crucial to construct a topological graph that accurately represents the timing path. The directed acyclic timing diagram proposed in this embodiment is composed of pins as nodes and the connection relationships between pins as edges. As shown in Figure 2(d), the types of pins include input pins IPIN of combinational logic components, output pins OPIN, input SINK of sequential logic components, output SOURCE, and clock input pins CPIN. The types of edges include clock_capture when CPIN points to SINK, clock_launch when CPIN points to SOURCE, and other combinations. Each node and edge has delay information obtained from a somewhat pessimistic analysis. This timing diagram not only contains rich timing information but also contains global topological structure features with a finer granularity than image data, which is beneficial to supplementing information not covered by image data and improving the prediction accuracy of congestion, wire length, and power consumption while predicting delay using the timing diagram.
[0068] S2: Generation of true labels:
[0069] As Figure 1 shown, with the help of EDA tools, the true values of wire length, congestion, and critical path delay can be obtained after routing, while the true value of power consumption can be obtained in the post-routing analysis stage. For each layout solution, the true values of wire length, power consumption, and critical path delay are scalars, while the true value of congestion is a hot spot map.
[0070] S3: Construction of the WCPTNet model:
[0071] Using a fully convolutional network (FCN) and a convolutional neural network (CNN) to simultaneously predict congestion, wire length, and power consumption not only has better accuracy than single-task models but also better prediction efficiency. At the same time, a graph neural network (GCN) is used to add a task branch for predicting the critical path delay, and by embedding global timing information into the global layout features, congestion, wire length, and power consumption are predicted. This model mainly consists of three parts: a CNN encoder, a GCN encoder, a decoder, and a parameter optimization strategy.
[0072] S301: CNN Encoder:
[0073] In WCPTNet, the encoder for processing image data is a CNN encoder, which consists of three convolutional layers and two max-pooling layers. Each convolutional layer uses two regular convolutional kernels and applies the LeakyReLU activation function and the InstanceNorm2d normalization layer. The role of the convolutional layer is to extract local features from the image, while the max-pooling layer helps improve the abstraction degree of the features by downsampling the feature map. The last convolutional layer combines batch normalization and the hyperbolic tangent activation function, which not only helps accelerate the network convergence process but also improves the gradient flow and reduces the sensitivity to the weight and bias initialization.
[0074] S302: GCN Encoder:
[0075] The encoder for processing image data in WCPTNet is the GCN encoder, which consists of three convolutional layers and one average-pooling layer. Each convolutional layer consists of a graph convolutional layer (GCNConv), a residual connection, a BatchNorm1d batch normalization layer, and a Relu activation function. Through multi-layer graph convolution, the features of the nodes in the graph are captured and the node out-degree is calculated, and the end nodes and their features are extracted. The global features and end features are obtained by averaging pooling all nodes and averaging pooling the end nodes respectively. The end features are used for the critical path delay prediction task, and the global features are upsampled to participate in the other three prediction tasks by fusing with the features of the CNN encoder. This encoder combines the aggregation of edge features, residual connections, and batch normalization techniques to improve the network stability and training efficiency.
[0076] S303: Decoder:
[0077] In WCPNet, the decoder consists of four parallel network structures, each of which is used to process a task. For the congestion prediction task, the decoder part adopts a typical FCN decoder architecture, including three transposed convolutional layers and two convolutional layers. In the second convolutional layer, a skip connection architecture is used to connect a certain feature layer of the encoder to the decoder, so as to fuse features of different scales, and the output of this part is the congestion heatmap. For the wire length and power consumption prediction tasks, the decoder part adopts a typical CNN decoder, including an AdaptiveAvgPool2d layer and four fully connected layers (LinearLayer), and these two outputs are continuous scalars. For the critical path delay prediction task, the decoder adopts two fully connected layers (Linear Layer), and its output is also a continuous scalar.
[0078] S304: Parameter optimization strategy:
[0079] WPCTNet trains a single model to perform multiple tasks simultaneously, and the loss function plays a key role in the training process. To capture the objectives and requirements of each individual task while considering joint optimization, a geometric loss strategy (GLS) is adopted. The total parameters of the WCPTNet to be trained are represented as θ total , and the equation of the total loss function is detailed as follows:
[0080]
[0081] where L w (θ w ), L p (θ p ), L c (θ c ), L t (θ t ) are the loss functions of wire length, power consumption, congestion, and critical path delay respectively, and at the same time θ w , θ p , θ c , θ t are the parameters that can be trained.
[0082] For wire length prediction, in formula (2), the real data is represented as a single continuous value, denoted as The loss function is defined as the mean square error (MSE) between the predicted wire length value (denoted as ) and the real label wire length value:
[0083]
[0084] Similarly, the loss function of power consumption prediction is defined as the predicted power consumption value and the real label power consumption value The mean squared error (MSE) between the predicted critical path delay value and the true label value is defined as the mean squared error (MSE) between the predicted critical path delay value and the true label value:
[0085]
[0086]
[0087] For congestion prediction, in Equation 2, the ground truth is a heatmap of continuous values, denoted as which contains the congestion value of each pixel. The loss function is defined as the mean squared error (MSE) between the predicted value and the ground truth value of the k-th pixel in the j-th map: where n
[0088]
[0089] represents the number of pixels in the congestion heatmap. Similar to Equation 2, note that n pixel = n 1 = n 2 = n 3 = n 4 . The relationship between different trainable parameters is described as follows:
[0090] θ total = θ w ∪ θ p ∪ θ c ∪ θ t (6)
[0091] To minimize L total , the trainable parameters can be optimized using techniques such as gradient descent or its variants.
[0092] Verification step:
[0093] To verify WPCTNet, 29 different designs were selected from the benchmark suite provided by VTR, and the resource utilization ranges of these designs are shown in Table 1. With the help of the VTR tool, multiple layout results were generated for each design using various different layout parameters (including the default settings), and then each layout result was routed using the default routing parameters. Finally, the generated dataset included 1600 samples, excluding those samples that were not successfully routed.
[0094] Table 1 Benchmark resource utilization ranges.
[0095] #LUTs #Nets #Adders #Multiply 1k - 25k 3k - 36k 0.25k - 2.5k 18-136
[0096] Meanwhile, to evaluate the generalization ability of the method described in this embodiment, a cross-design scheme was adopted for data splitting. Specifically, the training data included samples generated from 20 designs randomly selected from the above 29 designs, while the samples of the remaining designs were reserved for testing. This ensured that no design in the test data was seen during the training process. This embodiment used a computer equipped with an Intel Core i7-11800H CPU, 32GB of memory, and an Nvidia GTX4090 graphics card. The feature extractor and the ground-truth generator were implemented in C++ based on VTR. The WCPTNet and the baseline WCPNet models were implemented in Python based on PyTorch. The learning rate was set to 0.0001, the batch size was set to 4, the number of training epochs was set to 1000, and the AdamW optimizer was adopted.
[0097] To evaluate the prediction accuracy of WCPTNet, for congestion, wire length, and power consumption, the prediction results of the best MTL prediction model WCPNet in the FPGA field (Xian, Juming, et al. "WCPNet: Jointly Predicting Wirelength, Congestion and Power for FPGA Using Multi-Task Learning." ACM Transactions on Design Automation of Electronic Systems 29.5 (2024): 1-19.) were used as the comparison baseline. For the critical path delay, the critical path delay obtained by VPR through static timing analysis (STA) after the placement completion stage in VTR was extracted and normalized, and the MAPE was calculated with the final routed critical path delay value (already normalized) for use as the comparison benchmark for the critical path delay prediction task. WCPTNet and the baseline model WCPNet were trained on the training data and then their performance was evaluated on the test data that only included designs not seen in the training set. In Figure 4 it, the training loss curve showed that the loss curve converged well at around the 400th epoch. The prediction accuracies of WCPTNet and the baseline model WCPNet are listed in Table 2.
[0098] Table 2 Prediction Accuracy of WCPNet
[0099]
[0100] As can be seen from Table 1, the prediction accuracy of WCPTNet is 4.43%, 20.6%, and 8.57% higher than that of WCPNet in terms of wire length, power consumption, and congestion, respectively. Compared with STA in VTR, the prediction accuracy of the critical path delay is 56.45% higher. The WCPTNet in this embodiment shows superior performance in prediction accuracy compared to WCPNet.
[0101] Figure 4 The test losses of WCPNet (including the total loss and the losses of each task) and WCPNet were compared. When the training process converged, the three task-specific losses of congestion, wire length, and power consumption of WCPTNet were lower than those of the WCPNet model, and the delay task loss was lower than the single-task loss.
[0102] The method described in this embodiment is based on the class diagram image data after layout, combined with the directed acyclic graph (DAG) structure data that regards nodes as component pins and edges as the connection relationships between pins. Feature extraction is performed on these two data types through a supervised learning model, so as to fully explore their potential information. Finally, the feature information extracted from the class image data and the directed acyclic graph data is combined and jointly used to predict multiple performance indicators of the FPGA. This method of multi-source feature fusion is expected to improve the prediction accuracy and more comprehensively capture the complex dependencies in the layout design, thereby optimizing the design effect; at the same time, the method described in the present invention, by making full use of the rich information obtained in the layout stage, combines the supervised learning model for different data types for feature extraction, and realizes information sharing through multi-task collaborative training, so that the model can predict multiple performance indicators at the same time. In this way, the complex relationships and interactions in the design can be more comprehensively captured, the prediction accuracy can be improved, and more reliable decision-making support can be provided for FPGA design optimization.
[0103] The above discloses only several specific implementation scenarios of the present invention. However, the present invention is not limited thereto, and any changes that can be thought of by those skilled in the art should fall within the protection scope of the present invention. The above serial numbers of the present invention are only for description and do not represent the superiority or inferiority of the implementation scenarios.
Claims
1. A FPGA multi-performance index prediction method based on multi-task learning, characterized in that: include: Based on the extracted layout class diagram image data, the directed acyclic graph structure data of the connection relationship between the pins is obtained; Using a CNN-GCN combined model of a decoder-decoder structure to extract features from the image-like data and the directed acyclic graph data, to obtain image-like data features and directed acyclic graph data features; Performing feature fusion on the image-like data features and the directed acyclic graph data features; Different decoders are used to predict different prediction targets to simultaneously predict multiple performance indicators; The performance indicators include: congestion, line length, power consumption and critical path delay.
2. The FPGA multi-performance index prediction method based on multi-task learning according to claim 1 is characterized in that: The step of obtaining the directed acyclic graph structure data of the connection relationship between the pins includes: The pins are nodes, and the connections between the pins are edges to form a directed acyclic graph structure; The pins include: one or more of an input pin IPIN and an output pin OPIN of a combinational logic component, an input SINK and an output SOURCE of a sequential logic component, and a clock input pin CPIN; The edge includes: clock_capture when CPIN points to SINK; or clock_launch when CPIN points to SOURCE; or other combinations. Any of the nodes and edges contain delay information obtained through pessimistic analysis.
3. The FPGA multi-performance index prediction method based on multi-task learning according to claim 1 is characterized in that: The CNN-GCN combination model of the decoder-decoder structure includes: a CNN encoder, a GCN encoder, and a decoder; The CNN encoder consists of three convolutional layers and two maximum pooling layers, and each convolutional layer contains two regular convolution kernels, and applies the LeakyReLU activation function and the InstanceNorm2d normalization layer; the maximum pooling layer downsamples the feature map to improve the abstraction level of the feature; the last convolutional layer combines batch normalization and hyperbolic tangent activation function; The GCN encoder consists of three convolutional layers and an average pooling layer, and each convolutional layer consists of a graph convolution layer, a residual connection, a BatchNorm1d batch normalization layer and a Relu activation function; The decoder is composed of four parallel network structures, each of which is used to process one task. Different structures are used for decoding for different prediction tasks.
4. The FPGA multi-performance index prediction method based on multi-task learning according to claim 3 is characterized in that: The three convolutional layers in the GCN encoder are used to capture the features of the nodes in the graph, extract the terminal nodes and their features, and average pool all nodes and average pool the terminal nodes to obtain global features and terminal features; Among them, the endpoint features are used for the critical path delay prediction task, and the global features are used for feature fusion through upsampling to participate in the prediction tasks of congestion, line length, and power consumption.
5. The FPGA multi-performance index prediction method based on multi-task learning according to claim 3 is characterized in that: The decoder is composed of four parallel network structures, each of which is used to process one task. Different structures are used for decoding for different prediction tasks, including: For the congestion prediction task, the decoder adopts the FCN decoder architecture, which includes three transposed convolutional layers and two convolutional layers; and, in the second convolutional layer, a skip connection architecture is used to connect the encoder's same-shape features with the decoder, thereby fusing features of different scales, and the output is a congestion heat map.
6. The FPGA multi-performance index prediction method based on multi-task learning according to claim 3 is characterized in that: The decoder is composed of four parallel network structures, each of which is used to process one task. Different structures are used for decoding for different prediction tasks, including: For the line length and power consumption prediction tasks, the decoder uses a CNN decoder, including an AdaptiveAvgPool2d layer and four fully connected layers, and the output is a continuous scalar; or, For the critical path delay prediction task, the decoder uses two fully connected layers whose output is a continuous scalar.
7. The FPGA multi-performance index prediction method based on multi-task learning according to claim 3 is characterized in that: It also includes a parameter optimization process for the CNN-GCN combination model, including: Optimize the trainable parameters in the combined CNN-GCN model via gradient descent.
8. The FPGA multi-performance index prediction method based on multi-task learning according to claim 3 or 7, characterized in that: The parameter optimization process of the CNN-GCN combination model includes: The total parameters of WCPTNet to be trained are represented as θ total , the equation of the total loss function is as follows: Among them, L w (θ w )、L p (θ p )、L c (θ c )、L t (θ t ) are the loss functions of line length, power consumption, congestion and critical path delay respectively, and θ w ,θ p ,θ c ,θ t are parameters that can be trained; For line length prediction, the real data is represented as a single continuous value, denoted as The loss function is defined as the predicted line length value, expressed as And the mean square error between the actual label line length value is: The loss function for power consumption prediction is defined as the predicted power consumption value Compared with the actual tag power consumption value The mean square error between the two, the loss function of the critical path delay is defined as the predicted critical path delay value and the true label value The mean square error between: For congestion prediction, the true value is a heat map of continuous values, represented as It contains the congestion value of each pixel; the loss function is defined as the predicted value The ground truth value of the k-th pixel in the j-th image The mean square error between: where n pixel The number of pixels representing the congestion heat map; n1=n2=n3=n4. The relationship between different trainable parameters is described as follows: i total =θ w ∪θ p ∪θ c ∪θ t ; In order to minimize L total , the trainable parameters are optimized by gradient descent.
Citation Information
Cited By
Assessment device of FPGA interconnection architecture based on graph neural network
CN121212040A