VLSI congestion prediction method and device based on dual multi-modal fusion, and storage medium

By adopting the dual multimodal fusion method in VLSI design, grid and heterogeneous pattern features are constructed and deep feature fusion is carried out, the limitations of multimodal information processing and feature fusion are solved, and high accuracy and high efficiency VLSI congestion prediction is achieved.

CN119940275APending Publication Date: 2025-05-06SOUTHEAST UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510136297.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art performs poorly in processing multimodal information, and lacks effective fusion of layout and netlist features, resulting in limited accuracy and efficiency of VLSI congestion prediction.

Method used

Using a dual multimodal fusion method, we use the method to build grid-based two-dimensional layout features and heterogeneous graph-based netlist features, and input them into the congestion prediction model to perform feature extraction and fusion, and use deep feature fusion and cascade decoding modules to improve prediction accuracy.

Benefits of technology

It significantly improves the accuracy and efficiency of VLSI congestion prediction, and can quickly obtain high-quality and high-precision congestion prediction results in complex large-scale integrated circuit designs, guiding layout to achieve linear optimization and effectively reduce wiring congestion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940275A_ABST
    Figure CN119940275A_ABST
Patent Text Reader

Abstract

The invention discloses a VLSI congestion prediction method and device based on dual multi-modal fusion, and a storage medium. The method comprises the steps of 1, obtaining layout information and netlist information of a to-be-predicted circuit; 2, constructing two-dimensional layout features based on grids and netlist features based on heterogeneous graphs according to the layout information and the netlist information; and step 3, inputting the two-dimensional layout features and the netlist features into the congestion prediction model to obtain a congestion map of the circuit to be predicted. According to the VLSI congestion prediction method and device based on dual multi-modal fusion and the storage medium provided by the invention, a high-quality and high-precision congestion prediction result can be quickly obtained in a complex large-scale integrated circuit design, layout is guided to realize wiring optimization, and wiring congestion is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a VLSI congestion prediction method, device and storage medium based on dual multi-modal fusion, belonging to the technical field of integrated circuit electronic design automation (EDA). Background Art

[0002] With the advancement of manufacturing technology, modern very large scale integration (VLSI) circuit designs may contain millions of standard cells and thousands of macro modules to meet various design requirements. In the VLSI design flow, placement and routing have become critical but time-consuming stages. In addition, since placement has a significant impact on downstream stages such as clock tree synthesis and routing, placement significantly affects the quality and efficiency of chip design. Placers that do not consider routing congestion may cause routing failures and extend time to market. Therefore, it is critical to accurately and quickly predict routing congestion and optimize routing capabilities during the layout stage.

[0003] Traditional methods can be mainly divided into three categories: (1) static-based prediction methods, which estimate congestion directly based on layout attributes (such as pin density and network overlap) without actual routing, such as RUDY; (2) probability-based prediction methods, which use the probability of routing topology to estimate congestion, such as L-shaped or Z-shaped routing; (3) tool-based prediction methods, which rely on global routing tools to generate accurate congestion graphs, but the computational overhead is large and time-consuming. Although the first two methods are relatively fast, their accuracy is poor, while the third method is more accurate but has a high computational cost. Therefore, there is an urgent need for a congestion prediction method that is both accurate and efficient.

[0004] In recent years, machine learning (ML)-based methods have been widely used in VLSI congestion prediction due to their powerful prediction capabilities and fast decision-making characteristics. Various machine learning methods transform the congestion prediction problem into a computer vision problem, using neural networks to learn the relationship between layout and congestion graphs. For example, PROS2.0 is a machine learning-based plug-in tool that extracts grid-based layout features through convolutional layers and trains models using real global routing results to predict congestion graphs. However, such image-based models usually fail to effectively integrate netlist information, resulting in limited prediction accuracy.

[0005] At the same time, graph neural networks (GNNs) have also been used for congestion prediction. The proposed graph-based model treats the netlist as a graph and uses unit node features to predict congestion. However, traditional homogeneous GNN methods have performance bottlenecks when processing different netlists.

[0006] Therefore, in order to address the limitations of previous machine learning-based congestion prediction models in processing multimodal information and the lack of effective fusion of layout and netlist features, a new VLSI congestion prediction scheme based on multimodal fusion needs to be developed. Summary of the invention

[0007] Purpose: In order to overcome the deficiencies in the prior art, the present invention provides a VLSI congestion prediction method, device and storage medium based on dual multimodal fusion to solve the problems faced by the prior art in processing multimodal information poorly and lack of effective fusion of layout and netlist features.

[0008] Technical solution: To solve the above technical problems, the technical solution adopted by the present invention is:

[0009] In the first aspect, a VLSI congestion prediction method based on dual multi-modal fusion specifically includes:

[0010] Step 1: Obtain layout information and netlist information of the circuit to be predicted.

[0011] Step 2: constructing grid-based two-dimensional layout features and heterogeneous graph-based netlist features according to the layout information and netlist information.

[0012] Step 3: Input the two-dimensional layout features and netlist features into the congestion prediction model to obtain a congestion graph of the circuit to be predicted.

[0013] As a preferred solution, a method for constructing a grid-based two-dimensional layout feature specifically includes:

[0014] The circuit canvas to be predicted is divided into W×H grids, where each grid is defined as g i,j , i∈{1,2,…,W}, h∈{1,2,…,H}.

[0015] Four layout features are obtained from each grid and constructed into a grid-based two-dimensional layout feature.

[0016] Among them, four layout features include: macro module mask f Macro (i,j), uniform rectangular density mask f Rudy (i,j), pin density map f Pin (i,j) and the cell density map f Cell (i,j).

[0017] As a preferred solution, a method for constructing a netlist feature based on a heterogeneous graph specifically includes:

[0018] The circuit to be predicted is constructed as a heterogeneous graph, where the heterogeneous graph is defined as G Hetero = {V g ,Vn ,E gg ,E ng},V g and V n Denotes the node set of the grid and network respectively. gg represents the connection between grids, E ng Represents the connection between the grid and the wire net.

[0019] The features of each grid node and each wire network node are obtained to construct the netlist features based on the heterogeneous graph.

[0020] The grid node characteristics are defined as

[0021] Among them, x and y represent the horizontal and vertical coordinates of the relative position of the grid node in the placement area, respectively. Indicates the number of elements contained in the mesh node.

[0022] The network node feature is defined as

[0023] Among them, W bbox and h bbox Indicates B e The width and height of A bbox Indicates B e The area of is the number of cells belonging to the e-node of the network, B e is the minimum bounding box that contains all pins in net e.

[0024] As a preferred solution, the macroblock mask f Macro The expression of (i,j) is as follows:

[0025]

[0026] The uniform rectangular density mask f Rudy The expression of (i,j) is as follows:

[0027]

[0028] Among them, netlist represents netlist information.

[0029]

[0030] in, and Respectively represent the left, right, bottom, and top edges of the smallest bounding box that contains all the pins in net e.

[0031] The pin density diagram f PinThe expression of (i,j) is as follows:

[0032]

[0033] Among them, p e Indicates the pin of net e.

[0034] The cell density map f Cell The expression of (i,j) is as follows:

[0035] f Cell (i,j)=#Cells

[0036] Among them, #Cells represents the number of cells in the grid.

[0037] As a preferred solution, the congestion prediction model specifically includes: a feature extraction and fusion module, a deep feature fusion module and a cascade decoding module.

[0038] The feature extraction and fusion module includes: a first convolution layer, a second convolution layer, a third convolution layer, a fourth convolution layer, a first EFF module, a second EFF module, a third EFF module, a fourth EFF module, a first heterogeneous graph convolution layer, a second heterogeneous graph convolution layer, a third heterogeneous graph convolution layer, a fourth heterogeneous graph convolution layer and a fifth convolution layer.

[0039] Among them, the output end of the first convolutional layer is connected to the first input end of the first EFF module, the output end of the first heterogeneous graph convolutional layer is connected to the second input end of the first EFF module, the output end of the first heterogeneous graph convolutional layer is also connected to the input end of the second heterogeneous graph convolutional layer, and the output end of the first EFF module is connected to the input end of the second convolutional layer.

[0040] The output end of the second convolutional layer is connected to the first input end of the second EFF module, the output end of the second heterogeneous graph convolutional layer is connected to the second input end of the second EFF module, the output end of the second heterogeneous graph convolutional layer is also connected to the input end of the third heterogeneous graph convolutional layer, and the output end of the second EFF module is connected to the input end of the third convolutional layer.

[0041] The output end of the third convolutional layer is connected to the first input end of the third EFF module, the output end of the third heterogeneous graph convolutional layer is connected to the second input end of the third EFF module, the output end of the third heterogeneous graph convolutional layer is also connected to the input end of the fourth heterogeneous graph convolutional layer, and the output end of the third EFF module is connected to the input end of the fourth convolutional layer.

[0042] The output end of the fourth convolutional layer is connected to the first input end of the fourth EFF module, the first output end of the fourth heterogeneous graph convolutional layer is connected to the second input end of the fourth EFF module, and the second output end of the fourth heterogeneous graph convolutional layer is connected to the input end of the fifth convolutional layer.

[0043] The deep feature fusion module includes: an embedding layer and several multimodal Transformer layers.

[0044] The embedding layer is connected to several multimodal Transformer layers in sequence, the input end of the fourth EFF module is connected to the first input end of the embedding layer, and the output end of the fifth convolutional layer is connected to the second input end of the embedding layer.

[0045] The cascade decoding module includes: a sixth convolution layer, a seventh convolution layer, an eighth convolution layer, a ninth convolution layer, a first upsampling layer, a second upsampling layer, a third upsampling layer, and a fourth upsampling layer.

[0046] Among them, the first output end of the last multimodal Transformer layer is connected to the input end of the sixth convolutional layer, the output end of the sixth convolutional layer is connected to the input end of the first upsampling layer, the output end of the third EFF module and the output end of the first upsampling layer are respectively connected to the two input ends of the first fusion module, the output end of the first fusion module is connected to the input end of the seventh convolutional layer, the output end of the seventh convolutional layer is connected to the input end of the second upsampling layer, the output end of the second EFF module and the output end of the second upsampling layer are respectively connected to the two input ends of the second fusion module, the output end of the second fusion module is connected to the input end of the eighth convolutional layer, the output end of the eighth convolutional layer is connected to the input end of the third upsampling layer, the output end of the first EFF module and the output end of the third upsampling layer are respectively connected to the two input ends of the third fusion module, the output end of the third fusion module is connected to the input end of the ninth convolutional layer, and the output end of the ninth convolutional layer is connected to the input end of the fourth upsampling layer.

[0047] As a preferred solution, the first EFF module, the second EFF module, the third EFF module and the fourth EFF module all include: a tenth convolutional layer, a first global average pooling layer, a second global average pooling layer, a first fully connected layer, a second fully connected layer, a first activation function and a second activation function.

[0048] Among them, the first global average pooling layer, the first fully connected layer and the first activation function are connected in series, the input end of the first global average pooling layer is connected to the first input end of the first multiplication module, and the output end of the first activation function is connected to the second input end of the first multiplication module. The tenth convolutional layer, the second global average pooling layer, the second fully connected layer and the second activation function are connected in series, the output end of the second activation function is connected to the first input end of the second multiplication module, the output end of the tenth convolutional layer is also connected to the second input end of the second multiplication module, the output end of the first multiplication module is connected to the first input end of the fourth fusion module, and the output end of the second multiplication module is connected to the second input end of the fourth fusion module.

[0049] As a preferred solution, the multimodal Transformer layer includes: a first normalization layer, a second normalization layer, a third normalization layer, a fourth normalization layer, a MAAE module, a first MLP layer and a second MLP layer.

[0050] Among them, the output end of the first normalization layer is connected to the first input end of the MAAE module, the output end of the second normalization layer is connected to the second input end of the MAAE module, the first output end of the MAAE module is connected to the first input end of the fifth fusion module, the input end of the first normalization layer is connected to the second input end of the fifth fusion module, the second output end of the MAAE module is connected to the first input end of the sixth fusion module, the input end of the second normalization layer is connected to the second input end of the sixth fusion module, the output end of the fifth fusion module is connected to the input end of the third normalization layer, the output end of the third normalization layer is connected to the input end of the first MLP layer, and the sixth fusion module The output end of the fourth normalization layer is connected to the input end of the second MLP layer, the output end of the first MLP layer is connected to the first input end of the seventh fusion module, the output end of the fifth fusion module is also connected to the second input end of the seventh fusion module, the output end of the second MLP layer is connected to the first input end of the eighth fusion module, the output end of the sixth fusion module is also connected to the second input end of the eighth fusion module, the output end of the seventh fusion module serves as the first output end of the multimodal Transformer layer, and the output end of the eighth fusion module serves as the second output end of the multimodal Transformer layer.

[0051] As a preferred solution, the MAAE module includes: a first Q-key mapping module, a second Q-key mapping module, a first K-key mapping module, a second K-key mapping module, a first V-key mapping module, a second V-key mapping module, a first self-attention activation module, a second self-attention activation module, a first cross-attention activation module and a second cross-attention activation module.

[0052] The output end of the first normalization layer is connected to the input ends of the first Q key mapping module, the first K key mapping module and the first V key mapping module respectively, the second output end of the second normalization layer is connected to the input ends of the second Q key mapping module, the second K key mapping module and the second V key mapping module respectively, the output ends of the first Q key mapping module, the first K key mapping module and the first V key mapping module are connected to the input end of the first self-attention activation module respectively, the output ends of the first Q key mapping module, the second K key mapping module and the second V key mapping module are connected to the input end of the first cross-attention activation module respectively, the output end of the first self-attention activation module is connected to the first input end of the ninth fusion module, and the first cross-attention The output end of the activation module is connected to the second input end of the ninth fusion module, and the output end of the ninth fusion module serves as the first output end of the MAAE module. The output ends of the first K-key mapping module, the first V-key mapping module and the second Q-key mapping module are respectively connected to the input end of the second cross-attention activation module, and the output ends of the second Q-key mapping module, the second K-key mapping module and the second V-key mapping module are respectively connected to the input end of the second self-attention activation module. The output end of the second cross-attention activation module is connected to the first input end of the tenth fusion module, and the output end of the second self-attention activation module is connected to the second input end of the tenth fusion module, and the output end of the tenth fusion module serves as the second output end of the MAAE module.

[0053] In a second aspect, a computer-readable storage medium stores a computer program, which, when executed by a processor, implements a VLSI congestion prediction method based on dual multimodal fusion as described in any one of the first aspects.

[0054] According to a third aspect, a computer device includes:

[0055] Memory, used to store instructions.

[0056] The processor is used to execute the instructions so that the computer device performs the operations of the VLSI congestion prediction method based on dual multi-modal fusion as described in any one of the first aspects.

[0057] Beneficial effects: The present invention provides a VLSI congestion prediction method, device and storage medium based on dual multimodal fusion, which can effectively capture local and global placement information at multiple scales. Through multimodal fusion, the model can better understand the correlation between grid-based layout information and netlist knowledge, thereby significantly improving the accuracy of congestion prediction. In terms of feature extraction, the early feature fusion (EFF) method is adopted to fuse the netlist features extracted by HGCN with the grid-based layout features extracted by CNN, and these fused features are input into the cascade decoder to restore the details of the layout. In order to further enhance the fusion of multimodal features, a deep feature fusion (DFF) method is proposed, which combines multiple visual Transformer layers based on adaptive attention enhancement technology, uses self-attention (SA) to enhance intra-modal features, and cross-attention (CA) is used to perform cross-modal feature fusion between netlist and layout features. The present invention can quickly obtain high-quality and high-precision congestion prediction results in complex large-scale integrated circuit designs, guide layout to achieve routability optimization and effectively reduce wiring congestion. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 The present invention is a flowchart of a VLSI congestion prediction method based on dual multi-modal fusion.

[0059] Figure 2 It is a schematic diagram of the structure of the VLSI congestion prediction model proposed in the present invention.

[0060] Figure 3 Schematic diagram of four grid-based two-dimensional layout features constructed in the present invention.

[0061] Figure 4 Schematic diagram of the characteristics of the heterogeneous graph neural network constructed in the present invention.

[0062] Figure 5 is the node on the heterogeneous graph convolution in the present invention Diagram schematic of message update.

[0063] Figure 6 Schematic diagram of the network architecture for multi-scale feature extraction and fusion proposed in the present invention.

[0064] Figure 7 Schematic diagram of the EFF structure of multimodal feature fusion proposed in the present invention.

[0065] Figure 8 Detailed architecture diagram of a multimodal Transformer layer in the present invention.

[0066] Fig. 9 Detailed architecture diagram of the MAAE module proposed in DFF in the present invention. DETAILED DESCRIPTION

[0067] The following is a clear and complete description of the technical solutions in the examples of the present invention in conjunction with the accompanying drawings in the examples of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the protection scope of the present invention.

[0068] The present invention will be further described below in conjunction with specific embodiments.

[0069] Embodiment 1:

[0070] This embodiment introduces a VLSI congestion prediction method based on dual multi-modal fusion, including the following steps:

[0071] Step S1, preprocessing is performed according to the input VLSI circuit layout and netlist information.

[0072] Step S2, constructing grid-based two-dimensional layout features and heterogeneous graph-based netlist features according to the preprocessed input information.

[0073] In step S3, a convolutional neural network (CNN) is used to extract layout features and a heterogeneous graph neural network (HGNN) is used to extract netlist features to achieve multi-scale feature extraction.

[0074] Step S4, using the early feature fusion (EFF) method to merge the extracted features to achieve multimodal feature fusion.

[0075] In step S5, the deep feature fusion (DFF) method is used to enhance the intra-modal features and further perform cross-modal feature fusion based on the VisionTransformer layer of the adaptive attention enhancement technique.

[0076] Step S6, after deep feature fusion, upsampling and skip connection are performed to implement cascade decoding to restore the output features to the predicted congestion map results.

[0077] Furthermore, in step S1, preprocessing is performed according to the input VLSI circuit layout and netlist information as follows:

[0078] For a given layout, the congestion prediction model needs to be based on the input circuit layout information X placement and circuit diagram netlist G netlist Provides the highest possible accuracy in predicting congestion maps.

[0079] X placement ∈M W×H×C +G netlist →Y congestion ∈MW×H×1 (1)

[0080] Therefore, the circuit layout information X placement It is modeled as a grid-based two-dimensional module placement feature M, where W represents the feature width, H represents the feature height, and C represents the number of features. Each two-dimensional feature represents the circuit layout information obtained from different angles. Circuit diagram netlist G netlist Based on the circuit diagram netlist information, it is modeled to reflect the connections between the cells / macros in each line network. The final output is the same as the input circuit layout information X placement Congestion prediction results Y with the same size congestion , contains pixel values ​​reflecting the congestion level of each grid.

[0081] Furthermore, the specific method of step S2 is as follows:

[0082] Step 2.1, grid-based 2D layout feature construction: Generally speaking, the distribution and congestion of macro modules, cells, wire nets, and routing resources are strongly correlated. To this end, the placement canvas is divided into W×H grids to capture this strongly correlated information. Each grid g i,j (i∈{1,2,…,W},j∈{1,2,…,H}) all contain unique features. Specifically, the following four grid-based two-dimensional layout features are collected and constructed, such as Figure 3 shown.

[0083] f Macro (i,j)(Macro Mask) indicates whether a grid is occupied by a macro module. o,j The macroblock mask in is defined as follows:

[0084]

[0085] f Rudy (i,j)(RUDY Map) provides a fast congestion prediction method based on the current layout and netlist. and denotes the left, right, bottom, and top edges of the smallest bounding box that contains all pins in net e. Let For a single wire mesh e, its RUDY (Uniform Rectangular Density) is defined as:

[0086]

[0087] Then, for the grid g i,j , its RUDY mask is defined as follows:

[0088] f Rudy (i,j)=∑ e∈netlist fe (i,j) (4)

[0089] Among them, netlist represents netlist information.

[0090] f Pin (i,j)(Pin Density) represents the area utilization of the pins on each grid. A high pin density value indicates that the routing resources around the grid may be overused. Let p e Represents the pin of the line net e, and the grid g i,j The pin density map in is defined as follows:

[0091]

[0092] f Cell (i,j)(Cell Density) is defined as each grid g i,j The number of cells in , where the cell density map is represented as follows:

[0093] f Cell (i,j) = #Cells (6)

[0094] Among them, #Cells represents the number of cells in the grid.

[0095] Therefore, the size of the grid-based 2D layout feature is W×H×4.

[0096] Step 2.2, heterogeneous graph neural network feature construction: Considering the inconsistency of the netlists of different circuit designs, a heterogeneous graph is constructed to uniformly model the netlists and grid-based layouts of different circuit designs. Hetero = {V g ,V n ,E gg ,E ng},V g and V n Denotes the node set of the grid and network respectively. gg represents the connection between grids, E ng Represents the connection between the grid and the wire net. Figure 4 (a) illustrates the constructed heterogeneous graph neural network model. The heterogeneous graph includes two types of nodes: grid representing a grid and net representing a wire network.

[0097] Three different types of connection relations: (1) grid to grid, (2) grid to wirenet, and (3) wirenet to grid. This setting allows different types of nodes and edges to be distinguished in different representation spaces, thus facilitating message passing in heterogeneous graph structures.

[0098] Connection between grids and girds: Figure 4 (b) If the grid A cell in the grid is connected to the grid through the same wire net Another unit in the Add on and In addition, each grid node contains three kinds of information Where x and y are the horizontal and vertical coordinates of the relative position of the grid node in the placement area, Indicates the number of cells contained in the grid node. The edge set E between grids gg Representing cell connectivity relationships, it integrates netlists from different circuit designs. Therefore, it enhances the model to explore various circuit layouts and improves the prediction capability.

[0099] Connections between grids and nets: Figure 4 (c) shows that if a cell belongs to the grid and net There is an edge Add on and Each net node contains four types of information Let B e is the minimum bounding box that contains all pins in net e. Then, W bbox and H bbox Indicates B e The width and height of A bbox Indicates B e The area (i.e. B e The number of grids in is the number of cells belonging to the net node. The edge set between Grid and net and E ng It fills in the connection details of the wire network and integrates the logical diagram (netlist) and physical information (layout).

[0100] Furthermore, the specific method of step S3 is as follows:

[0101] The entire congestion prediction model, such as Figure 2 As shown in the figure, accurate congestion prediction is achieved through three main parts: feature extraction and fusion module, deep feature fusion module and cascade decoding module.

[0102] In the feature extraction and fusion module, the CNN branch uses convolutional neural networks to extract local features of the grid layer by layer, and captures features of different scales through downsampling. At the same time, the HGCN branch uses heterogeneous graph convolution and graph attention mechanisms to pass messages between grid nodes and wire network nodes, thereby effectively capturing graph structure information and contextual relationships.

[0103] Each layer of CNN convolution and the corresponding heterogeneous graph convolution output are used to achieve early feature fusion through the EFF method. The fused features are passed to the next layer of CNN convolution in turn and passed to the cascade decoding module through jump connections. For the HGCN branch, in addition to EFF, each layer of heterogeneous graph convolution output is also passed layer by layer. Finally, the heterogeneous graph convolution output of the last layer is passed to the embedding layer for processing together with the final EFF output as the input of the deep feature fusion module.

[0104] In the deep feature fusion stage, the self / cross attention mechanism is used to further fuse the features from CNN and HGCN, and the feature expression is enhanced through the multi-layer Transformer module to capture long-range dependencies and multimodal information.

[0105] Finally, in the cascade decoding module, multiple decoding layers gradually splice the early feature fusion output of EFF, and upsample them in turn to restore higher resolution feature maps, and finally generate an accurate congestion map.

[0106] Specifically including: feature extraction and fusion module, including: the first convolution layer, the second convolution layer, the third convolution layer, the fourth convolution layer, the first EFF module, the second EFF module, the third EFF module, the fourth EFF module, the first heterogeneous graph convolution layer, the second heterogeneous graph convolution layer, the third heterogeneous graph convolution layer, the fourth heterogeneous graph convolution layer and the fifth convolution layer.

[0107] Among them, the output end of the first convolutional layer is connected to the first input end of the first EFF module, the output end of the first heterogeneous graph convolutional layer is connected to the second input end of the first EFF module, the output end of the first heterogeneous graph convolutional layer is also connected to the input end of the second heterogeneous graph convolutional layer, and the output end of the first EFF module is connected to the input end of the second convolutional layer.

[0108] The output end of the second convolutional layer is connected to the first input end of the second EFF module, the output end of the second heterogeneous graph convolutional layer is connected to the second input end of the second EFF module, the output end of the second heterogeneous graph convolutional layer is also connected to the input end of the third heterogeneous graph convolutional layer, and the output end of the second EFF module is connected to the input end of the third convolutional layer.

[0109] The output end of the third convolutional layer is connected to the first input end of the third EFF module, the output end of the third heterogeneous graph convolutional layer is connected to the second input end of the third EFF module, the output end of the third heterogeneous graph convolutional layer is also connected to the input end of the fourth heterogeneous graph convolutional layer, and the output end of the third EFF module is connected to the input end of the fourth convolutional layer.

[0110] The output end of the fourth convolutional layer is connected to the first input end of the fourth EFF module, the first output end of the fourth heterogeneous graph convolutional layer is connected to the second input end of the fourth EFF module, and the second output end of the fourth heterogeneous graph convolutional layer is connected to the input end of the fifth convolutional layer.

[0111] Deep feature fusion module, including: embedding layer, several multimodal Transformer layers.

[0112] The embedding layer is connected to several multimodal Transformer layers in sequence, the input end of the fourth EFF module is connected to the first input end of the embedding layer, and the output end of the fifth convolutional layer is connected to the second input end of the embedding layer.

[0113] The cascade decoding module includes: a sixth convolution layer, a seventh convolution layer, an eighth convolution layer, a ninth convolution layer, a first upsampling layer, a second upsampling layer, a third upsampling layer, and a fourth upsampling layer.

[0114] Among them, the first output end of the last multimodal Transformer layer is connected to the input end of the sixth convolutional layer, the output end of the sixth convolutional layer is connected to the input end of the first upsampling layer, the output end of the third EFF module and the output end of the first upsampling layer are respectively connected to the two input ends of the first fusion module, the output end of the first fusion module is connected to the input end of the seventh convolutional layer, the output end of the seventh convolutional layer is connected to the input end of the second upsampling layer, the output end of the second EFF module and the output end of the second upsampling layer are respectively connected to the two input ends of the second fusion module, the output end of the second fusion module is connected to the input end of the eighth convolutional layer, the output end of the eighth convolutional layer is connected to the input end of the third upsampling layer, the output end of the first EFF module and the output end of the third upsampling layer are respectively connected to the two input ends of the third fusion module, the output end of the third fusion module is connected to the input end of the ninth convolutional layer, and the output end of the ninth convolutional layer is connected to the input end of the fourth upsampling layer.

[0115] Among them, Figure 7 As shown, the first EFF module, the second EFF module, the third EFF module and the fourth EFF module all include: a tenth convolutional layer, a first global average pooling layer, a second global average pooling layer, a first fully connected layer, a second fully connected layer, a first activation function and a second activation function.

[0116] Among them, the first global average pooling layer, the first fully connected layer and the first activation function are connected in series, the input end of the first global average pooling layer is connected to the first input end of the first multiplication module, and the output end of the first activation function is connected to the second input end of the first multiplication module. The tenth convolutional layer, the second global average pooling layer, the second fully connected layer and the second activation function are connected in series, the output end of the second activation function is connected to the first input end of the second multiplication module, the output end of the tenth convolutional layer is also connected to the second input end of the second multiplication module, the output end of the first multiplication module is connected to the first input end of the fourth fusion module, and the output end of the second multiplication module is connected to the second input end of the fourth fusion module.

[0117] Among them, Figure 8 As shown, the multimodal Transformer layer includes: a first normalization layer, a second normalization layer, a third normalization layer, a fourth normalization layer, a MAAE module, a first MLP layer (multi-layer perceptron) and a second MLP layer.

[0118] Among them, the output end of the first normalization layer is connected to the first input end of the MAAE module, the output end of the second normalization layer is connected to the second input end of the MAAE module, the first output end of the MAAE module is connected to the first input end of the fifth fusion module, the input end of the first normalization layer is connected to the second input end of the fifth fusion module, the second output end of the MAAE module is connected to the first input end of the sixth fusion module, the input end of the second normalization layer is connected to the second input end of the sixth fusion module, the output end of the fifth fusion module is connected to the input end of the third normalization layer, the output end of the third normalization layer is connected to the input end of the first MLP layer, and the sixth fusion module The output end of the fourth normalization layer is connected to the input end of the second MLP layer, the output end of the first MLP layer is connected to the first input end of the seventh fusion module, the output end of the fifth fusion module is also connected to the second input end of the seventh fusion module, the output end of the second MLP layer is connected to the first input end of the eighth fusion module, the output end of the sixth fusion module is also connected to the second input end of the eighth fusion module, the output end of the seventh fusion module serves as the first output end of the multimodal Transformer layer, and the output end of the eighth fusion module serves as the second output end of the multimodal Transformer layer.

[0119] Among them, Fig. 9 As shown, the MAAE module includes: a first Q-key mapping module, a second Q-key mapping module, a first K-key mapping module, a second K-key mapping module, a first V-key mapping module, a second V-key mapping module, a first self-attention activation module, a second self-attention activation module, a first cross-attention activation module and a second cross-attention activation module.

[0120] The output end of the first normalization layer is connected to the input ends of the first Q key mapping module, the first K key mapping module and the first V key mapping module respectively, the second output end of the second normalization layer is connected to the input ends of the second Q key mapping module, the second K key mapping module and the second V key mapping module respectively, the output ends of the first Q key mapping module, the first K key mapping module and the first V key mapping module are connected to the input end of the first self-attention activation module respectively, the output ends of the first Q key mapping module, the second K key mapping module and the second V key mapping module are connected to the input end of the first cross-attention activation module respectively, the output end of the first self-attention activation module is connected to the first input end of the ninth fusion module, and the first cross-attention The output end of the activation module is connected to the second input end of the ninth fusion module, and the output end of the ninth fusion module serves as the first output end of the MAAE module. The output ends of the first K-key mapping module, the first V-key mapping module and the second Q-key mapping module are respectively connected to the input end of the second cross-attention activation module, and the output ends of the second Q-key mapping module, the second K-key mapping module and the second V-key mapping module are respectively connected to the input end of the second self-attention activation module. The output end of the second cross-attention activation module is connected to the first input end of the tenth fusion module, and the output end of the second self-attention activation module is connected to the second input end of the tenth fusion module, and the output end of the tenth fusion module serves as the second output end of the MAAE module.

[0121] Step 3.1, for the CNN branch: Use four convolutional layers based on the ResNet50 architecture to extract grid-based 2D layout features. After each convolutional layer, the size of the hidden feature will become Among them l e ∈{1,2,3,4} represents the depth of the convolutional layer, Indicates the number of channels after each layer.

[0122] Step 3.2, for the HGCN branch: Use four heterogeneous graph convolution layers to pass messages between nodes and edges of the heterogeneous graph neural network. In order to reduce the computational complexity of HGCN, the heterogeneous graph is divided into 16 tiles according to the grid nodes. Each block contains W / 4×H / 4 grid nodes, and its corresponding net nodes are also divided. After each heterogeneous graph convolution layer, all grid node features are extracted from these tiles to form HGNN features in preparation for the next step. During the training process, the Graph Attention Network (GAT) is used as the heterogeneous graph convolution layer of HGCN to learn and update node features. The number of channels after the four heterogeneous graph convolution layers are {64, 256, 512, 1024}.

[0123] Figure 5Describes the mesh nodes Example of an update on a heterogeneous graph convolutional layer. This convolution operation can be divided into two parts: message generation and neighbor aggregation. For each network node in the graph, message generation involves collecting feature representations (messages) about all its neighbor nodes. Each node Through the learnable weight matrix W gg Transform to generate new feature representation:

[0124]

[0125] Based on the above formula, the feature representations of all neighbor nodes are collected, and then these feature representations are used to calculate the attention score weights of each neighbor node. and First, calculate the attention coefficient By Instruction right Relative importance of:

[0126]

[0127] Where: a is a learnable weight vector, [·||·] represents concatenation. and Respectively represent the connection E gg and E ng middle The set of neighbor nodes of . Then, After a layer of softmax function, at all neighbor nodes It is standardized on and The attention score between The calculation is as follows:

[0128]

[0129] In addition, the message aggregation part is introduced to aggregate all neighbor messages to the central node For message passing between grids, based on attention scores Grid nodes The message aggregation can be expressed as:

[0130]

[0131] where n i and n j They are and The number of neighbor nodes. Used to ensure the scale consistency of embeddings in graph convolution operations. For message passing between grids and nets, based on attention scores Grid nodes The message aggregation can be expressed as:

[0132]

[0133] Where W ng Indicates that on the side E ng The above two aggregation messages are used to update the central node representation as follows:

[0134]

[0135] in It represents the sum of two aggregate feature representations of a node.

[0136] Similarly, for message passing between grids and nets, based on the attention scores Net nodes The message aggregation is represented as:

[0137]

[0138] in Indicates the connection between grids and nets The set of neighbor nodes. gn Represents the weight matrix. The aggregated message is used to update the central node representation as follows:

[0139]

[0140] In HGNN The input node features of are three dimensions, and The input features of are four-dimensional. After each heterogeneous graph convolution layer, the features from neighboring nodes are combined with the features in the CNN branch. Maintain consistent dimensions.

[0141] Furthermore, the specific method of step S4 is as follows:

[0142] like Figure 7 As shown, since each EFF module (feature extraction layer) has two branches, the EFF module is used for multimodal feature fusion. Figure 7 The features extracted by CNN, the features extracted by HGCN, and the size of the output EFF features after each feature extraction layer are shown respectively. Specifically, the output of CNN is directly used as the layout feature, and the lth e The size of the feature extraction layer is At the same time, all grid nodes V output by HGCN g According to their original positions, they are organized into Then, a new convolutional layer is applied to extract the HGNN features, thus obtaining the size At this point, the sizes of the two features are consistent. The layout features and netlist features are then processed by two average pooling layers and two fully connected (FC) layers, followed by ReLU or sigmoid activation functions. By multiplying the input features with the processed features, the two modal features can effectively enhance important information and reduce less relevant information. In addition, the two modal features are element-wise aligned to form a multi-scale EFF feature, which integrates the netlist knowledge into the multi-scale layout features of the multi-modal interaction subspace, maintaining Finally, these fused features are directly fed back into the cascaded decoding layers with corresponding scales to recover the congestion map.

[0143] Furthermore, the specific method of step S5 is as follows:

[0144] like Figure 2 and Figure 7 As shown in the figure, two hidden features are selected as outputs after the feature extraction and fusion modules, which are the EFF features and HGNN features that incorporate netlist knowledge into the grid-based layout. In order to keep consistent with the size of the EFF features [W / 16, H / 16, 1024], the HGNN features are passed through a convolutional layer. Subsequently, the two modal features are used as inputs to the DFF (deep feature fusion module) to enhance and fuse the multimodal features. The DFF contains an embedding layer for processing input features and a multimodal Transformer layer for fusing multimodal features. First, the sizes of the two modal features are reduced to [W / 16, H / 16, C trans ], and then reshaped into [L,C trans ](L=W / 16×H / 16), where C trans is the number of channels designed specifically to better fit the Transformer layer based on the Vision Transformer architecture. After the above processing, the two multimodal features are defined as two-dimensional sequences x0 and y0, respectively, and then they are simultaneously input into the L d In a multimodal Transformer layer. Let l d ∈{1,2,…,L d} indicates the first d Multimodal Transformer layers. The input of each multimodal Transformer layer is The corresponding output is express.

[0145] In addition, the proposed multimodal Transformer layer is implemented based on the visual Transformer, and its structure is as follows Figure 8 As shown in Figure 2. The structure consists of a Multihead Adaptive Attention Enhancement (MAAE) block and two MLP blocks. The MAAE module includes multi-head self-attention (SA) for enhancing intra-modal features and multi-head cross-attention (CA) for performing cross-modal feature fusion on netlist and grid-based layout features. The output of the MAAE block is defined as follows:

[0146]

[0147] where LN(·) represents the normalization layer. Specifically, MAAE converts and Then, these two feature sequences are projected onto their corresponding linear and Multiply them together to get two sets of multimodal information matrices (Q x ,K x ,V x ) and (Q y ,K y ,V y ). In addition, these two sets of multimodal matrices are input into multi-head SA and multi-head CA, as Fig. 9 As shown. The multi-head SA block is used to capture the x ,K x ,V x ) or (Q y ,K y ,V y ). In this case, each element pays attention to all other elements in its own sequence, thereby enhancing the model's understanding of the global HGNN features and EFF features. The multi-head CA module allows two independent sequences to interact with each other, enabling the model to learn the relationship between the netlist and grid-based layout information. This enhances the robustness of the model in understanding complex features. The outputs of the multi-head SA and CA are:

[0148]

[0149] where d k is the normalization parameter, (·) T is the matrix transpose operator. The outputs of multi-head SA and CA can be calculated similarly. Next, the following adaptive mechanism is proposed to fuse the outputs of multi-head SA and CA:

[0150]

[0151] Where: λ is a learnable weighting coefficient used to balance the contributions of multi-head SA and CA.

[0152] Further, the MLP block consists of two FC layers followed by a ReLU activation function. The output equation of the multimodal Transformer layer processed by the MLP block can be summarized as:

[0153]

[0154] After all the multimodal Transformer layers, the last multimodal Transformer layer As the output of DFF. In summary, the proposed DFF further enhances the model's ability to predict congestion by fusing and enhancing the multimodal features of grid-based placement and netlist.

[0155] Furthermore, the specific method of step S6 is as follows:

[0156] The output features of DFF are fed into the cascaded decoding module to recover the final congestion graph Y congestion .like Figure 2 As shown in Figure 2, the proposed cascade decoding module has the same multi-scale features as the feature extraction and fusion modules. The cascade decoding module (Decoder) mainly consists of three parts: upsampling layer, skip-connected EFF features and CNN layer. Specifically, the CNN layer is first used to change the output of DFF to [W / 2 4 ,H / 2 4 ,512]. Next, the upsampling layer doubles the feature size to [W / 2 3 ,H / 2 3 ,512]. Further, this feature is connected with the corresponding skip-connected EFF feature, and the number of channels is reduced by 2 times through the CNN layer, and the feature size is [W / 2 3 ,H / 2 3 ,256]. This process is repeated 3 times, and the final CNN layer changes the number of output channels to 1. Finally, the cascade decoder passes through the upsampling layer to restore the accurate congestion map of size [W,H,1]. This design enables the model to make full use of the multi-model fusion function of the encoder and can effectively capture the layout and netlist information based on multi-scale grids to improve the accuracy of congestion prediction.

[0157] Example 2

[0158] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements a VLSI congestion prediction method based on dual multi-modal fusion as described in any one of Embodiment 1.

[0159] Example 3

[0160] A computer device comprising:

[0161] Memory, used to store instructions.

[0162] The processor is used to execute the instructions so that the computer device performs the operations of a VLSI congestion prediction method based on dual multi-modal fusion as described in any one of Embodiment 1.

[0163] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A VLSI congestion prediction method based on dual multi-modal fusion, characterized by: Specifically include: Step 1, obtaining layout information and netlist information of the circuit to be predicted; Step 2, constructing a grid-based two-dimensional layout feature and a heterogeneous graph-based netlist feature according to the layout information and the netlist information; Step 3: Input the two-dimensional layout features and netlist features into the congestion prediction model to obtain a congestion graph of the circuit to be predicted.

2. The VLSI congestion prediction method based on dual multi-modal fusion according to claim 1, characterized in that: The method of constructing a grid-based two-dimensional layout feature specifically includes: The circuit canvas to be predicted is divided into W×H grids, where each grid is defined as g i,j ,i∈{1,2,...,W}, j∈{1,2,...,H}; Obtain four layout features from each grid and construct a grid-based two-dimensional layout feature; Among them, four layout features include: macro module mask f Macro (i, j), uniform rectangular density mask f Rudy (i, j), pin density map f Pin (i, j) and the cell density map f Cell (i, j).

3. The VLSI congestion prediction method based on dual multi-modal fusion according to claim 1, characterized in that: The method for constructing a netlist feature based on a heterogeneous graph specifically includes: The circuit to be predicted is constructed as a heterogeneous graph, where the heterogeneous graph is defined as G Hetero = {V g , V n , E gg , E ng },V g and V n Represents the node set of the grid and network respectively; E gg represents the connection between grids, E ng Represents the connection between the grid and the wire net; Obtain the features of each grid node and each wire network node, and construct a netlist feature based on a heterogeneous graph; The grid node characteristics are defined as Among them, x and y represent the horizontal and vertical coordinates of the relative position of the grid node in the placement area, respectively. Indicates the number of cells contained in the grid node; The network node feature is defined as Among them, W bbox and H bbox Indicates B e The width and height of A bbox Indicates B e The area of is the number of cells belonging to the e-node of the network, B e is the minimum bounding box that contains all pins in net e.

4. The VLSI congestion prediction method based on dual multi-modal fusion according to claim 2, characterized in that: The macroblock mask f Macro The expression of (i, j) is as follows: The uniform rectangular density mask f Rudy The expression of (i,j) is as follows: f Rudy (i,j)=∑ e∈netlist f e (i,j); Among them, netlist represents netlist information; in, and denote the left, right, bottom and top edges of the smallest bounding box containing all pins in net e, respectively; The pin density diagram f Pin The expression of (i,j) is as follows: Among them, p e Indicates the pin of the line net e; The cell density map f Cell The expression of (i,j) is as follows: f Cell (i,j)=#Cells Among them, #Cells represents the number of cells in the grid.

5. The VLSI congestion prediction method based on dual multi-modal fusion according to claim 1, characterized in that: The congestion prediction model specifically includes: a feature extraction and fusion module, a deep feature fusion module and a cascade decoding module; The feature extraction and fusion module includes: a first convolution layer, a second convolution layer, a third convolution layer, a fourth convolution layer, a first EFF module, a second EFF module, a third EFF module, a fourth EFF module, a first heterogeneous graph convolution layer, a second heterogeneous graph convolution layer, a third heterogeneous graph convolution layer, a fourth heterogeneous graph convolution layer and a fifth convolution layer; The output end of the first convolutional layer is connected to the first input end of the first EFF module, the output end of the first heterogeneous graph convolutional layer is connected to the second input end of the first EFF module, the output end of the first heterogeneous graph convolutional layer is also connected to the input end of the second heterogeneous graph convolutional layer, and the output end of the first EFF module is connected to the input end of the second convolutional layer; The output end of the second convolutional layer is connected to the first input end of the second EFF module, the output end of the second heterogeneous graph convolutional layer is connected to the second input end of the second EFF module, the output end of the second heterogeneous graph convolutional layer is also connected to the input end of the third heterogeneous graph convolutional layer, and the output end of the second EFF module is connected to the input end of the third convolutional layer; The output end of the third convolutional layer is connected to the first input end of the third EFF module, the output end of the third heterogeneous graph convolutional layer is connected to the second input end of the third EFF module, the output end of the third heterogeneous graph convolutional layer is also connected to the input end of the fourth heterogeneous graph convolutional layer, and the output end of the third EFF module is connected to the input end of the fourth convolutional layer; The output end of the fourth convolutional layer is connected to the first input end of the fourth EFF module, the first output end of the fourth heterogeneous graph convolutional layer is connected to the second input end of the fourth EFF module, and the second output end of the fourth heterogeneous graph convolutional layer is connected to the input end of the fifth convolutional layer; The deep feature fusion module includes: an embedding layer and several multimodal Transformer layers; The embedding layer is connected to a plurality of multimodal Transformer layers in sequence, an input end of the fourth EFF module is connected to a first input end of the embedding layer, and an output end of the fifth convolutional layer is connected to a second input end of the embedding layer; The cascade decoding module includes: a sixth convolution layer, a seventh convolution layer, an eighth convolution layer, a ninth convolution layer, a first upsampling layer, a second upsampling layer, a third upsampling layer, and a fourth upsampling layer; Among them, the first output end of the last multimodal Transformer layer is connected to the input end of the sixth convolutional layer, the output end of the sixth convolutional layer is connected to the input end of the first upsampling layer, the output end of the third EFF module and the output end of the first upsampling layer are respectively connected to the two input ends of the first fusion module, the output end of the first fusion module is connected to the input end of the seventh convolutional layer, the output end of the seventh convolutional layer is connected to the input end of the second upsampling layer, the output end of the second EFF module and the output end of the second upsampling layer are respectively connected to the two input ends of the second fusion module, the output end of the second fusion module is connected to the input end of the eighth convolutional layer, the output end of the eighth convolutional layer is connected to the input end of the third upsampling layer, the output end of the first EFF module and the output end of the third upsampling layer are respectively connected to the two input ends of the third fusion module, the output end of the third fusion module is connected to the input end of the ninth convolutional layer, and the output end of the ninth convolutional layer is connected to the input end of the fourth upsampling layer.

6. The VLSI congestion prediction method based on dual multi-modal fusion according to claim 5, characterized in that: The first EFF module, the second EFF module, the third EFF module and the fourth EFF module each include: a tenth convolutional layer, a first global average pooling layer, a second global average pooling layer, a first fully connected layer, a second fully connected layer, a first activation function and a second activation function; Among them, the first global average pooling layer, the first fully connected layer and the first activation function are connected in series, the input end of the first global average pooling layer is connected to the first input end of the first multiplication module, and the output end of the first activation function is connected to the second input end of the first multiplication module. The tenth convolutional layer, the second global average pooling layer, the second fully connected layer and the second activation function are connected in series, the output end of the second activation function is connected to the first input end of the second multiplication module, the output end of the tenth convolutional layer is also connected to the second input end of the second multiplication module, the output end of the first multiplication module is connected to the first input end of the fourth fusion module, and the output end of the second multiplication module is connected to the second input end of the fourth fusion module.

7. The VLSI congestion prediction method based on dual multi-modal fusion according to claim 5, characterized in that: The multimodal Transformer layer includes: a first normalization layer, a second normalization layer, a third normalization layer, a fourth normalization layer, a MAAE module, a first MLP layer and a second MLP layer; Among them, the output end of the first normalization layer is connected to the first input end of the MAAE module, the output end of the second normalization layer is connected to the second input end of the MAAE module, the first output end of the MAAE module is connected to the first input end of the fifth fusion module, the input end of the first normalization layer is connected to the second input end of the fifth fusion module, the second output end of the MAAE module is connected to the first input end of the sixth fusion module, the input end of the second normalization layer is connected to the second input end of the sixth fusion module, the output end of the fifth fusion module is connected to the input end of the third normalization layer, the output end of the third normalization layer is connected to the input end of the first MLP layer, and the sixth fusion module The output end of the fourth normalization layer is connected to the input end of the second MLP layer, the output end of the first MLP layer is connected to the first input end of the seventh fusion module, the output end of the fifth fusion module is also connected to the second input end of the seventh fusion module, the output end of the second MLP layer is connected to the first input end of the eighth fusion module, the output end of the sixth fusion module is also connected to the second input end of the eighth fusion module, the output end of the seventh fusion module serves as the first output end of the multimodal Transformer layer, and the output end of the eighth fusion module serves as the second output end of the multimodal Transformer layer.

8. The VLSI congestion prediction method based on dual multi-modal fusion according to claim 7, characterized in that: The MAAE module includes: a first Q key mapping module, a second Q key mapping module, a first K key mapping module, a second K key mapping module, a first V key mapping module, a second V key mapping module, a first self-attention activation module, a second self-attention activation module, a first cross-attention activation module, and a second cross-attention activation module; The output end of the first normalization layer is connected to the input ends of the first Q key mapping module, the first K key mapping module and the first V key mapping module respectively, the second output end of the second normalization layer is connected to the input ends of the second Q key mapping module, the second K key mapping module and the second V key mapping module respectively, the output ends of the first Q key mapping module, the first K key mapping module and the first V key mapping module are connected to the input end of the first self-attention activation module respectively, the output ends of the first Q key mapping module, the second K key mapping module and the second V key mapping module are connected to the input end of the first cross-attention activation module respectively, the output end of the first self-attention activation module is connected to the first input end of the ninth fusion module, and the first cross-attention The output end of the activation module is connected to the second input end of the ninth fusion module, and the output end of the ninth fusion module serves as the first output end of the MAAE module. The output ends of the first K-key mapping module, the first V-key mapping module and the second Q-key mapping module are respectively connected to the input end of the second cross-attention activation module, and the output ends of the second Q-key mapping module, the second K-key mapping module and the second V-key mapping module are respectively connected to the input end of the second self-attention activation module. The output end of the second cross-attention activation module is connected to the first input end of the tenth fusion module, and the output end of the second self-attention activation module is connected to the second input end of the tenth fusion module, and the output end of the tenth fusion module serves as the second output end of the MAAE module.

9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, a VLSI congestion prediction method based on dual multi-modal fusion as described in any one of claims 1 to 8 is implemented.

10. A computer device, characterized in that: include: A memory for storing instructions; A processor is used to execute the instructions so that the computer device performs the operation of the VLSI congestion prediction method based on dual multi-modal fusion as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Wiring layout method based on reinforcement learning

    CN120706359A

  • A routable linear layout method based on reinforcement learning

    CN120706359B