A chip layout method and system based on reinforcement learning and backend optimization
Through the chip layout method of reinforcement learning and back-end optimization, the use of position, line and view masks to capture circuit features, combined with feature fusion and policy network, the existing chip layout methods are solved, and more efficient layout quality and speed are achieved.
Patent Information
- Application Number
- CN202510293674.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The existing chip layout methods are inefficient and lack generalization. Traditional methods rely on manual intensive work. Learning-based methods have shortcomings in feature fusion and network training, resulting in inaccurate line length estimation and poor generalization ability.
Using a chip layout method based on reinforcement learning and back-end optimization, by building a layout optimization model, using position mask, line mask and view mask to capture circuit features, combining feature fusion modules and policy networks, iterative optimization is performed, residual networks and channel attention mechanisms are introduced, and GPU parallel calculation is used to set reward functions to improve exploration efficiency.
It improves the efficiency and quality of chip layout, solves the problem of overlapping layout, improves the stability and training speed of the model, enhances the ability to capture complex modes, and realizes more effective layout solutions exploration.
Smart Images

Figure CN120235107B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a chip layout method in the technical field of integrated circuit design, and in particular to a chip layout method based on reinforcement learning and back-end optimization, and also to a chip layout system based on reinforcement learning and back-end optimization. Background Art
[0002] As integrated circuits (ICs) continue to grow in complexity in line with Moore's Law, IC system design is becoming increasingly challenging. Electronic Design Automation (EDA) tools play a leading role in addressing these challenges in IC design. Chip layout is a key task in modern chip design, involving the placement of millions of circuit blocks on a two-dimensional chip canvas. This process is crucial for minimizing latency and power consumption. Traditionally, chip layout requires months of intensive work by hardware engineers to produce the layout, a time-consuming and labor-intensive approach. Over time, deep reinforcement learning has emerged as an emerging automated tool for addressing the chip layout problem. A number of experiments and studies have shown that while learning-based methods are still in their early stages, they are already producing promising results, significantly advancing the chip design process in an automated manner.
[0003] However, existing learning-based methods have some flaws. For example, learning-based reinforcement learning methods such as Graph Placement and DeepPR use hypergraphs to represent networks, but hypergraphs have scalability issues in fully encoding network list information. They discard the relative position information of pins, which may lead to inaccurate line length estimates. The MaskPlace method has made great improvements in extracting node information from the dataset, but the processing in the subsequent reinforcement learning value network and feature fusion is relatively rough. Only traditional CNN encoders are used in the value network, which has limited ability to capture global dependencies, low training efficiency, and poor generalization ability. In the fusion of global and local features, only the feature information of the last layer is fused, and some important information may be lost during network transmission. Therefore, existing chip layout methods have problems with low efficiency and insufficient generalization of the solutions. Summary of the Invention
[0004] In order to solve the technical problems of low efficiency and insufficient generalization of existing chip layout methods, the present invention provides a chip layout method and system based on reinforcement learning and backend optimization.
[0005] The present invention is implemented using the following technical solution: a chip layout method based on reinforcement learning and backend optimization, which includes the following steps:
[0006] S1: Build layout optimization model:
[0007] S1.1: Feature encode the node, position, and line net features of the macromodule in the netlist to generate the corresponding position mask, line mask, and view mask;
[0008] S1.2: fusing the position mask and the line mask to generate local features, and fusing the line mask and the view mask to generate global features;
[0009] S1.3: feature fusion of the local features and the global features to generate a comprehensive feature;
[0010] S1.4: Based on the embedding vectors of the comprehensive features and the global features, generate the probabilities of various layout optimization actions that should be taken, and calculate the incentives and their values corresponding to the updated states after each round of actions;
[0011] S1.5: Update the parameters of step S1.4 based on the newly stored experience;
[0012] S2: First, multiple random netlists are used as sample data to form an original data set, and then the layout optimization model is trained using the original data set;
[0013] S3: First, the trained layout optimization model is used to process the original circuit diagram file to obtain an optimized netlist file, and then the optimized initial circuit layout diagram is generated accordingly. Finally, the overlapping parts of the layout are eliminated to obtain the final layout diagram.
[0014] The present invention captures feature information in the netlist file of the circuit through position masks, line masks and view masks, and continuously updates parameters through exploration randomization. More low-level feature information can be retained through fusion, which helps the network capture complex patterns in chip layout at a deeper level, eliminates the interference caused by different specifications of various standard units in the netlist data, makes it easier to extract relevant features, and more effectively meets the input of the network. At the same time, it can solve some overlapping layouts that may exist in the layout, and improve the layout quality to a certain extent, improve reliability and practicality, solve the technical problems of low efficiency and insufficient generalization of the existing chip layout method, improve the layout efficiency of the chip, make up for the generalization of the solution, and explore more effective layout solutions with improved exploration efficiency.
[0015] As a further improvement of the above scheme, the layout optimization model includes a feature extraction module, a feature fusion module, an experience pool, a policy network and a value network; wherein the feature extraction module is used to feature encode the node, position and line network features in the macro module to generate the position mask, line mask and view mask accordingly; the feature fusion module is used to perform feature fusion on the position mask and the line mask to generate the local feature; the feature fusion module is also used to perform feature fusion on the line mask and the view mask to generate the global feature; the feature fusion module is also used to unify the dimensions of the local features and the global features, and perform feature fusion on the channel dimension to generate the comprehensive feature; the policy network is used to generate the probability of various layout optimization actions to be taken according to the embedding vector of the comprehensive feature; the value network is used to calculate the incentive and its value corresponding to the state updated according to each round of action according to the embedding vector of the global feature; the experience pool is used to store new states, actions, rewards and values as experience, and update the policy network and the value network according to the stored experience.
[0016] Furthermore, the feature fusion module includes a first fusion unit, a second fusion unit and a third fusion unit; the first fusion unit adopts a 1×1 convolution layer, and is used to perform feature fusion on the position mask and the line mask to generate the local feature; the second fusion unit adopts an enhancement layer structure combining a residual network and a channel attention mechanism layer, and is used to perform feature fusion on the line mask and the view mask to generate the global feature; the second fusion unit includes four feature extraction layers composed of convolution layers and residual connections and four encoding layers, and the channel attention mechanism layer is alternately distributed in the feature fusion module; the third fusion unit is used to unify the dimensions of the local features and the global features, and perform feature fusion in the channel dimension to generate the comprehensive feature.
[0017] Furthermore, the value network includes a front fully connected layer Linear1, a global information obtained through a convolution module, and three rear fully connected layers Linear2, Linear3, and Linear4; the functional layers of the value network are connected through ReLU functions; the front fully connected layer Linear1 has 768 input channels and 512 output channels; the three rear fully connected layers all have 512 input channels and 1 output channel;
[0018] The strategy network includes two convolutional layers, a transformation layer and a Softmax normalization layer; the strategy network outputs the corresponding action probability according to the embedding vector of the comprehensive feature, which includes the following steps: dividing the embedding vector of the comprehensive feature into three paths, two of which are sliced and reconstructed at different scales to obtain coarse-grained features and fine-grained features, and the other path is used to extract mask features to obtain a mask vector; the coarse-grained features are processed by a convolution layer with a coarse convolution to obtain a first feature; the fine-grained features are processed by a convolution layer with an attention module to obtain a second feature; the first feature and the second feature are firstly spliced, and then merged by the transformation layer to transform them into a fusion feature matrix of a specified dimension; the fusion feature matrix is masked by using the mask vector to obtain an action feature matrix; the action feature matrix is processed by a Softmax normalization layer to obtain an action probability;
[0019] The objective function of the policy network is:
[0020] policy(θ)=E[min(r(θ)A,max(clip(r(θ),1-∈,1+∈)A))]
[0021] Wherein, policy(θ) represents the objective function, E represents the expected value, Represents the probability ratio of the new and old strategies; A=G t -V t , represents the advantage function estimate at time step t, G t is the action value function, V t is the state value function; clip(r(θ),1-∈,1+∈) represents the clipping function, which limits r(θ) to the interval [1-∈,1+∈].
[0022] As a further improvement of the above solution, the position mask is an N×N binary matrix, where cells where components can be placed are marked as 1 and cells where components cannot be placed are marked as 0;
[0023] The line mask is an N×N continuous matrix, and the elements in the matrix represent the increase in the wire length of the macro module placed in each unit relative to the position in the previous round;
[0024] The view mask is an N×N binary matrix, where cells where macro modules are placed are marked as 1, and cells where no macro modules are placed are marked as 0.
[0025] As a further improvement of the above scheme, in step S2, a large number of random netlists containing multiple components are first obtained, and the random netlists are used as the sample data to constitute the original data set, and then the original data set is divided into a training set and a test volume to train and test the layout optimization model, and the model parameters of the layout optimization model that meets the target after training are saved.
[0026] As a further improvement to the above scheme, in step S3, the diagram file of the original circuit diagram to be optimized is input into the trained layout optimization model, the layout optimization model outputs the optimized netlist file, and then the corresponding optimized initial layout diagram of the circuit is generated based on the netlist file. Finally, an algorithm guided by a mask mechanism is used to optimize the layout overlapping parts to obtain the final layout diagram; wherein, the algorithm guided by the mask mechanism generates a wire mesh mask based on the initial macro position, and the wire mesh mask records the HPWL increment after the current macro is placed in each candidate grid, and selects the grid set with the smallest increment value; if there are several grid sets, the algorithm guided by the mask mechanism selects the grid closest to the macro from all grid sets; the algorithm guided by the mask mechanism marks the already placed area as placed by binarization, and eliminates the layout overlapping parts by iteratively exchanging the positions of macro modules.
[0027] Furthermore, step S2 also tests the layout optimization model using the original data set; in the training phase, the policy network and the value network are updated at each time point; when updating the value network, gradient backpropagation of the global feature embedding is stopped; in the testing phase, a probability matrix is obtained from the policy network in each round, and actions are sampled according to the probability matrix; when the sampling exceeds a threshold, the corresponding value is obtained from the line mask, and the half-circle field line length is calculated before executing the action.
[0028] Furthermore, the experience pool, the policy network, and the value network constitute a reinforcement learning framework; in step S1, the reinforcement learning framework is initialized to set a state space, an action space, a reward function, a state transition function, and a reset condition before training; the state space is the range of states that a component can reach, and the action space is the range of all actions that a component can perform; the reward function is the feedback signal received by the agent when it achieves a goal in the environment; the semi-circle length HPW□ is regarded as a reward, and the reward function includes:
[0029] R t =-a t WL(P,H)-b t C(PH)
[0030] HPWL=(max{xi}-min{x i})+(max{y i}-min{y i})
[0031] Among them, R t is the global reward, a t 、b t are the total wiring length and the super parameters of congestion information and congestion importance, respectively. P is the number of macro modules that have been placed, H is the connection information of the corresponding macro module, WL() is the length of the wires of all placed macro modules, and C() is the cost function; max{x i} is the maximum horizontal coordinate of all pins in the network, min{x i} is the minimum horizontal coordinate of all pins in the network, max{y i} is the maximum value of the vertical coordinate of all pins in the network, min{y i} is the minimum vertical coordinate of all pins in the network;
[0032] The state transfer function executes: selecting an action through the policy network and executing it in the virtual environment, and the virtual environment returns a new state and reward; storing the new state, action and reward in the experience pool as experience, and updating the policy network and the value network based on the stored experience; taking the completion of placement of all components as the reset condition, and clearing the accumulated reward after the reset condition is met.
[0033] The present invention also provides a chip layout system based on reinforcement learning and back-end optimization, which includes:
[0034] A model building subsystem, which is used to build a layout optimization model; the model building subsystem includes a feature extraction module, a feature fusion module, an experience pool module, a strategy network module and a value network module; the feature extraction module is used to feature encode the nodes, positions and line network features of the macro modules in the netlist to generate corresponding position masks, line masks and view masks; the feature fusion module is used to fuse the position mask and the line mask to generate local features, and fuse the line mask and the view mask to generate global features; the feature fusion module is also used to feature fuse the local features and the global features to generate comprehensive features; the strategy network module is used to generate the probability of various layout optimization actions to be taken based on the embedding vector of the comprehensive feature; the value network module is used to calculate the incentives and their values corresponding to the states updated according to each round of actions based on the embedding vector of the global feature; the experience pool module is used to store new states, actions, rewards and values as experience, and update the parameters of the strategy network module and the value network module based on the stored experience;
[0035] a training subsystem, configured to first use a plurality of random netlists as sample data to form an original data set, and then train the layout optimization model using the original data set;
[0036] The optimization subsystem is used to first process the chart file of the original circuit diagram through the trained layout optimization model to obtain an optimized netlist file, then generate the optimized circuit initial layout diagram accordingly, and finally eliminate the layout overlapping parts to obtain the final layout diagram.
[0037] Compared with existing chip layout methods and systems, the chip layout method and system based on reinforcement learning and backend optimization of the present invention has the following beneficial effects:
[0038] 1. This chip layout method based on reinforcement learning and back-end optimization captures feature information in the circuit netlist file through position masks, line masks and view masks. By exploring randomization and continuously iteratively updating parameters, more low-level feature information can be retained through fusion, which helps the network capture complex patterns in chip layout at a deeper level, eliminates the interference caused by different specifications of various standard units in the netlist data, makes it easier to extract relevant features, and more effectively meets the network input. At the same time, it can solve some overlapping layouts that may exist in the layout, and to a certain extent improve the layout quality, reliability and practicality, and solve the technical problems of low efficiency and insufficient generalization of the existing chip layout methods, improve the chip layout efficiency, make up for the generalization of the solution, and explore more effective layout solutions with improved exploration efficiency.
[0039] 2. This chip layout method, based on reinforcement learning and back-end optimization, transforms circuit spatial layout into a spatial optimization problem for component placement within a canvas space, and iteratively optimizes the problem using reinforcement learning. An optimized residual network and channel attention mechanism layer architecture are introduced during the global feature extraction process. By integrating the residual network and channel attention mechanism layer architectures, the model's ability to understand and process complex placement information is enhanced. Furthermore, the channel attention mechanism adaptively adjusts the weight of each channel, allowing the model to focus more on features that have a significant impact on layout decisions.
[0040] 3. This chip layout method, based on reinforcement learning and back-end optimization, allows GPUs to enhance the model's ability to capture diverse information through parallel computing, improving model stability and accelerating training and inference. Compared to existing technologies, this method significantly improves key performance indicators, such as minimizing the half-circuit length (HPWL) of circuit layouts, and significantly accelerates model training and layout inference. A mask-guided algorithm is used to address overlapping layouts, ensuring the legitimacy of macromodule placement and improving layout quality to a certain extent.
[0041] 4. This chip layout method based on reinforcement learning and back-end optimization sets a series of reward functions during the layout process, increasing the "agent" (intelligent body) in the reinforcement learning process to explore unknown layouts, avoiding falling into local optimality due to repeatedly choosing the same action; at the same time, by maximizing rewards, it abandons layout plans that are obviously unpromising, thereby improving the exploration efficiency of layout plans. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 This is a flowchart of a chip layout method based on reinforcement learning and backend optimization according to Example 1 of the present invention;
[0043] Figure 2 for Figure 1 A flow chart of the stages of the chip layout method based on reinforcement learning and back-end optimization;
[0044] Figure 3 for Figure 1 Detailed diagram of the local feature extraction module of the chip layout method based on reinforcement learning and backend optimization;
[0045] Figure 4 for Figure 1 Detailed diagram of the global feature extraction module of the chip layout method based on reinforcement learning and backend optimization;
[0046] Figure 5 for Figure 1 A detailed introduction to the chip layout method based on reinforcement learning and backend optimization in
[15] , which mainly scores the actions generated by the current "agent";
[0047] Figure 6 for Figure 1 The overall structure diagram of the chip layout method based on reinforcement learning and back-end optimization;
[0048] Figure 7 for Figure 1 Detailed architecture diagram of the policy network in the chip layout method based on reinforcement learning and backend optimization;
[0049] Figure 8 for Figure 1 A schematic diagram showing the detailed steps of the algorithm optimization part of the chip layout method based on reinforcement learning and back-end optimization;
[0050] Figure 9 for Figure 4 The detailed structure diagram of Resnet18+SE mentioned in it;
[0051] Figure 10 for Figure 9 The structural diagram of SE mentioned in ;
[0052] Figure 11 For the Figure 7 The generated action probability matrix is described in detail in the figure. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0054] Example 1
[0055] See also Figures 1 to 11 This embodiment provides a chip layout method based on reinforcement learning and backend optimization. This chip layout method is based on a reinforcement learning chip layout solution and a backend layout optimization algorithm. In this embodiment, the chip layout method includes steps S1, S2, and S3. In other embodiments, additional steps may be added to these steps.
[0056] Step S1: Constructing a layout optimization model. In this embodiment, the method for constructing the layout optimization model includes: S1.1: Feature encoding the nodes, positions, and wire mesh features of the macromodules in the netlist to generate corresponding position masks, wire masks, and view masks; S1.2: Fusion of the position masks and wire masks to generate local features, and fusion of the wire masks and view masks to generate global features; S1.3: Feature fusion of local features and global features to generate comprehensive features; S1.4: Based on the embedding vectors of the comprehensive features and global features, generating the probabilities of various layout optimization actions that should be taken, and calculating the incentives and values corresponding to the updated states after each round of actions; S1.5: Updating the parameters of step S1.4 based on the newly stored experience.
[0057] The position mask is an N×N binary matrix, in which cells where components can be placed are recorded as 1, and cells where components cannot be placed are recorded as 0. In this embodiment, the method for generating the position mask is as follows: first, create an N×N temporary matrix with all values 1, and then create an N×N weight matrix with all values 1. Then, by traversing the horizontal and vertical coordinates and the length and width of each macro module, the values of the overlapping places in the temporary matrix are changed to 0. Finally, the temporary matrix and the weight matrix are multiplied to obtain the position matrix representing the position mask. In this embodiment, the position matrix is generated by the temporary matrix and the weight matrix, and then the position matrix can be synchronously updated while the temporary matrix is searched in real time.
[0058] The line mask is an N×N continuous matrix, and the elements in the matrix represent the increase in the wire length of the macromodule placed in each unit relative to the previous position. The purpose of designing the line mask in this embodiment is to find the optimal position with the minimum increase in wire length.
[0059] The view mask is an N×N binary matrix, where cells where macro modules are already placed are marked as 1, and cells where no macro modules are placed are marked as 0. The purpose of the view mask in this embodiment is to determine whether each grid cell is occupied by a module.
[0060] In this embodiment, the layout optimization model includes a feature extraction module, a feature fusion module, an experience pool, a policy network and a value network, and may also include an environment setting module. The environment setting module is used to calculate a portion of the data required for model training and store it in the experience pool. The policy network is used to make decisions to determine what action the environment setting module should take. The value network assists the policy network module in making decisions based on the sampling data of the experience pool. Specifically, the feature extraction module is used to feature encode the nodes, positions and wire network features in the macro module to generate corresponding position masks, wire masks and view masks. The above-mentioned feature extraction module is used to perform pixel-level feature fusion on the nodes, positions and wire network features contained in the macro module of the netlist, and then generate corresponding local information and global information. The feature extraction module includes a series of operations on the current layout information (macro module position information, netlist information, etc.) to generate local feature information and global feature information of the current "state".
[0061] The feature fusion module is used to fuse the position mask and line mask to generate local features. It is also used to fuse the line mask and view mask to generate global features. The feature fusion module is also used to unify the dimensions of local and global features and fuse features in the channel dimension to generate comprehensive features.
[0062] Combine Figure 3 、 Figure 4 as well as Figure 7 It can be seen that in this embodiment, the feature fusion module includes a first fusion unit, a second fusion unit and a third fusion unit; the first fusion unit adopts a 1×1 convolution layer, and is used to perform feature fusion of the position mask and the line mask to generate local features. The second fusion unit adopts an enhanced layer structure that combines a residual network (Resnet) and a channel attention mechanism layer (SE) to perform feature fusion of the line mask and the view mask to generate global features. The second fusion unit includes four feature extraction layers composed of convolution layers and residual connections and four encoding layers. The channel attention mechanism layer is alternately distributed in the feature fusion module. The third fusion unit is used to unify the dimensions of local features and global features, and perform feature fusion in the channel dimension to generate comprehensive features.
[0063] like Figure 9 The ResNet+SE architecture of this embodiment is shown. The function of this fusion structure is similar to that of a traditional convolutional neural network, but the channel attention mechanism of SE can adaptively adjust the weight of each channel, allowing the model to pay more attention to features that have an important impact on layout decisions. Figure 10 The detailed functions of the SE part are described in . For the SE part, the recalibration of feature channels is mainly achieved through two stages. First, we need to Figure X Extract global information from . This is usually achieved through global average pooling (GAP), resulting in a one-dimensional vector U, expressed as:
[0064] U=F sq (X) = GAP(X)
[0065] Next, we use two fully connected layers (FC) to generate channel attention weights. This process can be expressed as:
[0066] V=F cx (U)=σ(W2(σ(W1(U))))
[0067] Where: W1 and W2 are the weight matrices of the two fully connected layers. σ is another activation function, usually Sigmoid, to ensure that the output value is in the range [0,1]. Finally, the generated channel attention weight V is combined with the original feature Figure X Multiply them together to get the weighted feature map Y:
[0068] Y=F scale (X,V)
[0069] Through the above process, the features of each channel will be adjusted by its corresponding attention weight, enhancing important feature channels and suppressing unimportant features, thereby reducing the huge action-space complexity that may be encountered during the layout process.
[0070] The policy network generates the probabilities of various layout optimization actions based on the embedding vectors of comprehensive features. The value network calculates the incentives and values corresponding to the updated states based on the embedding vectors of global features. The experience pool stores new states, actions, rewards, and values as experience, and updates the policy and value networks based on this stored experience.
[0071] In this embodiment, Figure 5As shown in Figure 1, the value network consists of a leading fully connected layer, Linear1, a global information layer obtained through a convolutional module, and three trailing fully connected layers, Linear2, Linear3, and Linear4. Each functional layer of the value network is connected via Reluctant Unified Unit (ReLU) functions. The leading fully connected layer, Linear1, has 768 input channels and 512 output channels. The three trailing fully connected layers all have 512 input channels and 1 output channel.
[0072] like Figure 7 As shown in Figure 1, the policy network includes two convolutional layers, a transformation layer, and a Softmax normalization layer. The method for the policy network to output the corresponding action probability X based on the embedded vector input of the comprehensive feature includes the following steps (such as Figure 11 (as shown): The embedding vector input of the comprehensive feature is divided into three paths, two of which are sliced and reconstructed at different scales to obtain coarse-grained features F1 and fine-grained features F2, and the other path is used to extract mask features to obtain a mask vector mask. The coarse-grained feature F1 is processed by a convolution layer with a coarse convolution to obtain the first feature H1. The fine-grained feature F2 is processed by a convolution layer with an attention module to obtain the second feature H2. The first feature H1 and the second feature H2 are first concatenated, and then merged by the transformation layer to transform them into a fusion feature matrix FS1 of a specified dimension. The mask vector mas is used to perform a mask operation on the fusion feature matrix FS1 to obtain the action feature matrix A1. The action feature matrix A1 is processed by a Softmax normalization layer to obtain the action probability X.
[0073] The experience pool, policy network, and value network constitute a reinforcement learning framework. Step S1 also initializes the state space, action space, reward function, state transition function, and reset condition of the reinforcement learning framework before training. The state space is the range of states that a component can reach. Generally, it is set to a discrete state. In this embodiment, the state space contains two states: Unplaced modules that have not been laid out and Placed modules that have been laid out. Among them, whether the module has been laid out is evaluated through state characteristics such as Width, Height, Pin Offset, and Pin to Net Connection.
[0074] The action space is the range of all possible actions that a component can perform. In this embodiment, the chip canvas is viewed as a grid and divided into N×N cells, thereby generating N×N possible actions, that is, randomly placing an unplaced module in any cell where no module is placed.
[0075] The reward function is the feedback signal received by the agent when it achieves its goal in the environment, which directly affects the learning process of the agent. In this embodiment, the half-circle line length HPWL is regarded as a reward because the line length is the main optimization target among different performance indicators. This is different from the existing technology, which weights the half-circle line length and congestion as rewards and adjusts the weight coefficient as an additional hyperparameter. Specifically, a dense reward is achieved by defining a local HPWL, which only calculates the currently placed pins. For example, the part of HPWL after step t-1 is defined as HPWL(t-1), and HPWL(t) is calculated after taking action. The reward for step t is r(t) = HPWL(t) - HPWL(t-1), which is the opposite of the increase in HPWL. The reward function includes:
[0076] R t =-a t WL(P,H)-b t C(PH)
[0077] HPWL=(max{x i}-min{x i})+(max{y i}-min{y i})
[0078] Among them, R t is the global reward, a t 、b t are the total wiring length and the super parameters of congestion information and congestion importance respectively, P is the macro module that has been placed, H is the connection information of the corresponding macro module, WL() is the length of the wires of all placed macro modules, and C() is the cost function. max{x i} is the maximum horizontal coordinate of all pins in the network, min{x i} is the minimum horizontal coordinate of all pins in the network, max{y i} is the maximum value of the vertical coordinate of all pins in the network, min{y i} is the minimum vertical coordinate of all pins in the network.
[0079] State transfer function execution: The policy network selects an action and executes it in the virtual environment, which returns a new state and reward. The new state, action, and reward are stored in the experience pool as experience, and the policy network and value network are updated based on the stored experience. The placement of all components is used as a reset condition, and the accumulated reward is reset to zero after the reset condition is met, that is:
[0080] state[i,0]>=placed_num_marco-1.
[0081] In step S2, multiple random netlists are first used as sample data to form an original dataset, and then the layout optimization model is trained using the original dataset. Specifically, a large number of random netlists containing multiple components are first obtained and used as sample data to form an original dataset. The original dataset is then divided into a training set and a test set to train and test the layout optimization model. The model parameters of the trained layout optimization model that meets the target are saved.
[0082] The data set mainly includes information about macro modules and standard cells, which is represented as linked list information. In this embodiment, the training data set is a random netlist consisting of several components, which is processed into a series of linked lists through a series of processing steps. This embodiment performs multiple rounds of iterative updates on the exploration model of the initial layout solution based on the training data set. The "reward" is used to characterize the quality of the current module training. When the conditions are met, the training is stopped and the PPO algorithm is used to constrain the network. The objective function used for the policy network is:
[0083] policy(θ)=E[min(r(θ)A,max(clip(r(θ),1-∈,1+∈)A))]
[0084] Wherein, policy(θ) represents the objective function, E represents the expected value, Represents the probability ratio of the new and old strategies; A=G t -V t , represents the advantage function estimate at time step t, G t is the action value function, V t is the state value function; clip(r(θ), 1-∈, 1+∈) represents the clipping function, which restricts r(θ) to the interval [1-∈, 1+∈]. By obtaining the target layout that meets the requirements, the network model is explored. The quality of model training is primarily evaluated through "reward" evaluation; a larger "reward" value indicates a better model performance.
[0085] In this embodiment, the above steps also use the original data set to test the layout optimization model. Before training, it is necessary to preset the number of test rounds Epoch, the score obtained by the action score, the half-circle line length HPWL, the custom loss cost, the time spent on each test round, etc. During the training phase, the policy network and the value network are updated at each time point; when the value network is updated, the gradient back-stepping of the global feature embedding is stopped. During the testing phase, the probability matrix is obtained from the policy network in each round, and the actions are sampled according to the probability matrix. When the sampling exceeds a threshold, the corresponding value is obtained from the line mask, and the half-circle field line length is calculated (estimated) before executing the action.
[0086] In step S3, the trained layout optimization model is used to process the original circuit diagram's graph file to obtain an optimized netlist file, and then an optimized initial circuit layout diagram is generated accordingly. Finally, overlapping portions of the layout are eliminated to obtain the final layout diagram. In this embodiment, the graph file of the original circuit diagram to be optimized is input into the trained layout optimization model, which outputs an optimized netlist file. Based on the netlist file, a corresponding optimized initial circuit layout diagram is generated. Finally, an algorithm guided by a mask mechanism is used to optimize overlapping portions of the layout to obtain the final layout diagram.
[0087] like Figure 8 As shown, fine-tuning the RL layout as a subsequent process of the model can optimize the layout to a certain extent and eliminate possible layout overlaps. Using this algorithm alone to lay out the chip's macromodules can require a considerable amount of time to obtain a final layout. However, using the layout obtained by our model as the initial layout input can significantly save time. Furthermore, for different data sets, we determine the grid in the current placement netlist based on the number and size of macromodules in the current data set, which can reduce spatial complexity to a certain extent. Figure 5 The optimization process of the algorithm is described, including Figure 5 There is a certain overlap between any three macro modules (A, B, C). The algorithm guided by the mask mechanism generates a line mask based on the initial macro position, and the line mask records the current macro v i The increment of HPWL after each candidate grid is selected, and the grid set with the smallest increment value is selected, and its increment value is expressed as f i If there are several grid sets Q, the algorithm guided by the mask mechanism selects the distance macro v from all grid sets Q i The algorithm, guided by a mask mechanism, identifies already placed areas as placed by binary conversion (1 indicates unplaceable, 0 indicates placement possible). It then iteratively swaps the positions of macromodules to eliminate layout overlaps. This iterative swapping of macromodules allows for effective layout adjustments, improves placement quality to a certain extent, achieves zero overlap, and minimizes HPWL. Among them, E(g i ) indicates that its corresponding rectangular covering grid g i The first equality holds because the hyperedges e enumerated on the left side (LHS) are j The degree of ∈E is equal to the number of grids it covers, i.e., w j ,h j .
[0088] In summary, compared with existing chip layout methods, the chip layout method based on reinforcement learning and backend optimization in this embodiment has the following beneficial effects:
[0089] 1. This chip layout method based on reinforcement learning and back-end optimization captures feature information in the circuit netlist file through position masks, line masks and view masks. By exploring randomization and continuously iteratively updating parameters, more low-level feature information can be retained through fusion, which helps the network capture complex patterns in chip layout at a deeper level, eliminates the interference caused by different specifications of various standard units in the netlist data, makes it easier to extract relevant features, and more effectively meets the network input. At the same time, it can solve some overlapping layouts that may exist in the layout, and to a certain extent improve the layout quality, reliability and practicality, and solve the technical problems of low efficiency and insufficient generalization of the existing chip layout methods, improve the chip layout efficiency, make up for the generalization of the solution, and explore more effective layout solutions with improved exploration efficiency.
[0090] 2. This chip layout method, based on reinforcement learning and back-end optimization, transforms circuit spatial layout into a spatial optimization problem for component placement within a canvas space, and iteratively optimizes the problem using reinforcement learning. An optimized residual network and channel attention mechanism layer architecture are introduced during the global feature extraction process. By integrating the residual network and channel attention mechanism layer architectures, the model's ability to understand and process complex placement information is enhanced. Furthermore, the channel attention mechanism adaptively adjusts the weight of each channel, allowing the model to focus more on features that have a significant impact on layout decisions.
[0091] 3. This chip layout method, based on reinforcement learning and back-end optimization, allows GPUs to enhance the model's ability to capture diverse information through parallel computing, improving model stability and accelerating training and inference. Compared to existing technologies, this method significantly improves key performance indicators, such as minimizing the half-circuit length (HPWL) of circuit layouts, and significantly accelerates model training and layout inference. A mask-guided algorithm is used to address overlapping layouts, ensuring the legitimacy of macromodule placement and improving layout quality to a certain extent.
[0092] 4. This chip layout method based on reinforcement learning and back-end optimization sets a series of reward functions during the layout process, increasing the "agent" (intelligent body) in the reinforcement learning process to explore unknown layouts, avoiding falling into local optimality due to repeatedly choosing the same action; at the same time, by maximizing rewards, it abandons layout plans that are obviously unpromising, thereby improving the exploration efficiency of layout plans.
[0093] Example 2
[0094] This embodiment provides a chip layout method based on reinforcement learning and back-end optimization. Based on Example 1, this chip layout method selects a solution for simulation testing and compares it with multiple solutions.
[0095] 1. Experimental environment and parameter settings
[0096] In this experiment, the network training environment was based on the following hardware and software: Ubuntu 24.04 LTS operating system, Python 3.9.20, PyTorch 2.4.1, CUDA 12.1, an Intel 13790F processor with 32GB of RAM, and an NVIDIA RTX 4070ti GPU with 12GB of RAM. The Adam optimizer was used for model training. The Adam optimizer combines adaptive learning rate and momentum to effectively adapt to gradient variations across different parameters. Initially, Adam adaptively adjusts the learning rate, using a larger learning rate early in training to accelerate convergence, and then gradually reducing it later to refine the search for the optimal solution. Furthermore, Adam integrates momentum by accumulating the squared gradients of previous updates, thereby adjusting the direction and magnitude of parameter updates. This mechanism helps the optimizer avoid local minima and increases the likelihood of finding the global optimum.
[0097] 2. Control Group and Dataset
[0098] This experiment used the RSWPlacer solution of the present invention as the experimental group. To accurately evaluate the performance of the solution of this embodiment, technicians also compared the solution of the present invention with several state-of-the-art circuit layout optimization methods. The control groups included DREAMPlace, DeepPR, and MaskPlace, Graph.
[0099] Table 1: Dataset overview
[0100]
[0101] Each method was evaluated using various public circuit designs and evaluated according to their respective experimental setups. The optimized circuits were selected from public datasets, including the widely used ISPD2005 benchmark suite, as described in Table 1 above, and the Ariane RISC-V CPU design.
[0102] 3. Experimental data and analysis:
[0103] The optimization effects of the solution of this embodiment and the control group method on various sample circuits are shown in the following table:
[0104] Table 2 Comparison of HPWL results of each macro circuit after different layout optimization schemes
[0105]
[0106] Analyzing the data in the above table, we can find that:
[0107] The lower the HPWL, the better the performance. The layout optimization performance of the RSWplacer in each macro circuit provided by this embodiment is better than other methods, and only in one case is it worse than DREAMPlace, which shows the superiority of the solution of this embodiment.
[0108] Example 3
[0109] This embodiment provides a chip layout system based on reinforcement learning and back-end optimization, which includes a model building subsystem, a training subsystem, and an optimization subsystem.
[0110] The model building subsystem is used to construct the layout optimization model. It includes a feature extraction module, a feature fusion module, an experience pool module, a policy network module, and a value network module. The feature extraction module is used to feature encode the node, position, and line network features of the macromodule in the netlist to generate corresponding position masks, line masks, and view masks. The feature fusion module is used to fuse the position mask and line mask to generate local features, and to fuse the line mask and view mask to generate global features. The feature fusion module is also used to feature fuse local features and global features to generate comprehensive features. The policy network module is used to generate the probabilities of various layout optimization actions that should be taken based on the embedding vector of the comprehensive features. The value network module is used to calculate the incentives and values corresponding to the updated states after each round of actions based on the embedding vector of the global features. The experience pool module is used to store new states, actions, rewards, and values as experience and to update the parameters of the policy network module and value network module based on the stored experience.
[0111] The training subsystem uses multiple random netlists as sample data to form an original dataset, which is then used to train the layout optimization model. The optimization subsystem uses the trained layout optimization model to process the original circuit diagram file to obtain an optimized netlist file, then generates the optimized initial circuit layout diagram, and finally eliminates layout overlaps to obtain the final layout diagram.
[0112] Example 4
[0113] This embodiment provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the chip layout method based on reinforcement learning and back-end optimization described in Example 1 or Example 2, thereby creating the desired chip layout optimization system and performing circuit layout optimization on an input netlist file of a large integrated circuit design.
[0114] When applied, the method of Example 1 or Example 2 can be implemented in the form of software, such as a standalone program that can be installed on a computer device, which can be a computer, a smartphone, a control system, or other IoT device. The method of Example 1 or Example 2 can also be designed as an embedded program that can be installed on a computer device, such as a single-chip microcomputer.
[0115] The computer device can take various forms, including embedded chips or modules, or general-purpose data processing equipment, such as smart terminals that can execute programs, tablet computers, laptop computers, desktop computers, rack servers, blade servers, tower servers or cabinet servers (including independent servers, or server clusters consisting of multiple servers), etc.
[0116] The computer device of this embodiment includes at least, but is not limited to, a memory and a processor that can be interconnected via a system bus. Memory (i.e., readable storage media) includes flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, a magnetic disk, an optical disk, and the like. In some embodiments, the memory may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device.
[0117] In other embodiments, the memory may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc., equipped with the computer device. Of course, the memory may also include both the internal storage unit of the computer device and its external storage device. In this embodiment, the memory is generally used to store the operating system and various application software installed on the computer device. In addition, the memory may also be used to temporarily store various types of data that have been output or are about to be output.
[0118] In some embodiments, the processor may be a central processing unit (CPU), a graphics processing unit (GPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run program code stored in the memory or process data.
[0119] Example 5
[0120] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the chip layout method based on reinforcement learning and backend optimization of embodiment 1 or embodiment 2 are implemented.
[0121] When the method of Example 1 or Example 2 is applied, it can be applied in the form of software, such as a program designed as a computer-readable storage medium that can run independently. The computer-readable storage medium can be a USB flash drive designed as a USB shield, and the USB flash drive is designed to start the program of the entire method through external triggering.
[0122] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A chip layout method based on reinforcement learning and backend optimization, characterized in that: It includes the following steps: S1: Build layout optimization model: S1.1: Feature encode the node, position, and line net features of the macromodule in the netlist to generate the corresponding position mask, line mask, and view mask; S1.2: fusing the position mask and the line mask to generate local features, and fusing the line mask and the view mask to generate global features; S1.3: feature fusion of the local features and the global features to generate a comprehensive feature; S1.4: Based on the embedding vectors of the comprehensive features and the global features, generate the probabilities of various layout optimization actions that should be taken, and calculate the incentives and their values corresponding to the updated states after each round of actions; S1.5: Update the parameters of step S1.4 based on the newly stored experience; S2: First, multiple random netlists are used as sample data to form an original data set, and then the layout optimization model is trained using the original data set; S3: First, the trained layout optimization model is used to process the original circuit diagram file to obtain an optimized netlist file, and then the optimized circuit initial layout diagram is generated accordingly. Finally, the layout overlap is eliminated to obtain the final layout diagram; The layout optimization model includes a feature extraction module, a feature fusion module, an experience pool, a policy network and a value network; wherein the feature extraction module is used to feature encode the node, position and line network features in the macro module to generate the position mask, line mask and view mask respectively; the feature fusion module is used to perform feature fusion on the position mask and the line mask to generate the local feature; the feature fusion module is also used to perform feature fusion on the line mask and the view mask to generate the global feature; the feature fusion module is also used to unify the dimensions of the local features and the global features, and perform feature fusion on the channel dimension to generate the comprehensive feature; the policy network is used to generate the probability of various layout optimization actions to be taken according to the embedding vector of the comprehensive feature; the value network is used to calculate the incentive and its value corresponding to the state updated according to each round of action according to the embedding vector of the global feature; the experience pool is used to store new states, actions, rewards and values as experience, and update the policy network and the value network according to the stored experience; The value network includes a front fully connected layer Linear1, a global information obtained through a convolution module, and three rear fully connected layers Linear2, Linear3, and Linear4; the functional layers of the value network are connected by ReLU functions; the front fully connected layer Linear1 has 768 input channels and 512 output channels; the three rear fully connected layers all have 512 input channels and 1 output channel; The strategy network includes two convolutional layers, a transformation layer and a Softmax normalization layer; the strategy network outputs the corresponding action probability according to the embedding vector of the comprehensive feature, which includes the following steps: dividing the embedding vector of the comprehensive feature into three paths, two of which are sliced and reconstructed at different scales to obtain coarse-grained features and fine-grained features, and the other path is used to extract mask features to obtain a mask vector; the coarse-grained features are processed by a convolution layer with a coarse convolution to obtain a first feature; the fine-grained features are processed by a convolution layer with an attention module to obtain a second feature; the first feature and the second feature are firstly spliced, and then merged by the transformation layer to transform them into a fusion feature matrix of a specified dimension; the fusion feature matrix is masked by using the mask vector to obtain an action feature matrix; the action feature matrix is processed by a Softmax normalization layer to obtain an action probability; The objective function of the policy network is: policy(θ)=E[min(r(θ)A,max(clip(r(θ),1-∈,1+∈)A))] Wherein, policy(θ) represents the objective function, E represents the expected value, Represents the probability ratio of the new and old strategies; A=G t -V t , represents the advantage function estimate at time step t, G t is the action value function, V t is the state value function; Clip(r(θ),1-∈,1+∈) represents the clipping function, which limits r(θ) to the interval [1-∈,1+∈].
2. The chip layout method based on reinforcement learning and backend optimization according to claim 1, characterized in that: The feature fusion module includes a first fusion unit, a second fusion unit and a third fusion unit; the first fusion unit adopts a 1×1 convolution layer and is used to perform feature fusion on the position mask and the line mask to generate the local feature; the second fusion unit adopts an enhancement layer structure combining a residual network and a channel attention mechanism layer to perform feature fusion on the line mask and the view mask to generate the global feature; the second fusion unit includes four feature extraction layers composed of convolution layers and residual connections and four encoding layers, and the channel attention mechanism layer is alternately distributed in the feature fusion module; the third fusion unit is used to unify the dimensions of the local features and the global features, and perform feature fusion on the channel dimension to generate the comprehensive feature.
3. The chip layout method based on reinforcement learning and backend optimization according to claim 1, characterized in that: The position mask is an N×N binary matrix, where cells where components can be placed are marked as 1 and cells where components cannot be placed are marked as 0; The line mask is an N×N continuous matrix, and the elements in the matrix represent the increase in the wire length of the macro module placed in each unit relative to the position in the previous round; The view mask is an N×N binary matrix, where cells where macro modules are placed are marked as 1, and cells where no macro modules are placed are marked as 0.
4. The chip layout method based on reinforcement learning and backend optimization according to claim 1, characterized in that: In step S2, a large number of random netlists containing multiple components are first obtained, and the random netlists are used as the sample data to form the original data set. The original data set is then divided into a training set and a test volume to train and test the layout optimization model, and the model parameters of the layout optimization model that meets the target after training are saved.
5. The chip layout method based on reinforcement learning and backend optimization according to claim 1, wherein: Step S3: inputting the diagram file of the original circuit diagram to be optimized into the trained layout optimization model, the layout optimization model outputting the optimized netlist file, and then generating the corresponding optimized initial layout diagram of the circuit based on the netlist file. Finally, optimizing the layout overlap by an algorithm guided by a mask mechanism to obtain the final layout diagram; wherein, the algorithm guided by the mask mechanism generates a wire mesh mask based on the initial macro position, and the wire mesh mask records the HPWL increment after the current macro is placed in each candidate grid, and selects the grid set with the smallest increment value; if there are several grid sets, the algorithm guided by the mask mechanism selects the grid closest to the macro from all grid sets; the algorithm guided by the mask mechanism marks the already placed area as already placed by binarization, and eliminates the layout overlap by iteratively exchanging the positions of the macro modules.
6. The chip layout method based on reinforcement learning and backend optimization according to claim 1, characterized in that: Step S2, further testing the layout optimization model using the original dataset; during the training phase, the policy network and the value network are updated at each time point; when updating the value network, gradient backpropagation of the global feature embedding is stopped; During the testing phase, a probability matrix is obtained from the policy network in each round, and actions are sampled according to the probability matrix; when the sampling exceeds a threshold, the corresponding value is obtained from the line mask, and the half-cycle field line length is calculated before executing the action.
7. The chip layout method based on reinforcement learning and backend optimization according to claim 1, characterized in that: The experience pool, the policy network, and the value network constitute a reinforcement learning framework; step S1 also initializes the reinforcement learning framework before training, setting the state space, action space, reward function, state transition function, and reset condition; the state space is the range of states that a component can reach, and the action space is the range of all actions that a component can perform; the reward function is the feedback signal received by the intelligent agent when it achieves a goal in the environment; The half-circle length HPWL is regarded as a reward, and the reward function includes: R t =-a t WL(P,H)-b t C(P.H) HPWL=(max{x i }-min{x i })+(max{y i }-min{y i }) Among them, R t is the global reward, a t 、b t are the total wiring length and the super parameters of congestion information and congestion importance, respectively. P is the number of macro modules that have been placed, H is the connection information of the corresponding macro module, WL() is the length of the wires of all placed macro modules, and C() is the cost function; max{x i } is the maximum horizontal coordinate of all pins in the network, min{x i } is the minimum horizontal coordinate of all pins in the network, max{y i } is the maximum value of the vertical coordinate of all pins in the network, min{y i } is the minimum vertical coordinate of all pins in the network; The state transfer function executes: selecting an action through the policy network and executing it in the virtual environment, and the virtual environment returns a new state and reward; storing the new state, action and reward in the experience pool as experience, and updating the policy network and the value network based on the stored experience; taking the completion of placement of all components as the reset condition, and clearing the accumulated reward after the reset condition is met.
8. A chip layout system based on reinforcement learning and backend optimization, characterized in that: It includes: A model building subsystem, which is used to build a layout optimization model; the model building subsystem includes a feature extraction module, a feature fusion module, an experience pool module, a strategy network module and a value network module; the feature extraction module is used to feature encode the nodes, positions and line network features of the macro modules in the netlist to generate corresponding position masks, line masks and view masks; the feature fusion module is used to fuse the position mask and the line mask to generate local features, and fuse the line mask and the view mask to generate global features; the feature fusion module is also used to feature fuse the local features and the global features to generate comprehensive features; the strategy network module is used to generate the probability of various layout optimization actions to be taken based on the embedding vector of the comprehensive feature; the value network module is used to calculate the incentives and their values corresponding to the states updated according to each round of actions based on the embedding vector of the global feature; the experience pool module is used to store new states, actions, rewards and values as experience, and update the parameters of the strategy network module and the value network module based on the stored experience; a training subsystem, configured to first use a plurality of random netlists as sample data to form an original data set, and then train the layout optimization model using the original data set; The optimization subsystem is used to first process the chart file of the original circuit diagram through the trained layout optimization model to obtain an optimized netlist file, then generate the optimized circuit initial layout diagram accordingly, and finally eliminate the layout overlapping parts to obtain the final layout diagram.
Citation Information
Patent Citations
Reinforcement learning macro-module layout method based on curiosity driving
CN118536458A
Perception model self-recommendation method and device based on hyper-parameter intelligent optimization
CN118917412A