Simulated IC layout wiring sequence optimization and wiring method based on reinforcement learning
By using reinforcement learning-based line sequence selection model and bidirectional wiring algorithm in simulated IC automatic wiring technology, the wiring sequence and results are automatically optimized, and the problems of high computational complexity and non-optimal wiring sequence in the existing technology are solved, achieving efficient and high-quality wiring effects.
Patent Information
- Application Number
- CN202510028139.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-08
AI Technical Summary
The existing analog IC automatic wiring technology has high computational complexity when dealing with complex circuit structures, making it difficult to find a feasible solution, and the artificially specified wiring sequence may be non-optimal, affecting circuit performance.
Using reinforcement learning-based method, by constructing a line sequence selection model and a bidirectional wiring algorithm, the wiring sequence is automatically learned, the wiring results are optimized, and reward settings are carried out from multiple aspects of line length, number of through holes and coupling noise to improve wiring efficiency and layout performance.
It significantly improves the rationality and effectiveness of the wiring sequence, improves the quality of wiring results, reduces line length, number of through holes and coupling noise, improves layout performance, and improves wiring efficiency.
Smart Images

Figure CN119990052A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic design automation, and in particular to an analog IC layout wiring sequence optimization and wiring method based on reinforcement learning. Background Art
[0002] At present, the field of integrated circuits (ICs) is developing rapidly. With the continuous advancement of science and technology, the market has higher and higher performance requirements for integrated circuits, and the complexity and scale of circuits are also increasing. In this context, the design assistance of electronic design automation (EDA) technology is particularly important. Digital signals have clear high and low level states and are easy to process and analyze. After years of development, digital EDA tools are powerful and reliable in logic synthesis, layout and routing, and their development has been relatively complete. However, analog ICs are very sensitive to noise, interference, etc. The difference between the two leads to the need to meet more performance constraints in the wiring process of analog IC layouts, which brings certain difficulties to the design of automatic wiring technology for analog ICs.
[0003] Today's automatic routing technology mainly uses analytical algorithms and heuristic algorithms. A key advantage of analytical routing is that it can use mathematical formulas to accurately determine the optimal routing path, thereby ensuring a certain level of routing accuracy and reliability. However, the high computational complexity of this method limits its application in complex circuit structures, and it may not be possible to find a feasible solution when the problem scale is large. In this case, heuristic routing becomes a more practical choice. Similarly, in the heuristic routing method, when the problem scale reaches a certain level, different routing orders will affect the overall routing results.
[0004] Artificial intelligence (AI) is a powerful technology that plays an increasingly important role in our daily lives. AI can automatically extract features and discover patterns through the analysis and learning of large amounts of data, thereby achieving efficient decision-making and prediction. Due to these outstanding features, machine learning has gradually begun to emerge and be applied in the field of IC physical design. AI technology can assist in optimizing chip layout, wiring and other aspects, improve design efficiency and performance, and bring new ideas and methods to IC physical design. Summary of the invention
[0005] The purpose of the present invention is to provide an analog IC layout wiring sequence optimization and wiring method based on reinforcement learning, a wire sequence selection method based on reinforcement learning, and a bidirectional Routing algorithms effectively improve the efficiency of analog IC routing and improve the performance results of the layout.
[0006] The technical solution adopted by the present invention is: In the first aspect of the present invention, there is provided a method for optimizing and routing an analog IC layout routing sequence based on reinforcement learning, the method comprising: Obtain wiring data, including wiring areas, obstacle areas, wire net sets, and the starting and ending points of each wire net; Gridding the wiring space and constructing a multi-channel image based on the wiring data; Using multi-channel images as states, we solve the analog IC net sorting problem based on a Markov decision process, including: Construct a line sequence selection model. At each action decision, the line sequence selection model takes the current multi-channel image as input and outputs a probability vector and a prediction value. The probability vector is a vector composed of the probabilities of each line network being selected. Each time the line sequence selection model makes an action decision, it selects the line network with the highest probability in the probability vector for line network sorting, and updates the multi-channel image at the same time, so as to complete the sorting of all line networks; the prediction value refers to the value of the current line network sorting result after selecting the line network with the highest probability for line network sorting; According to the sorting results of all the nets, combined with the routing area, obstacle area and the starting point and end point of each net, each net is routed in turn to obtain the routing result of the simulated IC layout; The performance index of the simulated IC layout wiring result is calculated as a reward, and the total loss function of the line sequence selection model is constructed in combination with the predicted value to optimize the line sequence selection model; wherein the total loss function includes the loss function of the strategy network and the loss function of the value network.
[0007] In some embodiments, gridding the wiring space and constructing a multi-channel image based on the wiring data includes: Grid the wiring space of the analog IC, and the number of grids in the length and width directions corresponds to the length and width of the multi-channel image; If the analog IC uses a double-layer wiring method, that is, different metal layers are used for wiring in the horizontal direction and the vertical direction and connected through through holes, then the number of channels of the multi-channel image is twice the number of wire nets plus two. Different metal layers of each wire net correspond to two channels respectively. The last two channels represent the wiring area and the obstacle area; If the analog IC uses a single-layer wiring method, the number of channels of the multi-channel image is the number of wire nets plus one. The metal layer of each wire net corresponds to a channel, and the last channel represents the wiring area and obstacle area.
[0008] In some embodiments, the wiring space is gridded and a multi-channel image is constructed based on the wiring data, further comprising: The grid cells are marked with a binary method, where the available routing grid area is marked as 0, and the unroutable grid area that has been routed or occupied by other obstacles is marked as 1; The grid area where the start and end points of each line net are located is marked as 0; if a grid is selected for line net sorting, the grid area where the start and end points of the grid are located is marked as 1, thereby updating the multi-channel image. Therefore, the multi-channel image contains information about the currently arranged line net and the remaining unarranged line nets.
[0009] In some embodiments, building a line sequence selection model includes: (1) The feature map of the multi-channel image is processed through the attention mechanism to adaptively adjust the weight of each channel, including: ① Feature map of multi-channel image x Perform global average pooling and global maximum pooling operations to generate two different channel descriptors:
[0010] in, represents the feature map of global average pooling, Represents the feature map of global maximum pooling, represents global average pooling, represents global maximum pooling; ②The feature map of global average pooling and the feature map of global maximum pooling are transmitted through a shared fully connected network; the network first reduces the number of channels to , and then restore the number of channels to C , and finally use the Sigmoid function to generate the channel attention weight:
[0011] in, C is the original number of channels of the feature map of the multi-channel image, r is the reduction ratio, FC represents a fully connected network, s is the Sigmoid activation function, is the average attention weight, is the maximum attention weight; ③Calculate attention weight A :
[0012] ④Add attention weight A Feature maps with multi-channel images x Multiply by channel to get the weighted feature map output:
[0013] in, is the weighted feature map; (2) The weighted feature map Feature maps with multi-channel images x Splicing, obtaining a spliced feature map; (3) Construct a neural network that takes the concatenated feature map as input and outputs a probability vector and a prediction value.
[0014] In some embodiments, constructing the line sequence selection model further includes: Process multi-channel images and scale them into images of specific length, width, and number of channels.
[0015] In some of these embodiments, the neural network is a residual neural network.
[0016] In some of these embodiments, for each net, a unidirectional Algorithmic or Bidirectional The algorithm routes the net and marks the routed grid area as 1 after the routing is completed; If the analog IC uses a double-layer wiring method, then unidirectional Algorithmic or Bidirectional When searching at the current node, the algorithm will determine whether the adjacent nodes in the six search directions of the current node are available routing grid areas. The six search directions include up, down, left, right, front, and back, and select the better moving direction through cost function calculation; When a path transitions between layers at the same planar location on different metal layers, that location is a via location.
[0017] In some of the embodiments, a performance index of a simulated IC layout and routing result is calculated as a reward, including: Calculate various performance indicators of analog IC layout routing results, including: (1) Line length length , that is, the total length of the wiring path; (2) Number of through holes , that is, the total number of through holes in the wiring path when double-layer wiring; (3) Coupling noise ; According to the performance indicators of the simulated IC layout and routing results, the relative rewards are calculated respectively:
[0018] in, Indicates the base value; Indicates specific variable values, corresponding to various performance indicators of the analog IC layout wiring results; Indicates the relative reward value of the corresponding performance indicator, including the relative reward value of the line length , Relative reward value of the number of through holes and coupling noise relative reward value ; The total reward value is:
[0019] in, is the total reward value, a、b、c is the weight coefficient; If the analog IC uses a single-layer wiring method, the total reward value does not include the relative reward value for the number of through holes.
[0020] In some embodiments, constructing a total loss function of a line order selection model includes: The discounted cumulative reward calculation method is used to allocate rewards to the states of each time step as follows:
[0021] In the formula, t is the time step, the value range is [0, N], N is the total number of all wire networks; γ is the discount factor; Rewards for simulating IC layout and routing results; Indicates that at time step t Discount accumulated rewards when Based on the discounted cumulative reward, the loss functions of the policy network and the value network are calculated; for the policy network, the gradient-based clipping policy loss function is used as follows:
[0022]
[0023]
[0024] in, i represents the weight parameters of the policy network, represents the loss function of the policy network, Represents the time step t is the ratio of the probability that the old model selects a certain network for network sorting in a certain state to the probability that the new model selects the network for network sorting in the same state; Represents the time step t The predictive value of Indicates the difference; is a hyperparameter, Indicates that Restricted to [ , ] range, Indicates the limit value; For the value network, the mean square error is used as the loss function, as shown below:
[0025] In the formula, Φ is the weight parameter of the value network, represents the loss function of the value network; Total loss function As shown below:
[0026] in, is a regular term used to prevent parameter overfitting. Represents the weight matrix of the model.
[0027] According to a second aspect of the present invention, an analog IC is provided, which is wired using the analog IC layout wiring sequence optimization and wiring method based on reinforcement learning as described in any one of the first aspects.
[0028] Compared with the prior art, the present invention has the following advantages and beneficial effects: (1) Optimize the wiring sequence. The analog IC line network sorting problem is modeled as a Markov process. Using a reinforcement learning framework with an attention mechanism, the model automatically learns key channel information, enhances feature expression, and suppresses irrelevant or redundant information interference, providing a more accurate basis for wiring decisions. This solves the problem of non-optimal wiring sequence specified by humans, thereby improving the rationality and effectiveness of the wiring sequence.
[0029] (2) Improve the quality of routing results. Rewards are set from three aspects: line length, number of vias, and coupling noise to optimize routing results. Reducing line length can reduce parasitic capacitance, inductance, resistance, and power consumption, improving electrical performance and energy efficiency; reducing the number of vias can reduce the negative impact of parasitic parameters on high-frequency performance, signal transmission quality, power consumption, stability, manufacturability, and reliability; reducing coupling noise can reduce the impact of electrical or magnetic interactions between adjacent wires. Through these optimizations, the quality of routing results is significantly improved, and layout performance is improved.
[0030] (3) Improve wiring efficiency. Use bidirectional The routing algorithm searches for paths from both the start and end points of a two-ended network. Under reasonable cost function settings, the search process is accelerated. Ideally, the routing time can be shortened to one-way The algorithm is half-finished, which greatly improves the routing efficiency. At the same time, a higher routing layer change cost is added to the heuristic function, which reduces the number of through holes and further optimizes the routing effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 An overall framework diagram of an analog IC layout wiring sequence optimization and wiring method based on reinforcement learning provided in an embodiment of the present application; Figure 2 A two-way method provided in the embodiment of the present application Flowchart of the algorithm. DETAILED DESCRIPTION
[0032] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Based on the embodiments provided in this application, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention.
[0033] Obviously, the drawings described below are only some examples or embodiments of the present application. For ordinary technicians in this field, the present application can also be applied to other similar scenarios based on these drawings without creative work. In addition, it can also be understood that although the efforts made in this development process may be complicated and lengthy, for ordinary technicians in this field related to the content disclosed in this application, some changes in design, manufacturing or production based on the technical content disclosed in this application are just conventional technical means, and should not be understood as insufficient content disclosed in this application.
[0034] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those of ordinary skill in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0035] Unless otherwise defined, the technical terms or scientific terms involved in this application should be understood by people with ordinary skills in the technical field to which this application belongs. The words "one", "a", "a", "the" and the like involved in this application do not indicate a quantitative limitation, and may represent the singular or plural. The terms "include", "comprise", "have" and any of their variations involved in this application are intended to cover non-exclusive inclusions; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units that are not listed, or may also include other steps or units inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there may be three relationships, for example, "A and / or B" can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.
[0036] At present, the automatic routing technology of analog IC is difficult. Analog IC is extremely sensitive to noise and interference, and the integrity and accuracy of its internal signals are easily disturbed by external factors. During the routing process, it is necessary not only to ensure the correct electrical connection between each device, but also to strictly control the noise level and electromagnetic interference during signal transmission. These special requirements make the automatic routing technology of analog IC face more challenges in the design and implementation process.
[0037] Commonly used automatic wiring algorithms are mainly divided into analytical algorithms and heuristic algorithms. Analytical wiring algorithms rely on complex mathematical models and formulas to calculate the optimal wiring path. When dealing with complex circuit structures, as the scale of the circuit increases, the factors that need to be considered increase, resulting in a sharp increase in computational complexity. In this case, the application of analytical algorithms in complex circuit structures will be limited, and feasible solutions may not be found when dealing with large-scale problems. When the heuristic wiring algorithm deals with wiring problems, when the scale of the problem reaches a certain level, the wiring order becomes a key factor affecting the overall wiring results. At present, the wiring order is mostly manually specified, but the manually specified method often lacks comprehensive optimization considerations. Different wiring orders may lead to completely different wiring results, and will also have a certain impact on the performance of the circuit.
[0038] According to the characteristics of analog IC wiring, this application provides an analog IC layout wiring sequence optimization and bidirectional The algorithmic wiring method improves the efficiency of analog IC wiring and improves the performance of the layout. The method includes the following steps: 1) Obtaining wiring data and performing corresponding processing; 2) Constructing a multi-channel image of the wire network information as the environment state; 3) The intelligent agent performs actions, explores and learns from experience in the environment, and records the corresponding actions and states; 4) If the training indicators are met, the existing data is input into the neural network to start the training process; 5) After completing a round of wire sequence arrangement, a bidirectional method is used according to the wiring order at this time. The algorithm is used for path planning. This application solves the problem that the artificially specified wiring sequence is not always optimal by modeling the analog IC wire network sorting problem as a Markov process and using a specific reinforcement learning framework, so as to realize the autonomous learning of the wiring sequence by the intelligent agent. In the reinforcement learning reward setting, the wiring results are optimized from multiple aspects such as wire length, number of through holes and coupling noise, and the relative value reward method is used to solve the problem of large differences in the order of magnitude of the reward indicators, so as to obtain a more reasonable basis for wiring decisions.
[0039] The present invention provides an analog IC layout wiring sequence optimization and bidirectional The wiring method of the algorithm includes the following steps: 1) Obtaining the wiring data and performing corresponding processing, wherein the wiring data includes the wiring area and obstacle area (such as the area occupied by the device) obtained from the layout results, the device connection relationship netlist and the corresponding pin coordinates, and splitting all the wire nets into double-ended wire nets; 2) Construct a multi-channel image and its feature map based on the line network information, process the feature map through the attention mechanism, and then input it into the neural network; 3) The agent (i.e., the line sequence selection model) performs actions, explores, and learns from experience in the environment, and records the corresponding actions and states; 4) If the training indicators are met, the existing data is input into the neural network to start the training process; 5) After completing a round of line sequence arrangement, use bidirectional wiring according to the wiring sequence at this time Algorithm for path planning.
[0040] From the above, this application models the analog IC line network sorting problem as a Markov decision process, and uses a reinforcement learning framework with an attention mechanism to process it. The reinforcement learning framework splits all circuit lines into two-pin nets. The feature map is processed by the attention mechanism, and the multi-channel image is processed by the residual neural network. The attention mechanism highlights important channel information by adaptively adjusting the weights of each channel in the feature map, thereby enhancing the representation ability of the model. In the reinforcement learning reward setting, the wiring results are optimized from three aspects: line length, number of through holes, and coupling noise, and rewards are set for the agent in the form of relative value rewards.
[0041] With the help of reinforcement learning, the intelligent agent can autonomously learn the wiring sequence, effectively solve the problem of non-optimal wiring sequence specified by humans, improve wiring efficiency and layout performance, and its reward setting comprehensively optimizes the wiring results by multiple factors, which is bidirectional. The algorithm accelerates the search process and provides new ideas for analog IC layout routing.
[0042] Specifically, Figure 1 As shown, the present application maps the wiring sequence selection problem to a Markov decision problem, and the action selection of the current wiring sequence depends only on the current state.
[0043] The present invention firstly performs uniform gridding on the layout space and constructs a multi-channel image of the wiring information. M Considering that analog ICs use double-layer wiring, that is, different metal layers are used for wiring in the horizontal and vertical directions and connected through vias, the number of channels in the multi-channel image will be set to twice the number of double-ended nets plus two, and the grid cells will be marked using a binary method. The bottom two layers represent the wiring area and the obstacle area, and the remaining layers contain the marking information of each net.
[0044] In the embodiment of the present application, the wiring adopts a grid wiring method. First, the coordinate information of the layout image is obtained to know the range of the wiring space, and then the wiring space is divided into uniform grids. Taking into account factors such as the space utilization of the wiring layer, circuit performance and circuit manufacturability, the analog IC adopts a double-layer wiring mode, that is, the horizontal and vertical directions of the wiring use different metal layers and are connected through through holes. When the grid unit is marked by a binary method, the available wiring area can be marked as 0, and the grid resources that cannot be wired because they have been wired or occupied by other obstacles are marked as 1. The bottom two layers of the image mark the pin information of all wire nets and the final wiring information after the obstacles are formed after the subsequent actual wiring. The remaining layers will be compared one by one according to the pin information of all the original wire nets to mark the positions of the starting and ending points. Whenever an action decision is made, the image layer group where the corresponding wire net is located will be marked.
[0045] It should be noted that the present application is also applicable to analog ICs with a single-layer wiring method. For analog ICs with a single-layer wiring method, each wire net corresponds to only one layer, that is, it only occupies one channel of a multi-channel image.
[0046] Set state S = {M, J, H} ,in J Represents the current set to be routed. H Represents a sequence of arranged wire nets. Action space A The size corresponds to the current state J After each decision, A The size of will be reduced by one. Define the currently selected network action as at , the wiring strategy is π : S → A . Build a line sequence selection model and make action decisions based on the line sequence selection model.
[0047] The line sequence selection model process is as follows: (1) After obtaining the wiring-related information, the constructed wiring multi-channel image is scaled to adapt to the input format of the neural network, and the feature map is obtained after scaling x , and sent to the attention mechanism for corresponding processing. After the corresponding processing, it is combined with the original feature map x After splicing, it is passed into the neural network.
[0048] (2) The neural network outputs a probability vector and an estimated value for each action decision. The probability vector refers to the probability that each line is selected to be added to the network. H The vector of probabilities in the , select the line network with the largest probability value as the current actual action a t , the estimated value is to choose this action a t Then the value will be generated according to the current line sequence.
[0049] (3) After selecting the result of a round of line sequence, use bidirectional The algorithm performs wiring and calculates the actual reward value based on the wiring results, which serves as reward feedback for reinforcement learning and is used to optimize the quality of wire sequence selection.
[0050] Among them, the attention mechanism can highlight important channel information by adaptively adjusting the weights of each channel in the feature map, thereby enhancing the representation ability of the model. In this work, the embodiment of the present application uses average pooling and maximum pooling to extract global channel features.
[0051] For the input feature map ,in C is the number of channels, H and W is the spatial dimension, then the output of the global average pool (GAP) The average value for each channel is calculated as: (1) In the formula, Represents the position in the feature map aisle c The value at .
[0052] Use the global max pool (GMP) method to calculate the maximum value of each channel and output , as shown below: (2) By inputting feature maps x Perform average pooling and maximum pooling operations to generate two different channel descriptors: (3) Both pooled features are passed through a shared fully connected network. This network first reduces the number of channels to (where r is the reduction ratio), and then restore the number of channels to C , and finally use the Sigmoid function to generate the channel attention weight: (4) Among them, FC represents the fully connected layer, and σ is the Sigmoid activation function, which ensures that the output is in the range of [0, 1].
[0053] Then add these two attention maps: (5) Attention Weight A With the original input feature map x Multiply channel-wise to produce a reweighted output: (6) In the overall reinforcement learning framework, the weighted output The attention mechanism is integrated into the feature processing of the model to enhance the model's recognition and utilization of important features of multi-channel images.
[0054] Since the length and width of the circuit layouts in different cases are inconsistent, the example of this application processes the multi-channel image, scales it into an image of uniform length, width, and height (i.e., number of channels), and then passes it into the residual neural network. Of course, other neural network structures can also be tried to replace the residual neural network to process multi-channel images.
[0055] The fusion of the attention mechanism and the neural network enables the model to automatically learn which channels are more important for the current task, thereby enhancing the feature expression of key channels and suppressing irrelevant or redundant information.
[0056] It should be noted that rewards can only be calculated based on the wiring results after at least one round of wire sequence arrangement is completed. Because if one round of wire network arrangement is not completed, it is impossible to perform wiring based on the wire network arrangement results, and the wiring results cannot be obtained.
[0057] For the reward setting in reinforcement learning, the present invention example evaluates the wiring results from the following perspectives: (1) Cable length: Excessively long metal wires increase parasitic capacitance and inductance, inhibiting signal transmission speed and integrity, and ultimately degrading electrical performance. In addition, longer wiring leads to higher resistance and power consumption, reducing energy efficiency. In the reward setting, the total length of the wiring result is directly calculated as a component of the reward.
[0058] To calculate the wire length, this application will take the total length Wire_length It is defined as the sum of the Manhattan distances between points in each routing path. and The path, Wire_length As shown in formula (7): (7) in, n Represents the total number of segments on the path. The z coordinate is ignored in this calculation since only the 2D distance between each segment is considered.
[0059] (2) Number of through holes: In analog ICs, vias will introduce parasitic capacitance, inductance, and resistance. A large number of vias will increase parasitic parameters, negatively affecting high-frequency performance, signal transmission quality, power consumption, and stability. Vias will also affect the manufacturability of the circuit. Considering the physical discontinuity between the vias and the surrounding materials, and the requirement for additional processing steps, the presence of vias will lead to increased manufacturing complexity and costs, and ultimately reduce the yield rate. In addition, vias are the cause of thermal stress concentration during circuit operation, thereby reducing the reliability of the circuit. In the example of this application, the number of times the metal wire layer changes in the wiring result is approximated to the number of vias, which is another bonus indicator.
[0060] This is illustrated by calculating the z-coordinate change between each turning point of the routing path as the number of vias. and ,when When , a through hole is added, as shown in equation (8): (8) in, is an indicator function, when (represents a layer change, i.e. a via should be added) is 1, otherwise it is 0.
[0061] (3) Coupling noise: In analog IC wiring, the smaller the spacing between metal wires, the more obvious the increase in coupling noise. Coupling noise usually comes from the electrical or magnetic interaction between adjacent wires, and the degree of coupling depends on the proximity of the wires. In order to reduce the coupling noise of the wiring result, the present invention example determines the penalty term by calculating the coupling noise of parallel lines that are too close. Formula (9) (parameter meaning) gives the correlation noise calculation method.
[0062] (9) in, and A function of the coupled noise between two edges, MD represents the maximum distance between two parallel line segments that have coupled noise effects. and It is a variable related to the grid line segment in the wiring diagram, which is used to determine whether the line segment is selected for wiring. The logical AND operation (&) is used to determine whether two specific line segments are selected for wiring at the same time.
[0063] Considering the inconsistency of the sizes of different reward components, this application uses a relative value method to calculate each reward category, as shown in formula (10). The standard value is calculated based on the heuristic network order and the corresponding wiring results in the current advanced wiring work.
[0064] (10) in, Indicates the reference value, x is a specific variable value. This calculation is used to determine the variable value x The change relative to the baseline value is converted into a reward value. , then the reward value is positive, indicating an improvement relative to the reference value. x Greater than , then the reward value is negative, indicating a deterioration relative to the reference value.
[0065] According to the above analysis, the total reward value is shown in formula (11).
[0066] (11) If the analog IC uses a single-layer wiring method, the total reward value does not include the relative reward value for the number of through holes.
[0067] Considering that reinforcement learning requires the agent to continuously explore the environment, the wiring process needs to be executed multiple times, which makes the time cost relatively high.
[0068] Based on this, the present application example designs a bidirectional Routing algorithms, such as Figure 2As shown, the search starts from the start and end of each double-ended net at the same time to improve the search efficiency. When the paths at both ends reach the same intersection, they will be backtracked and combined into a complete path. Ideally, the routing process takes one-way time. In addition, to reduce the path irregularity that may be caused by bidirectional search, a higher routing layer change cost is added to the heuristic function to minimize the number of vias.
[0069] Figure 2 It is bidirectional The flow of the routing algorithm. Bidirectional The algorithm first creates two sets of nodes, representing the starting point and the end point. Next, the algorithm checks whether the open set (i.e., the set of nodes to be processed) is empty. If the open set is empty, it means that no path can be found to connect the starting point and the end point, and the algorithm will declare the routing failed. If the open set is not empty, the algorithm will select the node with the smallest total cost (including actual cost and estimated cost) from each of the two sets. Then, the two nodes are expanded, their neighbor nodes are explored, and the total cost of these neighbor nodes is updated. At the same time, the algorithm will update the cost of the neighbor nodes and add them to the open set for subsequent processing. During the expansion process, the algorithm will constantly check whether a common node of the two sets, i.e., the confluence point, is found. If a confluence point is found, the algorithm will start from this node and search backwards to the starting point and the end point respectively to build a complete path. Once the path is built, the algorithm will output the routing result and end the run.
[0070] It should be noted that this application can also use one-way The algorithm is just more time-consuming than the bidirectional Of course, this application can also use other path planning algorithms. In addition, if the analog IC is a double-layer wiring method, then the one-way Algorithms and Bidirectional The routing algorithm needs to meet the design rules, that is, when searching for the current node, the node at the same position on the other metal layer will also be added to the open set, provided that the node at the same position on the other metal layer can be routed. If the routing path passes through the same position in two metal layers, then this position is the through-hole position.
[0071] Specific, one-way Algorithmic or Bidirectional When searching at the current node, the algorithm determines whether the adjacent nodes in the six search directions of the current node are available routing grid areas. The six search directions include up, down, left, right, front, and back. The algorithm selects the optimal moving direction through cost function calculation. When the path undergoes inter-layer transition at the same plane position on different metal layers, this position is the through-hole position.
[0072] It should be noted that when the current node is located at the upper metal layer, its search direction does not include the upper area or the upper area is directly set as a non-routable area; the same applies to the lower area.
[0073] Finally, this application example uses a proximal strategy optimization algorithm to update the neural network model parameters. The reward is calculated based on the wiring results obtained by the algorithm, and the discounted cumulative reward calculation method is used to allocate the reward to the state of each time step, as shown in formula (12).
[0074] (12) Where t is the time step, and its value range is [0, N]. This time step means that after the network is arranged in sequence, the bidirectional The reward value obtained by the wiring algorithm for data calculation based on the wiring results generated by this wiring sequence. γ is the discount factor.
[0075] Based on the discounted cumulative reward, the present invention example calculates the loss function of the policy network and the value network. For the policy network, a gradient clipping policy loss function is used, as shown in equations (13) to (15).
[0076] (13) (14) (15) in, i represents the weight parameters of the network, Yes ϵ is a hyperparameter used to limit the update range of the policy. (x, A, B) means restricting x to the range [A, B]. It represents the probability ratio of the old model output to the current model output. The probability ratio is explained as follows: Current Policy Network It means in the state Next select action The probability of is the probability of the old policy network in the same state and action, express and The ratio reflects the relative change between the new and old strategies.
[0077] For the value network, the mean square error is used as the loss function, as shown in formula (16).
[0078] (16) Where Φ is the weight parameter of the value network. The total loss function is shown in formula (17).
[0079] (17) in, It is a regularization term used to prevent parameter overfitting. represents the model parameters to be regularized (the weight matrix of the neural network), is the L2 norm of the parameter, that is, the square root of the sum of the squares of the weights, and γ is a hyperparameter of the regularization strength.
[0080] In addition, the example of this application adopts an adaptive learning rate method to dynamically adjust the step size according to the gradient information, so as to avoid the fixed learning rate from converging too slowly or oscillating during the training process, thereby improving the optimization effect of the model at different stages. The use of an adaptive learning rate can ensure stability and generalization ability during the training process, making the model more resistant to noise and outliers. The use of the RMSProp optimizer in the example of this application can minimize the loss and optimize the update of the weight parameters of the neural network, thereby maximizing the reward of the line sequence selection result, so that the model can adaptively learn to arrange the line network sequence with high quality, and further improve the wiring quality.
[0081] The reinforcement learning-based analog IC layout wiring order optimization and wiring method of the present application is based on a reinforcement learning framework for line sequence selection, which models the analog IC wire network sorting problem as a Markov process. The framework splits all circuit wire networks into two-pin nets and processes them using a reinforcement learning framework with an attention mechanism, thereby solving the problem that the current wiring order is not always optimal due to human regulations. In the reinforcement learning reward setting, the wiring results are optimized from three aspects: wire length, number of through holes, and coupling noise, and rewards are set for the intelligent agent in the form of relative value rewards to solve the problem of excessive differences in the order of magnitude of reward indicators. Based on bidirectional The routing algorithm designs a path planning method that starts searching from both the starting point and the end point of routing. By setting the cost function reasonably, it can speed up the routing and improve the routing efficiency while meeting the design rule check of the corresponding process and ensuring the routing effect.
[0082] An embodiment of the present application also provides an analog IC, which is wired using the analog IC layout wiring sequence optimization and wiring method based on reinforcement learning in the above method embodiment.
[0083] In summary, the present invention integrates the attention mechanism into the analog IC layout wiring order optimization and bidirectional The algorithm's routing method realizes the refinement and intelligence of the network information processing. The attention mechanism enables the model to automatically focus on key channel information, thereby effectively enhancing feature expression, suppressing the interference of irrelevant or redundant information, and providing a more accurate basis for subsequent routing decisions. Under the reinforcement learning framework, the routing results are comprehensively optimized from multiple aspects such as line length, number of vias, and coupling noise, and with the help of bidirectional The algorithm accelerates the search process, which not only solves the problem that the artificially specified wiring order is not always optimal, but also significantly improves the wiring efficiency and layout performance.
[0084] It should be pointed out that the technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above-mentioned embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification. In addition, according to the needs of implementation, the various steps / components described in this application can be split into more steps / components, and two or more steps / components or partial operations of steps / components can be combined into new steps / components to achieve the purpose of the present invention.
[0085] It is easy for a person skilled in the art to understand that the above-mentioned embodiments only express several implementation methods of the present application, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention patent. It should be pointed out that for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. An analog IC layout wiring sequence optimization and wiring method based on reinforcement learning, characterized in that: The method includes: Obtain wiring data, including wiring areas, obstacle areas, wire net sets, and the starting and ending points of each wire net; Gridding the wiring space and constructing a multi-channel image based on the wiring data; Using multi-channel images as states, we solve the analog IC net sorting problem based on a Markov decision process, including: Construct a line sequence selection model. At each action decision, the line sequence selection model takes the current multi-channel image as input and outputs a probability vector and a prediction value. The probability vector is a vector composed of the probabilities of each line network being selected. Each time the line sequence selection model makes an action decision, it selects the line network with the highest probability in the probability vector for line network sorting, and updates the multi-channel image at the same time, so as to complete the sorting of all line networks; the prediction value refers to the value of the current line network sorting result after selecting the line network with the highest probability for line network sorting; According to the sorting results of all the nets, combined with the routing area, obstacle area and the starting point and end point of each net, each net is routed in turn to obtain the routing result of the simulated IC layout; The performance index of the simulated IC layout wiring result is calculated as a reward, and the total loss function of the line sequence selection model is constructed in combination with the predicted value to optimize the line sequence selection model; wherein the total loss function includes the loss function of the strategy network and the loss function of the value network.
2. The method for optimizing and routing analog IC layout wiring sequence based on reinforcement learning according to claim 1, characterized in that: The wiring space is meshed and a multi-channel image is constructed from the wiring data, including: Grid the wiring space of the analog IC, and the number of grids in the length and width directions corresponds to the length and width of the multi-channel image; If the analog IC uses a double-layer wiring method, that is, different metal layers are used for wiring in the horizontal direction and the vertical direction and connected through through holes, then the number of channels of the multi-channel image is twice the number of wire nets plus two. Different metal layers of each wire net correspond to two channels respectively. The last two channels represent the wiring area and the obstacle area; If the analog IC uses a single-layer wiring method, the number of channels of the multi-channel image is the number of wire nets plus one. The metal layer of each wire net corresponds to a channel, and the last channel represents the wiring area and obstacle area.
3. The method for optimizing and routing analog IC layout wiring sequence based on reinforcement learning according to claim 2, characterized in that: Gridding the routing space and constructing a multi-channel image from the routing data, also includes: The grid cells are marked with a binary method, where the available routing grid area is marked as 0, and the unroutable grid area that has been routed or occupied by other obstacles is marked as 1; The grid area where the starting point and the end point of each line grid are located is marked as 0; if a grid is selected for line grid sorting, the grid area where the starting point and the end point of the grid are located is marked as 1, thereby updating the multi-channel image.
4. The method for optimizing and routing analog IC layout wiring sequence based on reinforcement learning according to claim 2, characterized in that: Build a line sequence selection model, including: (1) The feature map of the multi-channel image is processed through the attention mechanism to adaptively adjust the weight of each channel, including: ① Feature map of multi-channel image x Perform global average pooling and global maximum pooling operations to generate two different channel descriptors: in, represents the feature map of global average pooling, Represents the feature map of global maximum pooling, represents global average pooling, represents global maximum pooling; ②The feature map of global average pooling and the feature map of global maximum pooling are transmitted through a shared fully connected network; the network first reduces the number of channels to , and then restore the number of channels to C , and finally use the Sigmoid function to generate the channel attention weight: in, C is the original number of channels of the feature map of the multi-channel image, r is the reduction ratio, FC represents a fully connected network, σ is the Sigmoid activation function, is the average attention weight, is the maximum attention weight; ③Calculate attention weight A : ④Add attention weight A Feature maps with multi-channel images x Multiply by channel to get the weighted feature map output: in, is the weighted feature map; (2) The weighted feature map Feature maps with multi-channel images x Splicing, obtaining a spliced feature map; (3) Construct a neural network that takes the concatenated feature map as input and outputs a probability vector and a prediction value.
5. The method for optimizing analog IC layout wiring sequence and wiring based on reinforcement learning according to claim 4, characterized in that: Constructing a line sequence selection model also includes: Process multi-channel images and scale them into images with a specific length, width, and number of channels.
6. The analog IC layout wiring sequence optimization and wiring method based on reinforcement learning according to claim 4 or 5, characterized in that: The neural network is a residual neural network.
7. The analog IC layout wiring sequence optimization and wiring method based on reinforcement learning according to claim 3, characterized in that: For each net, combine the routing area, obstacle area, and its start and end points, using a one-way Algorithmic or Bidirectional The algorithm routes the net and marks the routed grid area as 1 after the routing is completed; If the analog IC uses a double-layer wiring method, then unidirectional Algorithmic or Bidirectional When searching at the current node, the algorithm will determine whether the adjacent nodes in the six search directions of the current node are available routing grid areas. The six search directions include up, down, left, right, front, and back, and select the better moving direction through cost function calculation; When a path transitions between layers at the same planar location on different metal layers, that location is a via location.
8. The method for optimizing and routing analog IC layout wiring sequence based on reinforcement learning according to claim 2, characterized in that: Calculate performance metrics of analog IC layout results as a reward, including: Calculate various performance indicators of analog IC layout routing results, including: (1) Line length length , that is, the total length of the wiring path; (2) Number of through holes , that is, the total number of through holes in the wiring path when double-layer wiring; (3) Coupling noise ; According to the performance indicators of the simulated IC layout and routing results, the relative rewards are calculated respectively: in, Indicates the base value; Indicates specific variable values, corresponding to various performance indicators of the analog IC layout wiring results; Indicates the relative reward value of the corresponding performance indicator, including the relative reward value of the line length , Relative reward value of the number of through holes and coupling noise relative reward value ; The total reward value is: in, is the total reward value, α, β, γ is the weight coefficient; If the analog IC uses a single-layer wiring method, the total reward value does not include the relative reward value for the number of through holes.
9. The method for optimizing and routing analog IC layout wiring sequence based on reinforcement learning according to claim 1, characterized in that: Construct the total loss function of the line order selection model, including: The discounted cumulative reward calculation method is used to allocate rewards to the states of each time step as follows: In the formula, t is the time step, the value range is [0, N], N is the total number of all wire networks; γ is the discount factor; Rewards for simulating IC layout and routing results; Indicates that at time step t Discount accumulated rewards when Based on the discounted cumulative reward, the loss functions of the policy network and the value network are calculated; for the policy network, the gradient-based clipping policy loss function is used as follows: in, θ represents the weight parameters of the policy network, represents the loss function of the policy network, Represents the time step t is the ratio of the probability that the old model selects a certain network for network sorting in a certain state to the probability that the new model selects the network for network sorting in the same state; Represents the time step t The predictive value of Indicates the difference; is a hyperparameter, Indicates that Restricted to [ , ] range, Indicates the limit value; For the value network, the mean square error is used as the loss function, as shown below: In the formula, Φ is the weight parameter of the value network, represents the loss function of the value network; Total loss function As shown below: in, is a regular term used to prevent parameter overfitting. Represents the weight matrix of the model.
10. An analog IC, characterized in that: The analog IC is wired using the analog IC layout wiring sequence optimization and wiring method based on reinforcement learning as described in any one of claims 1 to 9.
Citation Information
Patent Citations
PCB automatic wiring method and system based on line sequence simulation
CN113569523A
System and method for optimizing chip layout based on deep reinforcement learning
CN114154412A
Layout element automatic layout method and device based on mixed strategy reinforcement learning
CN117408215A
Global path planning method and device for an unmanned vehicle
US20220196414A1
Cited By
PCB intelligent design method and system
CN122133595A
PCB intelligent design method and system
CN122133595B