A global routing method based on route generation and sequential selection for reinforcement learning
By introducing multi-agent reinforcement learning methods in global routing, using imitation learning and reverse thread removal technology, the problems of global routing algorithm performance bottlenecks and high computing overhead in the existing technology are solved, and more efficient routing performance and significant reduction in computing overhead are achieved.
Patent Information
- Application Number
- CN202310378097.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-07
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-04-07
AI Technical Summary
The existing global routing algorithms have performance bottlenecks and high computing overhead, which is difficult to effectively solve the problems of traditional algorithms, and can only support the wiring of extremely small-scale global routing units or single networks.
A multi-agent reinforcement learning global routing method based on line generation and sequence selection is proposed. By constructing a reinforcement learning model including the first agent and the second agent, the convergence of the agent is generated by a simulated learning acceleration line, and the convergence of the agent is selected by reverse-removing the thread order.
It significantly improves the performance of the reinforcement learning cabling model, reduces computing overhead, shortens bus length, and reduces the total time overhead by about 5 times.
Smart Images

Figure CN116402012B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automated chip design, and more specifically, to a reinforcement learning global routing method based on route generation and sequential selection. Background Art
[0002] In the field of electronic design automation, the global routing task is a key task. The challenge lies in the need to connect a large number of circuit components (usually multiple pins) without violating printed circuit or integrated circuit rules. In addition, for the optimization of subsequent PPA objectives (performance, power consumption, area), the global routing task needs to control the routing bus length as short as possible and have less capacity overflow.
[0003] Currently, global routing algorithms mainly adopt traditional heuristic algorithms, and there are bottlenecks in both their performance and time efficiency. In recent years, routing algorithms based on artificial intelligence have been successively proposed, which are expected to break through the bottleneck of global routing algorithms. They can be roughly divided into generative routing models and reinforcement learning routing models. For example, the generative routing model (such as PRNet) trains a CGAN (Conditional Generative Adversarial Network) model. The generator learns the routing results, and the discriminator judges the authenticity and connectivity of the generated routing. The generator and the discriminator are alternately optimized and improved. Finally, the routing results of different networks are sequentially generated by the trained generator. The existing reinforcement learning routing models regard the global routing cell (GRC) as the state space, the routing position as the action space, and the combination of wirelength (WL) and capacity overflow as the reward of reinforcement learning, so as to optimize the reinforcement learning routing effect.
[0004] For example, the solution proposed by Liao H et al. (Liao H, Zhang W, Dong X, et al. A deep reinforcement learning approach for global routing [J]. Journal of Mechanical Design, 2020, 142(6)) mainly includes: transforming the multi-pin problem into multiple two-pin problems; using the Deep Q-Network (DQN) to optimize the solution strategy for the two-pin problem; and combining the results of multiple two-pin problems into a multi-pin problem. This solution has a high time complexity.
[0005] After analysis, it is found that the existing generative routing models still rely on the training results of traditional routing tools, and to some extent, they do not solve the problems of traditional algorithms. In addition, although the existing reinforcement learning routing models have made breakthroughs in performance, due to their very large training overhead, they currently only support very small-scale global routing units, or only solve the routing of a single network. Summary of the invention
[0006] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a reinforcement learning global routing method based on line generation and sequence selection. The method comprises the following steps:
[0007] Divide the target chip into multiple grids to form a grid graph, where the grid graph includes multiple nodes, and each node represents a unit;
[0008] Determine the path connecting the nodes in the grid graph by using the constructed reinforcement learning wiring model to achieve global wiring of the target chip;
[0009] The reinforcement learning model includes a first agent and a second agent. The first agent completes the preliminary wiring results of each network for the grid diagram. The second agent selects a network from the preliminary wiring results generated by the first agent to dismantle, and then uses the first agent to regenerate the wiring results of the network and calculate the updated reward value. The second agent optimizes the corresponding strategy network according to the updated reward value.
[0010] Compared with the prior art, the advantage of the present invention is that it proposes a reinforcement learning global routing method for multiple agents based on line generation and sequence selection, accelerates the convergence of the line generation agent through simulation learning, and accelerates the convergence of the sequence selection agent through reverse wiring, thereby improving the performance of the reinforcement learning routing model and significantly reducing the computational overhead.
[0011] Further features and advantages of the present invention will become apparent from the following detailed description of exemplary embodiments of the present invention with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.
[0013] Figure 1 is a flow chart of a reinforcement learning global routing method based on line generation and sequence selection according to one embodiment of the present invention;
[0014] Figure 2 is a global routing schematic diagram according to an embodiment of the present invention;
[0015] Figure 3 It is a schematic diagram of the overall process of the global routing method based on route generation and sequential selection according to an embodiment of the present invention;
[0016] Figure 4 It is a schematic diagram of the routing process of a single network according to an embodiment of the present invention. Detailed implementation manners
[0017] Now, various exemplary embodiments of the present invention will be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present invention.
[0018] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended as a limitation on the present invention, its application, or its use.
[0019] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be regarded as part of the specification.
[0020] In all the examples shown and discussed herein, any specific values should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.
[0021] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.
[0022] See Figure 1 As shown, the provided global routing method based on route generation and sequential selection includes the following steps:
[0023] Step S110: Divide the target chip into multiple grids to form a grid graph, and design a global routing task based on the grid graph.
[0024] As Figure 2 shown, first divide the chip into multiple tiles to form a grid graph GCell, and each node represents a tile. Figure 2 In the left figure, the three dark nodes represent the tiles where the pins required to be connected by a certain net are located. The task of global routing is to find a path connecting these tiles, as Figure 2 shown by the route in the right figure.
[0025] Step S120: Construct a reinforcement learning routing model including two agents, and obtain a global routing result by jointly optimizing the two agents.
[0026] Combine Figure 3 As shown, the constructed reinforcement learning wiring model adopts a multi-agent architecture, which are respectively called the agent for line generation and the agent for sequential selection.
[0027] In the agent for line generation, the problem is regarded as a Rectilinear Steiner Tree (RST) problem. For example, a number of nodes are interconnected with the shortest line length. Specifically, a fast heuristic Steiner Tree generation algorithm with obstacles is designed. The agent for line generation quickly approximates this heuristic wiring algorithm through imitation learning, and then WL and Overflow are used as optimization metrics (or rewards) and fed back to the agent for line generation. The policy network corresponding to the agent for line generation is updated through Proximal Policy Optimization (PPO). The agent for line generation first completes the preliminary wiring of all nets.
[0028] In the sequential selection agent, for the preliminary wiring of all nets completed by the agent for line generation, one net is selected for reverse removal and rewiring. The updated WL and Overflow are used as optimization metrics and fed back to the sequential selection agent, and the policy network corresponding to the sequential selection agent is updated through PPO.
[0029] Specifically, the reinforcement learning wiring method includes the following steps:
[0030] Step S1, for the agent for line generation, design a Steiner Tree generation algorithm with obstacles as a heuristic algorithm.
[0031] The Steiner Tree problem can be understood as a tree that connects the points in a specified set of points. This algorithm does not pursue finding the minimum spanning tree, but only pursues a relatively small spanning tree with a relatively fast generation speed. Therefore, using the Steiner Tree algorithm as a heuristic algorithm, the agent for line generation can obtain a better policy in a shorter time through imitation learning.
[0032] For example, given a source / target tiles set T, randomly select a target / source tile v, and calculate the score of its neighbor node u. The calculation formula is:
[0033] score u = cap u / min node∈T (|x u - x node | + |y u - ynode |) (1)
[0034] Among them, cap u represents the remaining routing capacity of node u, represents the minimum value of A obtained after traversing all nodes in set T, x u and x node respectively represent the abscissas of node u and node, y u and y node represent the ordinates of node u and node. The abscissa and ordinate here refer to the horizontal and vertical positions of the node in the GCell.
[0035] Find the node u' with the highest score, u' = argmax u (score u ). After that, connect the corresponding nodes as part of the routing, as Figure 4 shown.
[0036] It should be understood that other heuristic algorithms can also be used, such as the ant colony algorithm or neural network, etc.
[0037] Step S2: Use the agent for route generation to generate the routing results of each network.
[0038] For example, for the agent for route generation, its corresponding Reward (reward) includes the total wire length WL and Overflow, the state space is the remaining capacity Capacity of all current edges, source / target tile nodes and the routed routes, and the action space is to select the next edge to be connected. In this agent, first use imitation learning to quickly approximate the generation structure of heuristic routing, and then use the PPO policy to optimize this agent.
[0039] Step S3: Combine the agent for sequential selection and obtain the global routing result through joint optimization.
[0040] For example, for the agent for sequential selection, its corresponding Reward includes the total wire length WL and Overflow, the state space is the remaining capacity Capacity of all current edges, and the action space is to select the next net to be removed. After removing a certain net through this agent, then use the route generation agent in Step S2 to regenerate the routing result of this net, obtain the new WL and Overflow, and feedback them to the agent for sequential selection. This agent can also be optimized through the PPO policy.
[0041] It should be understood that, without departing from the spirit and scope of the present invention, those skilled in the art can make appropriate changes or modifications to the above embodiments. For example, the PPO algorithm can also be replaced by the A3C algorithm, etc. Also, the reward values or optimization objectives of each agent can further include other performance indicators.
[0042] To further verify the effectiveness of the present invention, a comparison was made using the benchmark generated by the prior art (Liao H, Zhang W, Dong X, et al. A deep reinforcement learning approach for global routing[J]. Journal of Mechanical Design, 2020, 142(6)). The results show that the present invention shortens the bus length, and the total time overhead is reduced by approximately 5 times.
[0043] In summary, compared with the prior art, the present invention has the following advantages:
[0044] 1) A fast heuristic Steiner Tree generation algorithm with obstacles is designed. This algorithm does not seek to find the minimum spanning tree, but only aims for a relatively small spanning tree with a fast generation speed, which is much smaller in terms of time overhead than the A* used in the best technology. In addition, the prior art transforms the multi-pin problem into multiple two-pin problems, resulting in a high time complexity of the problem. The present invention directly implements the Steiner Tree generation algorithm without the need to divide the problem.
[0045] 2) The present invention accelerates the convergence of the wire generation agent by imitating the above heuristic algorithm through imitation learning.
[0046] 3) In the sequential selection agent of the present invention, the method of reverse wire removal is adopted to accelerate the convergence of the sequential selection agent, and the wire generation agent and the sequential selection agent are combined to achieve global routing through joint optimization.
[0047] The present invention can be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.
[0048] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0049] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded to an external computer or an external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0050] The computer program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, Python, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present invention.
[0051] Aspects of the present invention are described herein with reference to the flowchart and / or block diagram of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer - readable program instructions.
[0052] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions comprises a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0053] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0054] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions. As will be apparent to those of ordinary skill in the art, implementation by hardware, implementation by software, and implementation by a combination of software and hardware are equivalent.
[0055] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of technology in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.
Claims
1. A global routing method based on route generation and sequential selection in reinforcement learning, comprising the following steps: Dividing a target chip into multiple grids to form a grid graph, which contains multiple nodes, and each node represents a cell; Using the constructed reinforcement learning routing model to determine the paths connecting the nodes in the grid graph to achieve global routing of the target chip; Among them, the reinforcement learning model includes a first agent and a second agent. The first agent completes the preliminary routing results of each network for the grid graph. The second agent selects a network to be removed from the preliminary routing results generated by the first agent, and then uses the first agent to regenerate the routing results of this network and calculate the updated reward value. The second agent optimizes the corresponding policy network according to the updated reward value; Among them, for the first agent, the corresponding reward includes the total wire length and capacity overflow. The state space is the remaining capacity of all current edges, the source node, the target node, and the routed routes. The action space is to select the next edge to be connected; Among them, for the second agent, the corresponding reward includes the total wire length and capacity overflow. The state space is the remaining capacity of all current edges, and the action space is to select the next network to be removed.
2. The method according to claim 1, characterized in that The first agent completes the preliminary routing results of each network by imitating and learning the partial routing results generated based on the Steiner tree.
3. The method according to claim 2, characterized in that The partial routing results generated based on the Steiner tree are obtained according to the following steps: Given a set of original nodes and target nodes , randomly select a target node or a source node , calculate its neighbor nodes 's score, expressed as: Find the node with the highest score and connect the corresponding nodes as the partial routing results; Among them, represents the score of node . represents the remaining routing capacity of node . represents the minimum value obtained after traversing all nodes in the set . . and respectively represent the abscissa of node and node . and respectively represent the ordinate of node and node .
4. The method according to claim 1, characterized in that The first agent and the second agent use proximal policy optimization to update the corresponding policy networks.
5. The method according to claim 1, characterized in that When the second agent selects a network to be removed from the routing results generated by the first agent, the reverse wire removal method is adopted.
6. A computer-readable storage medium, on which a computer program is stored, wherein When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
7. A computer device, including a memory and a processor, and a computer program capable of running on the processor is stored on the memory, characterized in that When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Multi-agent fault tolerance consistency method and system based on reinforcement learning
CN113919495A
System and method for optimizing chip layout based on deep reinforcement learning
CN114154412A