An end-to-end reinforcement learning hybrid scale layout method based on post-processing
By introducing a post-processing module into the end-to-end reinforcement learning layout model, wefting post-processing is performed to target component overlapping problems in hybrid scale layout scenarios, and the problem of serious overlapping layout results in the existing technology is solved, achieving better layout results and performance improvements.
Patent Information
- Application Number
- CN202310404878.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-17
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-04-17
AI Technical Summary
The existing end-to-end reinforcement learning layout model fails to effectively consider the actual size of macro components in hybrid scale layout scenarios, resulting in serious component overlap problems in layout results, affecting the performance of the chip.
A hybrid scale layout method based on post-processing is designed. By introducing a post-processing module into the layout model, weft post-processing is performed to optimize the overlapping area and component regularity of the layout results.
By introducing a post-processing module, the overlap area of layout results is significantly reduced, component regularity is improved, and the chip performance indicators are optimized.
Smart Images

Figure CN116415541B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automated chip design, and more specifically, to an end-to-end reinforcement learning hybrid-scale layout method based on post-processing. Background Art
[0002] The goal of chip layout is to place macro elements and standard elements on a chip board, requiring optimization of PPA (performance, power consumption, area) targets while satisfying constraints such as density and congestion. Traditional algorithms for solving the layout problem usually include heuristic algorithms, genetic algorithms, Monte Carlo sampling, or force-directed optimization algorithms. However, these algorithms have bottlenecks in terms of time efficiency and design complexity for specific layout problems, and cannot effectively utilize the rapidly developing hardware computing power in the field of artificial intelligence.
[0003] In recent years, intelligent layout algorithms based on reinforcement learning have been gradually proposed, achieving good fitting effects by modeling the state space through convolutional neural networks and graph neural networks. For example, the input of the end-to-end reinforcement learning layout model includes the sizes of each element and the connection relationships between elements (i.e., the netlist graph), makes sequential decisions for each element, obtains global embedding features and node embedding features with different coarse and fine granularities respectively through the policy network module (convolutional neural network and graph neural network), fuses the feature vectors obtained by the two networks, and finally predicts the probability distribution of the possible placement positions of elements at the current moment. Current end-to-end reinforcement learning layout models include Google's and DeepPlace, etc.
[0004] Upon analysis, existing end-to-end reinforcement learning layout models all default that each macro element occupies one grid unit. However, in the context of hybrid-scale layout, there are often large size differences between macro elements, that is, the actual size information of macro elements is not considered during the layout process, resulting in relatively serious element overlap problems in the layout results, affecting indicators such as actual wire length, and ultimately affecting the performance of the chip. Summary of the Invention
[0005] The object of the present invention is to overcome the defects of the above-mentioned prior art and provide an end-to-end reinforcement learning hybrid-scale layout method based on post-processing. The method includes the following steps:
[0006] For the layout problem of target elements on a chip, construct an end-to-end reinforcement learning layout model, where the reinforcement learning layout model includes a policy network and a post-processing module, and the target elements include macro elements and standard elements;
[0007] Use the policy network to achieve the overall layout of all macro elements;
[0008] The post - processing module aims to reduce the overlapping area between components, performs edge - sticking post - processing on all target components together, and obtains the layout results of all target components;
[0009] Calculate the reward value corresponding to the layout result using the set reward function, and update the policy network periodically until the set optimization goal is met.
[0010] Compared with the prior art, the advantages of the present invention are as follows. For the mixed - scale layout scenario with large differences in macro - component sizes, an end - to - end reinforcement learning layout model including a post - processing module is designed. After the overall layout of the agent is completed, post - processing is performed on all components together. Thus, by introducing the prior knowledge of placing components as close to the edge as possible, while reducing the overlapping area of the layout result, indicators such as component regularity are optimized.
[0011] Through the following detailed description of the exemplary embodiments of the present invention with reference to the accompanying drawings, other features and advantages of the present invention will become clear. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings incorporated in and forming a part of this specification illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.
[0013] Figure 1 is a flowchart of an end - to - end reinforcement learning - based mixed - scale layout method with post - processing according to an embodiment of the present invention;
[0014] Figure 2 is an overall framework diagram of an end - to - end reinforcement learning - based mixed - scale layout method with post - processing according to an embodiment of the present invention;
[0015] Figure 3 is an overall simulation flowchart of an end - to - end reinforcement learning - based mixed - scale layout method with post - processing according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] Now, various exemplary embodiments of the present invention will be described in detail with reference to the accompanying drawings. It should be noted that: Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0017] The following description of at least one exemplary embodiment is merely illustrative in nature and in no way serves as a limitation to the present invention, its application, or its use.
[0018] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be regarded as part of the specification.
[0019] In all the examples shown and discussed here, any specific values should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.
[0020] It should be noted that like reference numerals and letters refer to like items in the following figures, and thus, once an item is defined in one figure, it will not be discussed further in subsequent figures.
[0021] See Figure 1 and 2 As shown, the provided post - processing - based end - to - end reinforcement learning hybrid - scale layout method includes the following steps:
[0022] Step S110, for the layout problems of macro - elements and standard elements, construct an end - to - end reinforcement learning layout model, and this reinforcement learning layout model includes a post - processing module.
[0023] In the present invention, a reinforcement learning method is adopted to place macro - elements (such as SRAM) and standard elements (such as logic gates, including NAND, NOR) on the chip board, so as to optimize the power, performance, area (PPA), etc. of the chip, while complying with restrictions such as layout density and routing congestion.
[0024] The designed end - to - end reinforcement learning layout model, in addition to the policy network, also additionally includes a post - processing module. Generally speaking, the policy network places macro - elements in sequence, and then combines with the post - processing module to generate the placement of standard elements, and the positions of macro - elements will also be further adjusted and optimized during the post - processing.
[0025] In the reinforcement learning method provided by the present invention, the layout area on the chip is used as the environment. The state is each possible partial placement of the netlist on the chip canvas, that is, the layout situation of macro - elements in the layout area is used as the state. For a given macro - element to be placed currently, the available actions are the set of all positions in the discrete canvas space (grid cells) where the macro - element can be placed, and this position set does not violate any hard restrictions on density or blockage, etc. The reward refers to the reward for taking an action in a certain state.
[0026] Step S120, for all the macro - elements to be placed, use the policy network to achieve the overall layout of the macro - elements.
[0027] The learning objective of the hybrid - scale macro - element layout strategy is to learn a policy π, so that it can make layout decisions according to the given netlist graph and component size and other information, that is, determine the placement positions of macro - elements in sequence. For example, the policy network includes a convolutional neural network to obtain the global feature vector of the layout image, and a graph convolutional network to process the graph structure extracted from the netlist graph to obtain the local embedding feature of the netlist graph. Then, the two feature vectors are concatenated to complete the macro - element layout.
[0028] In one embodiment, for the layout of macro elements, the reward function adopted is the negative weighted sum of wire length, congestion, and density. The weights can be used to explore the trade - offs between various metrics.
[0029] Step S130, aiming to reduce the overlapping area between elements, perform edge - fitting post - processing on all elements together to obtain the layout result of all elements.
[0030] To alleviate the problem of element overlap in the intermediate layout result and improve the element regularity, a post - processing algorithm is specifically proposed to incorporate the prior knowledge of placing elements as close to the edge as possible in traditional methods into the reinforcement learning layout framework, that is, introduce a post - processing module to participate in the entire end - to - end learning process, as Figure 2 shown. When the overall layout of macro elements is completed, combine the post - processing module to perform edge - fitting post - processing on all elements together to optimize the overlapping area of the layout result.
[0031] Step S140, estimate the wire length and congestion metrics based on the layout result obtained from the post - processing, and feedback them to the agent, and periodically update the policy network through the proximal policy optimization algorithm.
[0032] In this step, calculate the reward value corresponding to the layout result using the set overall reward function, and periodically update the policy network until the set optimization goal is met.
[0033] In summary, the present invention first determines the placement position of each macro element in turn through the agent. After the overall layout of the macro elements is completed, perform post - processing on all elements together. Estimate the wire length and congestion metrics based on the layout result and feedback them to the agent to update the policy network, obtaining the agent for the next iteration. Specifically, it includes: (1) For the macro element at the current moment, obtain the global embedding features and node embedding features with different coarse and fine granularities respectively according to the policy network module (convolutional neural network and graph neural network), fuse the feature vectors obtained by the two networks, and predict the probability distribution of the possible placement positions of the element; (2) After the overall layout of the macro elements is completed, perform edge - fitting post - processing on all elements together to optimize the overlapping area of the layout result; (3) Estimate the wire length and congestion metrics based on the layout result obtained from the post - processing and feedback them to the agent, and periodically update the policy network through (Proximal Policy Optimization).
[0034] For clarity, Figure 3 schematically shows the simulation process of the provided end - to - end reinforcement learning hybrid - scale layout method based on post - processing, including the following steps:
[0035] Step S31, read the component data of the circuit to be laid out.
[0036] For example, the data preprocessing module receives a data file describing the components and netlist diagram information in the circuit to be laid out, and saves the extracted status to the computer memory for processing by the reinforcement learning layout model.
[0037] Step S32: Select the component to be placed at the current moment.
[0038] For example, construct a blank chessboard grid, initialize the result list, and select a macro component to be placed.
[0039] Step S33: Call the policy network to predict the position probability of the component.
[0040] For example, use the policy network to predict the placement position of the component and check whether the result is valid. If the selected position is already occupied, search for an alternative placement position and select the next component to be placed.
[0041] Step S34: Determine whether all macro components have been placed.
[0042] When there are still macro components not placed, go back to step S33; otherwise, enter step S35.
[0043] Step S35: Run the post-processing algorithm to adjust the component positions and calculate the offset loss.
[0044] The process of running the post-processing algorithm includes: for each macro component, calculate the distance from its boundary to the nearest component (if there is no component, calculate the distance to the chip board boundary); move the component in the direction of the nearest neighbor until it overlaps with another component or approaches the chip board boundary; calculate the moving distance of each component in the horizontal and vertical directions respectively, and add the sum of the squares of all moving distances as the offset loss to the reward function.
[0045] Step S36: Call the gradient-based layout optimization algorithm and calculate the reward value.
[0046] For example, run the gradient-based layout optimization algorithm DREAMPlace to read the positions of the macro components in the result list, and complete the placement of the standard components after iterative optimization. Calculate the reward value for this time (including the offset loss) according to the complete layout result, appropriately update the policy for selecting the agent's actions, go back to step S32 when the training is not over, and complete the training when the set number of rounds is reached.
[0047] To further verify the effectiveness of the present invention, tests were conducted on the specific dataset of ISPD-2005, which included a total of 8 different circuits with the total number of components ranging from 200,000 to 2 million. The experimental results show that compared with the existing baseline model without using the post-processing method, for the chip layout results obtained by the present invention, the average overlapping area is reduced by 2.14, the component regularity is improved by 2.67 times, and there is also a 0.2% reduction in the bus length index.
[0048] In summary, compared with the prior art, the present invention has the following advantages:
[0049] 1) The present invention first proposes a post-processing method based on the edge placement of the traditional placer in the end-to-end reinforcement learning layout model, thereby reducing the overlapping area of the layout results and simultaneously optimizing indicators such as component regularity.
[0050] 2) The present invention takes the offset loss of the computational post-processing algorithm as part of the reward function, thereby organically combining the reinforcement learning agent with the post-processing algorithm and improving the overall performance of the chip layout.
[0051] The present invention can be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.
[0052] The computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0053] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0054] The computer program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, Python, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via an Internet service provider through the Internet). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present invention.
[0055] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0056] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0057] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0058] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur in a different order than noted in the figures. For example, two consecutive boxes may, in fact, be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box of the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or can be implemented by a combination of dedicated hardware and computer instructions. It is well-known to those skilled in the art that implementations by hardware, by software, and by a combination of software and hardware are equivalent.
[0059] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to technologies in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.
Claims
1. An end-to-end reinforcement learning hybrid-scale layout method based on post-processing, comprising the following steps: For the layout problem of target components on a chip, construct an end-to-end reinforcement learning layout model, wherein the reinforcement learning layout model includes a policy network and a post-processing module, and the target components include macro components and standard components; Use the policy network to achieve the overall layout of all macro components; The post-processing module aims to reduce the overlapping area between components, and performs edge-attaching post-processing on all target components together to obtain the layout results of all target components; Calculate the reward value corresponding to the layout results using a set overall reward function, and periodically update the policy network until the set optimization goal is met; Wherein, the overall reward function further includes an offset loss calculated based on the component movement distance during the post-processing; Wherein, the post-processing module performs the following process: For each macro component, calculate the distance from its boundary to the nearest component, and in the case of no adjacent component, calculate the distance to the chip board boundary; Move the macro component in the direction of the nearest neighbor until it overlaps with another component or approaches the chip board boundary; Calculate the movement distance of each macro component in the horizontal and vertical directions respectively, take the sum of the squares of all movement distances as the offset loss and add it to the reward function to calculate the overall reward function; Run a gradient-based layout optimization algorithm to obtain the positions of macro components, and complete the placement of standard components after iterative optimization to obtain the layout results of all target components; Calculate the corresponding reward value according to the layout results, and this reward value includes the offset loss, and then update the layout policy according to the reward value.
2. The method according to claim 1, characterized in that the policy network includes a convolutional neural network and a graph convolutional network. The convolutional neural network is used to extract the global feature vector of the layout image, and the graph convolutional network is used to process the graph structure extracted from the netlist graph to obtain the local embedding feature vector of the netlist graph, and the overall layout of the macro components is obtained by splicing the global feature vector and the local embedding feature vector.
3. The method according to claim 1, characterized in that During the process of using the policy network to achieve the overall layout of all macro components, the reward function is the negative weighted sum of wire length, congestion and density.
4. The method according to claim 1, characterized in that the optimization goals of the reinforcement learning layout model include power consumption, performance and chip area.
5. The method according to claim 1, characterized in that the policy network is updated periodically using proximal policy optimization.
6. A computer-readable storage medium, on which a computer program is stored, wherein when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
7. A computer device, including a memory and a processor, and a computer program capable of running on the processor is stored on the memory, characterized in that when the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
System and method for optimizing chip layout based on deep reinforcement learning
CN114154412A
Chip macro-cell layout method and system based on lightweight deep reinforcement learning
CN114372438A