Generating integrated circuit placement using neural networks

Through the node placement neural network of reinforcement learning training, combined with encoder and strategy neural network, the problem of generating high-quality integrated circuit chip layout planning in the existing technology is solved, and high-quality layout planning with fast and low resource consumption is achieved, adapting to different netlists and chip sizes, and optimizing the placement of macro components.

CN120303665APending Publication Date: 2025-07-11GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202280102590.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-12-15
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to quickly generate high-quality integrated circuit chip layout plans without consuming excessive computing resources and requiring the participation of human experts, especially in the case of crowded chip surfaces.

Method used

The node placement neural network is adopted for reinforcement learning training, combined with the encoder and policy neural network, and through the cooperation of course learning and the default placement, a high-quality chip layout plan is gradually generated to avoid overlapping macro components.

Benefits of technology

It realizes the generation of high-quality chip layout plans that exceed human level in a few hours, reduces power consumption, increases computing power, adapts to new netlists and chip sizes, reduces computing resource requirements, and effectively places crowded blocks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303665A_ABST
    Figure CN120303665A_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating computer chip placement. One of the methods includes training a node placement neural network through reinforcement learning, the node placement neural network configured to: receive an input representation at each of a plurality of time steps, the input representation includes data representing a current state of placement up to the time step node netlist on the surface of the integrated circuit chip, and the input representation is processed to generate a fractional distribution over a plurality of locations on the surface of the integrated circuit chip.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION

[0001] This specification relates to using neural networks to implement electronic design automation, and more specifically, to generate computer chip placement.

[0002] Computer chip placement is a schematic representation of the placement of some or all of the circuits in the circuitry of a computer chip on the surface of the computer chip (i.e., the chip area).

[0003] A neural network is a machine learning model that employs one or more layers of non-linear units to predict an output for a received input. In addition to the output layer, some neural networks also include one or more hidden layers. The output of each hidden layer is used as an input to the next layer in the network (i.e., the next hidden layer or the output layer). Each layer of the network generates an output from the received input based on the current values of a corresponding set of parameters. SUMMARY OF THE INVENTION

[0004] This specification describes a system implemented as a computer program on one or more computers in one or more locations that generates a chip placement for an integrated circuit. In this specification, the integrated circuit for which a chip placement is generated will be referred to as a "computer chip", but should generally be understood to mean any collection of electronic circuits fabricated on a piece of semiconductor material. The chip placement places each node in a node netlist at a corresponding location on the surface of the computer chip.

[0005] Certain embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.

[0006] Floorplanning, which involves placing the components of a chip on the surface of the chip, is a key step in the chip design process. The placement of the components should optimize metrics such as area, total wire length, and congestion. If the floorplanning performs poorly in terms of these metrics, the integrated circuit chip generated based on the floorplanning will perform poorly. For example, the integrated circuit chip may not function properly, may consume excessive power, may have unacceptable latency, or may have any of various other undesirable properties resulting from a suboptimal placement of components on the chip.

[0007] The techniques described allow for the automatic generation of high-quality chip floorplans with minimal user involvement by leveraging the described node placement neural network and the described training techniques. As a specific example, when distributed training is employed, high-quality (i.e., superhuman) placements can be generated within hours without any human expert involvement.

[0008] Unlike the described system, conventional layout planning schemes employ a weeks-long process that requires significant human involvement. Due to the vast space of potential node placement combinations, conventional automated methods have been unable to reliably generate high-quality layout plans without consuming excessive computational power and wall-clock time, without the involvement of human experts, or both. However, by effectively leveraging reinforcement learning to train the described node placement neural network, the described technique is able to rapidly generate high-quality layout plans.

[0009] In addition, integrated circuit chips produced using this method can have reduced power consumption compared to integrated circuit chips produced by conventional methods. For a given surface area, the integrated circuit chips can also have increased computational power, or, from another perspective, for a given amount of computational power, fewer resources can be used to produce the integrated circuit chips.

[0010] Furthermore, when trained as described in this specification, i.e., when the encoder neural network is trained through supervised learning and the policy neural network is trained through reinforcement learning, the described node placement neural network can rapidly generalize to new netlists and new integrated circuit chip dimensions. This significantly reduces the amount of computational resources required to generate placements for new netlists, as generating high-quality layout plans for new netlists requires little or no computationally expensive fine-tuning.

[0011] Moreover, this specification describes techniques for generating high-quality layout plans even in the presence of congested blocks. A "congested block" is a chip, i.e., the entire chip or a portion of a larger chip, where macros consume most of the surface area of the chip. This makes it difficult to generate a valid placement, where a valid placement is one in which no macro nodes overlap on the surface of the chip. Thus, when training the node placement neural network through reinforcement learning, due to the congested nature of the chip surface, it is possible that the neural network only places a subset of nodes before the placement enters an infeasible state. This makes it difficult for the node placement neural network to receive meaningful rewards (i.e., rewards related to placement quality), e.g., because the placement process terminates once an infeasible state is entered and no rewards can be computed or received, or because the same default reward is received whenever the placement enters an infeasible state, and thus makes it difficult to train the node placement neural network through reinforcement learning to accurately place congested blocks.

[0012] By modifying the reinforcement learning pipeline to use curriculum learning, to place both macro nodes and standard cell nodes using a default placer when termination criteria are met, or both, the system can ensure that meaningful rewards are generated throughout training, thereby improving the resulting performance of the node placement neural network when placing congested blocks.

[0013] Details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 An example placement generation system is shown.

[0015] Figure 2 Processing of a node placement neural network at a time step is shown.

[0016] Figure 3 Is a flowchart of an example process for training a node placement neural network.

[0017] Figure 4 Is a flowchart of an example process for training a node placement neural network by reinforcement learning.

[0018] Figure 5 Is a flowchart of an example process for placing macro nodes at a given time step during training a node placement neural network by reinforcement learning.

[0019] Like reference numerals and names in the various figures indicate like elements. DETAILED DESCRIPTION

[0020] Figure 1 An example placement generation system 100 is shown. The placement generation system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations, in which the systems, components, and techniques described below are implemented.

[0021] The system 100 receives netlist data 102 of a computer chip (i.e., a very large scale integration (VLSI) chip) that is to be fabricated and includes a plurality of integrated circuit components, such as transistors, resistors, capacitors, etc. Depending on the desired function of the chip, the plurality of integrated circuit components can be different. For example, the chip can be a dedicated chip for machine learning computing, video processing, encryption, or another computationally intensive function, i.e., an application specific integrated circuit (ASIC). In some cases, the computer chip can be a part of a larger computer chip, e.g., a part of an ASIC that includes a certain subset of the components of the ASIC.

[0022] Netlist data 102 is data that describes the connectivity of integrated circuit components of a computer chip. In particular, netlist data 102 specifies the connectivity on a computer chip among a plurality of nodes, where each of the plurality of nodes corresponds to one or more of the integrated circuit components of the computer chip. That is, each node corresponds to a respective appropriate subset of the integrated circuit components, and these subsets do not overlap. In other words, for each of the plurality of nodes, netlist data 102 identifies which other nodes (if any) that node needs to be connected to by one or more wires in the manufactured computer chip. In some cases, the integrated circuit components have been clustered into clusters, for example, by an external system or by using existing clustering techniques, and each node in the netlist data represents a different integrated circuit component in a cluster.

[0023] System 100 generates as output a final computer chip placement 152 that places some or all of the nodes in the netlist data 102 at corresponding locations on the surface of the computer chip. That is, the final computer chip placement 152 identifies corresponding locations on the surface of the computer chip for some or all of the nodes in the netlist data 102 and thus for the integrated circuit components represented by the nodes. For convenience, this specification refers to the placement of a given set of components represented by nodes as placing the nodes that represent the components.

[0024] As an example, netlist data 102 can identify two types of nodes: nodes that represent macro components and nodes that represent standard cell components.

[0025] A macro component is a large block of IC components, for example, a static random access memory (SRAM) or other memory block, which is represented as a single node in the netlist. For example, a node that represents a macro component can include nodes that each represent a corresponding instance of an SRAM. As another example, a node that represents a macro component can include a hard macro composed of a fixed number of standard cells, for example, a macro composed of a fixed number of instances of a register file. As another example, a node that represents a macro component can include one or more nodes that each represent a phase-locked loop (PLL) circuit to be placed on the chip. As yet another example, a node that represents a macro component can include one or more nodes that each represent a sensor to be placed on the chip.

[0026] A standard cell component is a group of transistors and interconnect structures, for example, a group that provides a Boolean logic function (e.g., AND, OR, XOR, XNOR, inverter) or a group that provides a storage function (e.g., a flip-flop or a latch).

[0027] In some implementations, a node in the netlist data represents a single standard cell component. In some other implementations, a node in the netlist data represents a clustered standard cell component.

[0028] Typically, placer 152 assigns each node to a grid square in an N×M grid that covers the surface of the chip, where N and M are integers. In some cases, some or all of the nodes may be larger than a single grid square. In these cases, assigning a node to a grid square (or, more generally, to a site or location on the surface of the chip) means assigning a given location on the node (e.g., the center of the node) to a given location within that site or location (e.g., the center of the grid square).

[0029] In some implementations, the values of N and M are provided as an input to system 100.

[0030] In other implementations, system 100 generates the values of N and M.

[0031] For example, system 100 can view the problem of selecting the optimal number of rows and columns as a bin-packing problem and rank different combinations of rows and columns based on the amount of wasted space on the surface of the chip caused by the different combinations. Then, system 100 can select the combination that results in the least amount of wasted space as the values of N and M.

[0032] As another example, system 100 can use a grid generation machine learning model to process inputs derived from the netlist data, data representing the surface of the integrated circuit chip, or both, the grid generation machine learning model being configured to process the inputs to generate an output that defines how to divide the surface of the integrated circuit chip into an N×M grid.

[0033] System 100 includes a node placement neural network 110 and a graph placement engine 130.

[0034] System 100 uses the node placement neural network 110 to generate a macro node placement 122.

[0035] Specifically, the macro node placement 122 places each macro node (i.e., each node representing a macro) in the netlist data 102 at a corresponding location on the surface of the computer chip.

[0036] System 100 generates the macro node placement 122 by placing the corresponding macro nodes in the netlist data 102 at each time step in a sequence of multiple time steps.

[0037] That is, system 100 generates macro-node placements node-by-node over multiple time steps, where each macro-node is placed at a location at different time steps within the time step according to the macro-node order. The macro-node order sorts the macro-nodes, where each node in the macro-node order that comes before any given macro-node is placed before that given macro-node.

[0038] At each particular time step in the sequence, system 100 generates an input representation for that particular time step and processes the input representation using the node-placement neural network 110.

[0039] The input representation for a particular time step typically characterizes the placement state up to that time step.

[0040] For example, the input representation can at least characterize (i) the respective locations on the surface of the chip of any macro-node in the macro-node order that comes before the particular macro-node to be placed at the particular time step, and (ii) the particular macro-node to be placed at that particular time step.

[0041] The input representation can also optionally include data characterizing the connectivity specified in the netlist data 102 between the nodes. For example, for some or all of the nodes, the input representation can characterize one or more other nodes to which the node is connected according to the netlist. For example, the input representation can represent each connection between any two nodes as an edge connecting the two nodes.

[0042] Reference is made below Figure 2 to describe the example input representation in more detail.

[0043] In the first time step of the sequence, the input representation indicates that no nodes have been placed yet, and thus for each node in the netlist, indicates that the node does not yet have a location on the surface of the chip.

[0044] The node-placement neural network 110 is a neural network that has parameters (referred to in this specification as "network parameters") and is configured to process the input representation according to the current values of the network parameters to generate a score distribution over multiple locations on the surface of the computer chip, e.g., a probability distribution or a distribution of logits. For example, the distribution can be over grid squares in an N×M grid covering the surface of the chip.

[0045] Then, system 100 uses the score distribution generated by the neural network to assign the macro-node to be placed at a particular time step to a location among the multiple locations.

[0046] Reference is made below Figures 2 to 4 to describe in more detail the operations performed by the neural network 110 at a given time step and the use of the score distribution to place nodes at that time step.

[0047] By adding macro nodes to the placement one by one, after the last time step in the sequence, the macro node placement will include the corresponding placements for all the macro nodes in the macro nodes of the netlist data 102.

[0048] Once the system 100 has generated the macro node placement 122, the graphics placement engine 130 generates an initial computer chip placement 132 by placing each standard cell in the standard cells at corresponding positions on the surface of the partially placed integrated circuit chip, where the partially placed integrated circuit chip includes macro components represented by the macro nodes placed according to the macro node placement (i.e., placed as in the macro node placement 122).

[0049] In some implementations, the engine 130 clusters the standard cells into a set of standard cell clusters (or obtains data identifying the clusters that have been generated), and then uses a default placer to place each cluster of standard cells at corresponding positions on the surface of the partially placed integrated circuit chip.

[0050] As a specific example, the engine 130 can use a partitioning technique based on a normalized min - cut objective to cluster the standard cells. An example of such a technique is hMETIS, which is described in: Karypis, G. and Kumar, V. A hypergraph partitioning package. In HMETIS, 1998.

[0051] In some other implementations, the engine 130 does not cluster the standard cells and uses a default placer to directly place each standard cell at a corresponding position on the surface of the partially placed integrated circuit chip.

[0052] The default placer can be any suitable software for placing the nodes of the netlist on the surface of the chip starting from an initial placement.

[0053] For example, the default placer can be an analytical placer that generates an analytical solution to the placement problem given an initial placement and the remaining nodes on the netlist.

[0054] For example, the default placer may utilize a graphical placement technique for placing graphical nodes. For example, the default placer may use a force-based technique, i.e., a force-directed technique. In particular, when using the force-based technique, the engine 130 represents the netlist as a spring system that applies forces to each node according to the weight × distance formula, causing closely connected nodes to attract each other. Optionally, the engine 130 also introduces repulsive forces between overlapping nodes to reduce the placement density. After applying all the forces, the engine 130 moves the nodes in the direction of the force vectors. To reduce oscillation, the engine 130 may set a maximum distance for each movement. The use of the force-directed technique for placing nodes is described in more detail in the following literature: Shahookar, K. and Mazumder, P. Vlsi cell placement techniques. ACM Comput.Surv., 23(2):143220, June 1991. ISSN 0360-0300. doi: 10.1145 / 103724.103725.

[0055] As another example, the default placer may be an analytical placer for placing nodes on a netlist, which models placement and the netlist as an electrostatic system. Examples of such placers include ePlace and RePlace. These types of placers are described in more detail in, for example, the following literature: Lue, et al, ePlace: Electrostatics-based Placement using Fast Fourier Transform and Nesterov’s Method.

[0056] As yet another example, the default placer may be an analytical placer that solves an analytical placement problem, e.g., an analytical placer that models placement and the netlist as an electrostatic system, which is equivalent to training a neural network. Examples of such placers are described in the following literature: Lin, et al, DREAMPlace: Deep Learning Toolkit-Enabled GPU Acceleration for Modern VLSI Placement.

[0057] In some implementations, the system 100 uses the initial placement 132 as the final placement 152.

[0058] In some other implementations, the system 100 provides the initial placement 132 as an input to the legalization engine 150, which adjusts the initial placement 132 to generate the final placement 152.

[0059] Specifically, the legalization engine 150 can generate a legalized integrated circuit chip placement by applying a greedy legalization algorithm to an initial integrated circuit chip placement. For example, the engine 150 can perform a greedy legalization step to snap macros to the nearest legal position while complying with minimum spacing constraints.

[0060] Optionally, the engine 150 can further refine the legalized placement to generate the final placement 152, or can directly refine the initial position 132 without generating a legalized placement to generate the final placement 152, e.g., by performing simulated annealing on a reward function.

[0061] Example reward functions are described in more detail below.

[0062] As a specific example, the engine 150 can perform simulated annealing by applying a hill climbing algorithm to iteratively adjust the placement in the legalized placement or the initial placement 132 to generate the final computer chip placement 152. Hill climbing algorithms and other simulated annealing techniques that can be used to adjust the macro node placement 122 are described in more detail in the following: S. Kirkpatrick, C. D. Gelatt, and M. P. Vecchi. Optimization by simulated annealing, SCIENCE, 220(4598):671–680, 1983. As another example, the system 100 can further refine the legalized placement, or can directly refine the initial placement 132 without generating a legalized placement by providing the legalized placement or the initial placement 132 to an electronic design automation (EDA) software tool for evaluation and fine-tuning.

[0063] Optionally, the system 100 or an external system can then fabricate (produce) a chip (integrated circuit) based on the final placement 152. Such integrated circuits can exhibit improved performance, e.g., have one or more of lower power consumption, lower latency, or smaller surface area compared to integrated circuits designed using a conventional design process, and / or can be produced using fewer resources.

[0064] Fabrication can use any known technique.

[0065] In some cases, fabricating a chip based on the final placement can include presenting data identifying the placement to a user to allow the user to modify the final placement 152 before fabrication, or providing the final placement 152 to electronic design automation (EDA) for fine-tuning before fabrication.

[0066] The system 100 can receive the netlist data 102 in any of a variety of ways.

[0067] For example, system 100 can receive netlist data 102 as an upload from a remote user of the system via a data communication network, for example, using an application programming interface (API) available to system 100. In some cases, system 100 can then provide the final placement 152 to the remote user via an API provided by system 100, for example, for manufacturing a chip according to the final placement 152.

[0068] As another example, system 100 can be part of an electronic design automation (EDA) software tool and can receive netlist data 102 from a user of the tool or from another component of the tool. In this example, system 100 can provide the final placement 152 for evaluation by another component of the EDA software tool before a computer chip is manufactured.

[0069] To use neural network 110 to generate a high-quality placement, the system (or another system) trains neural network 110 on training data.

[0070] In some implementations, the system uses reinforcement learning to train neural network 110 end-to-end to maximize the received expected reward as measured by a reward function. The reward function generally measures the quality of the placement generated using node placement neural network 110. The reward function will be described in more detail below with reference to Figure 3 more details.

[0071] In some implementations, to improve the generalization of neural network 110 to new netlists after training, the system can pre-train some components of neural network 110 via supervised learning and then train other components via reinforcement learning. Such training processes will be described in more detail below with reference to Figure 3 more details.

[0072] Generally, when training neural network 110 via reinforcement learning, for example, from scratch or after pre-training, the system repeatedly uses neural network 110 to generate placements for a netlist from a training data set and computes the corresponding value of the reward function for each placement.

[0073] To improve the effectiveness of training, for example, to improve the performance of the neural network when generating placements for crowded chip blocks, the system can make one or more modifications to how neural network 110 is used to generate placements during training relative to the above techniques.

[0074] As an example, during training, the system can determine whether the placement of each macro node will cause the placement to enter an infeasible state after the macro node is placed.

[0075] The placed "infeasible state" is the state where two macro components overlap each other. That is, the corresponding parts of the two in the two macro components are placed at the same point on the surface of the chip.

[0076] Once the system determines that placing a given macro node will cause the placement to enter an infeasible state, instead of continuing to use the neural network 110 to place the macro node, the system alternatively uses a default placer to place the remaining macro nodes in the macro node. That is, instead of using the default placer to place only standard cells, the system uses the default placer to place both the remaining macro nodes and standard cells in the macro node order.

[0077] This is described in more detail below with reference to Figure 4 and Figure 5 for a more detailed description.

[0078] Alternatively or additionally, the system can utilize a schedule that defines how many macro nodes in the macro nodes are to be placed by the neural network 110 and how many macro nodes in the macro nodes are to be placed using the default placer. That is, before training on a given training netlist, the system identifies a subset of the macro nodes in the training netlist that are to be placed using the neural network 110. Once those nodes have been placed using the neural network 110, the system uses the default placer to place any remaining macro nodes and standard cells. As the training progresses, the system can increase the proportion of macro nodes to be placed by the neural network 110.

[0079] Using curriculum learning results in a method where the placement task starts with easy tasks and then gradually becomes more difficult. That is, as the number of macros to be placed increases, the layout planning task becomes more difficult for the neural network and starts "easier" because using the default placer reduces the number of macros to be placed by the neural network. However, as will be described in more detail below, using the default placer may result in macro overlaps and thus an illegal placement that cannot actually be fabricated. By gradually increasing the percentage of macros placed by the neural network until the percentage reaches 100, at which point all macros in the macro are placed by the neural network and the generated placement will be guaranteed to be legal, the system can improve the exploration and placement quality, for example, for crowded blocks, or more generally, for blocks with a large number of macros.

[0080] This is described in more detail below with reference to Figure 5 for a more detailed description.

[0081] Figure 2 Shows an example of the processing of the node placement neural network 110 at a given time step.

[0082] As referred to above with reference to Figure 1As described, at each time step during placement generation, the node placement neural network 110 is configured to receive an input representation and process the input representation to generate a score distribution over a plurality of locations on the surface of a computer chip, e.g., a probability distribution or a distribution of logits.

[0083] Typically, the input representation at least includes (i) data characterizing the respective locations on the surface of the chip of any macro nodes preceding a particular macro node to be placed at a particular time step in the macro node order, and (ii) data characterizing the particular macro node to be placed at the particular time step.

[0084] As Figure 2 shown, the node placement neural network 110 includes an encoder neural network 210, a policy neural network 220, and optionally a value neural network 230.

[0085] The encoder neural network 210 is configured to process the input representation at each particular time step to generate an encoded representation 212 of the input representation. The encoded representation is a numerical representation in a fixed-dimensional space, i.e., an ordered set of a fixed number of numerical values. For example, the encoded representation can be a vector or matrix of floating-point values or other types of numerical values.

[0086] The policy neural network 220 is configured to process the encoded representation 212 at each particular time step to generate a score distribution.

[0087] Typically, the policy neural network 220 can have any suitable architecture that allows the policy neural network 220 to map the encoded representation 212 to a score distribution. As Figure 2 shown in the example, the policy neural network 220 is a deconvolutional neural network that includes a fully-connected neural network followed by a set of deconvolutional layers. The policy neural network 220 can optionally include other types of neural network layers, e.g., batch normalization layers or other kinds of normalization layers. However, in other examples, the policy neural network 220 can be, e.g., a recurrent neural network, i.e., a neural network that includes one or more recurrent neural network layers (e.g., long short-term memory (LSTM) layers, gated recurrent unit (GRU) layers, or other types of recurrent layers) having an output layer that generates scores for locations. For example, when the scores are probabilities, the output layer can be a softmax layer.

[0088] The value neural network 230, when in use, is configured to process the encoded representation 212 at each particular time step to generate a value estimate that estimates the value of the current state of the placement as of that particular time step. The value of the current state is an estimate of the output of a reward function for the placement generated starting from the current state (i.e., starting from the current partial placement). For example, the value neural network 230 can be a recurrent neural network or can be a feed-forward neural network, e.g., a neural network including one or more fully-connected layers.

[0089] This value estimate can be used during the training of the neural network 110 (i.e., when using reinforcement learning techniques that rely on the available value estimates). In other words, when the reinforcement learning technique used to train the node placement neural network requires a value estimate, the node placement neural network 110 also includes the value neural network 230 that generates the value estimate required by the reinforcement learning technique.

[0090] The training of the node placement neural network 110 will be described in more detail below.

[0091] As Figure 2 shown in the example of, the input feature representation includes the respective vectorized representations (“macro features”) of some or all of the nodes in the netlist, the “netlist graph data” representing the connectivity between the nodes in the netlist as edges each connecting two respective nodes in the netlist data, and the “current macro id” identifying the macro node being placed at a particular time step. As a particular example, the input feature representation can include the respective vectorized representations of only macro nodes, of macro nodes and standard cell clusters, or of macro nodes and standard cell nodes.

[0092] Each vectorized representation characterizes the corresponding node. In particular, for each node that has been placed, the vectorized representation includes data identifying the position of the node on the surface of the chip, e.g., the coordinates of the center of the node or the coordinates of some other marked part of the node, and for each node that has not been placed, the vectorized representation includes data indicating that the node has not been placed, e.g., including default coordinates indicating that the node has not been placed on the surface of the chip. The vectorized representation can also include other information characterizing the node, e.g., the type of the node, the size of the node (e.g., the height and width of the node), etc.

[0093] In Figure 2 the example of, the encoder neural network 210 includes a graph encoder neural network 214 that processes the vectorized representations of the nodes in the netlist to generate (i) a netlist embedding of the vectorized representations of the nodes in the netlist and (ii) a current node embedding representing the macro node to be placed at a particular time step. An embedding is a numerical representation in a fixed-dimensional space, i.e., an ordered set of a fixed number of numerical values. For example, an embedding can be a vector or matrix of floating-point values or other types of numerical values.

[0094] Specifically, the graph encoder neural network 214 randomly initializes, for example, the respective edge embeddings of each edge in the netlist data and initializes the respective node embeddings of each node in the netlist data, i.e., such that the node embedding is equal to the respective vectorized representation for that node.

[0095] Then, the graph encoder neural network 214 repeatedly updates the node embeddings and edge embeddings by updating the embeddings at each message passing iteration among multiple message passing iterations.

[0096] After the last message passing iteration, the graph encoder neural network 214 generates a netlist embedding and a current node embedding from the node embeddings and edge embeddings.

[0097] As a specific example, the neural network 214 can generate a netlist embedding by combining the edge embeddings after the last message passing iteration. For example, the system can compute the netlist embedding by applying a reduce mean function to the edge embeddings after the last message passing iteration.

[0098] As another specific example, the neural network 214 can set the current node embedding of the current node to be equal to the embedding of the current node after the last message passing iteration.

[0099] The neural network 214 can use any one of various message passing techniques to update the node embeddings and edge embeddings at each message passing iteration.

[0100] As a specific example, at each message passing iteration, the neural network 214 uses the respective node embeddings of the two nodes connected by an edge to update the edge embedding of each edge.

[0101] At each iteration, to update the embedding of a given edge, the network 214 generates an aggregated representation at least from the node embeddings of the two nodes connected by that edge and uses a first fully connected neural network to process the aggregated representation to generate an updated edge embedding for the given edge. In some implementations, each edge has the same weight in the netlist data, i.e., one. In some other implementations, each edge is associated with a respective weight in the netlist data, and the system generates an aggregated representation from the node embeddings of the two nodes connected by the edge and the weight associated with the edge in the netlist data. For example, the weight of each edge can be learned jointly with the training of the neural network.

[0102] To update the embedding of a given node at a given message passing iteration, the system uses the respective edge embeddings of the edges connected to that node to update the node embedding of that node. For example, the system can average the respective edge embeddings of the edges connected to that node.

[0103] The input feature representation may also optionally include "netlist metadata" characterizing the node netlist. The netlist metadata may include any suitable information characterizing the netlist. For example, the information may include any of the following information: underlying semiconductor technology (horizontal and vertical routing capacity); total number of nets (edges), macros, and standard cell clusters in the netlist; canvas size, i.e., the size of the surface of the chip; or number of rows and columns in the grid.

[0104] When the input feature representation includes netlist metadata, the encoder neural network 210 may include a fully connected neural network that processes the metadata to generate a netlist metadata embedding.

[0105] The encoder neural network 210 generates an encoded representation at least from the netlist embedding of the vectorized representation of the nodes in the netlist and the current node embedding representing the macro node to be placed at a particular time step. When the encoder neural network 210 also generates a netlist metadata embedding, the system also uses the netlist metadata embedding to generate the encoded representation.

[0106] As a specific example, the neural network 210 may concatenate the netlist embedding, the current node embedding, and the netlist metadata embedding, and then use a fully connected neural network to process the concatenation to generate the encoded representation.

[0107] The system also tracks the density of the locations on the chip, i.e., the density of the squares in the grid. In particular, the system maintains a density value for each location, which indicates the degree to which the location is occupied. When a node has been placed at a given location, the density value of that location is set to equal one (or set to a different maximum value indicating that the location is fully occupied). When no node has been placed at a given location, the density value of that location indicates the number of edges passing through that location. The density value of a given location may also reflect blockages that block certain parts of the chip surface, such as clock straps or other structures, by setting the values of those locations to one.

[0108] Once the policy neural network 220 has generated a score distribution at a time step, the system uses the density to generate a modified score distribution, and then uses the modified score distribution to assign a node corresponding to that time step. In particular, the system modifies the score distribution by setting the score of any location having a density value that meets (e.g., exceeds) a threshold to zero.

[0109] For example, the system may assign the node to the location having the highest score in the modified score distribution, or sample a location from the modified score distribution, i.e., such that each location has an equal likelihood of being selected, and then assign the node to the sampled location.

[0110] This is represented in Figure 2 as a grid density mask that can be applied to a fractional distribution, i.e., a mask in which the value of any position with a density above a threshold is zero and the value of any position with a density not above the threshold is one, to generate a modified fractional distribution.

[0111] As a specific example, the threshold can be equal to one, and the system can set the fraction of any position where a node has been placed (i.e., a position with a density value of one) to zero. As another example, the threshold can be less than one, indicating that the system also sets the fraction of any position that does not have a node but has too many wires passing through it (i.e., the number of wires associated with the position is above the threshold) to zero.

[0112] As described above, in order to use the neural network 110 to generate a high-quality placement, the system (or another system) trains the neural network on training data.

[0113] In some implementations, the system uses reinforcement learning to train the neural network 110 end-to-end to maximize the expected reward received as measured by a reward function. The reward function generally measures the quality of the placement generated using the node placement neural network 110. This will be described in more detail below with reference to Figure 3 the reward function.

[0114] To improve the generalization of the neural network 110, the system can train the encoder neural network 210 through supervised learning and then train the policy neural network 220 through reinforcement learning. This will be described in more detail below with reference to Figure 3 such training processes.

[0115] Figure 3 is a flowchart of an example process 300 for training a node placement neural network. For convenience, process 300 will be described as being performed by a system of one or more computers located at one or more locations. For example, a suitably programmed placement generation system (e.g., Figure 1 the placement generation system 100) can perform process 300.

[0116] The system can perform process 300 to train the node placement neural network, i.e., to determine the trained values of the network parameters.

[0117] In some implementations, the system distributes the training of the node-placement neural network across many different workers (i.e., across many different homogeneous or heterogeneous computing devices, i.e., devices that use CPUs, GPUs, or ASICs to perform training computations). In some of these implementations, some or all of the steps in step 300 can be performed in parallel by many different workers that operate asynchronously with respect to each other in order to accelerate the training of the node-placement neural network. In other implementations, the different workers operate synchronously to perform some or all of the steps of process 300 in parallel in order to accelerate the training of the neural network.

[0118] The system can use process 300 to train any node-placement neural network, where any node-placement neural network includes: (i) an encoder neural network that is configured to receive, at each of a plurality of time steps, an input representation that includes data representing the current state of the placement of a node netlist on the surface of an integrated circuit chip, and to process the input representation to generate an encoder output; and (ii) a policy neural network that is configured to receive, at each of the plurality of time steps, an encoded representation generated from the encoder output generated by the encoder neural network, and to process the encoded representation to generate a score distribution over a plurality of locations on the surface of the integrated circuit chip.

[0119] An example of such a neural network is the neural network described above with reference to Figure 2 the neural network described above.

[0120] Another example of such a neural network is described in the following application: Application No. 16 / 703,837, filed on December 4, 2019, entitled "GENERATING INTEGRATED CIRCUIT FLOORPLANS USING NEURAL NETWORKS", the entire content of which is hereby incorporated herein by reference in its entirety.

[0121] The system obtains supervised training data (step 302).

[0122] The supervised training data includes (i) a plurality of training input representations, each training input representation representing a corresponding placement of a corresponding node netlist, and (ii) for each training input representation, a corresponding target value that includes a reward function that measures the quality of the placement of the corresponding node netlist.

[0123] More specifically, the reward function measures certain characteristics of the generated placement, which when optimized cause a chip fabricated using the generated placement to exhibit good performance in one or more of, for example, power consumption, heat generation, or timing performance.

[0124] In particular, the reward function includes corresponding terms for one or more characteristics. For example, when there are multiple terms, the reward function can be the sum or weighted sum of the multiple terms.

[0125] As an example, the reward function can include a wire length measurement, i.e., a term that measures the wire length of the wires on the surface of the chip, and the higher this wire length measurement is when the wire length between nodes on the surface of the chip is shorter.

[0126] For example, the wire length can be the Manhattan distance or the negative value of other distance measurements between all adjacent nodes among adjacent nodes on the surface of the chip.

[0127] As another example, the wire length measurement can be based on the half-perimeter wire length (HPWL), which approximates the wire length using the half-perimeter of the bounding box of all nodes in the netlist. When calculating the HPWL, the system can assume that all wires leaving the standard cell cluster originate from the center of the cluster. In particular, the system can calculate the HPWL of each edge in the netlist and then calculate the wire length measurement as the negative value of the normalized sum equal to the HPWL of all edges in the edges of the netlist.

[0128] Including a term that measures the wire length in the reward function has the advantage that the wire length roughly measures the routing cost and is also related to other important metrics such as power and timing.

[0129] As another example, the reward function can include a congestion measurement, i.e., a term that measures congestion, and the higher this congestion measurement is when the congestion on the surface of the computer chip is lower. Congestion is a measure of the difference between the available routing resources in a given area (not necessarily a continuous area) on the chip compared to the actual wires passing through that area. For example, congestion can be defined as the ratio of the wires passing through the area in the generated placement to the available routing resources (e.g., the maximum number of wires that can pass through that area). As a specific example, the congestion measurement can track the density of wires across the horizontal and vertical edges of the surface.

[0130] In particular, the system can utilize a routing model for the netlist (e.g., net bounding box, upper L, lower L, A*, minimum spanning tree, or the network of the actual routing, etc.). Based on this routing model, the congestion measurement can be calculated by determining the ratio of the available routing resources in the placement to the routing estimate from the routing model for each position on the surface.

[0131] As another example, the system can calculate congestion measurements by separately keeping track of the vertical and horizontal allocations at each location, e.g., the congestion measurements calculated as described above. The system can then smooth the congestion estimate by running a convolutional filter (e.g., a 5×1 convolutional filter or a filter of a different size depending on the number of locations in each direction) in both the vertical and horizontal directions. The system can then calculate the congestion measurement as the negative of the average of the top 10%, 15%, or 20% of the congestion estimate.

[0132] As another example, the reward function can include a timing term, i.e., a term that measures the timing of the digital logic, which is higher when the performance of the chip is better (e.g., for a placement of a corresponding chip that takes less time to perform a certain computational task, the reward function takes a correspondingly higher value). Static timing analysis (STA) can be used to measure the timing or performance of the placement. The measurement can include calculating the stage delays on the logic paths (including internal cell delays and wire delays) and finding the critical path that will determine the maximum speed at which the clock can run for safe operation. For a realistic view of the timing, logic optimization may be necessary to accommodate paths getting longer or shorter as the node placement progresses.

[0133] As another example, the reward function can include one or more terms that measure the power or energy to be consumed by the chip, i.e., one or more terms that are higher when the power to be consumed by the chip is lower.

[0134] As another example, the reward function can include one or more terms that measure the area of the placement, i.e., one or more terms that are higher when the area occupied by the placement is lower.

[0135] In some cases, the system receives supervised training data from another system.

[0136] In other cases, the system generates supervised training data. As a specific example, based on the output of a different node placement neural network (e.g., a node placement neural network having a simpler architecture than the architecture described above with reference to Figure 2 placements represented by multiple training inputs can be generated at different time points during the training of the different node placement neural networks on different netlists. This can ensure that the placements have varying quality.

[0137] For example, the system can generate supervised training data by selecting a set of different accelerator netlists and then generating placements for each netlist. To generate diverse placements for each netlist, the system can train a simpler policy network on the netlist data with various congestion weights (ranging from 0 to 1) and random seeds, for example, through reinforcement learning, and collect snapshots of each placement during the process of policy training. Each snapshot includes a representation of the placement and a reward value generated for the placement by a reward function. The untrained policy network starts with random weights, and the generated placements have low quality, but as the policy network is trained, the quality of the generated placements improves, allowing the system to collect a diverse dataset using placements with varying quality.

[0138] In some implementations, the training input representation can all represent the final placement, i.e., the placement that places all the macro nodes in the corresponding netlist. In some other implementations, the training input representation can represent the placements at various stages of the placement generation process, i.e., some representations can represent partial placements that place only some of the macro nodes.

[0139] The system jointly trains an encoder neural network on the supervised training data with a reward prediction neural network through supervised learning (step 304).

[0140] The reward prediction neural network is configured to: for each training encoder input, receive the encoder output generated by the encoder neural network from the training input representation, and process the encoded representation to generate a predicted value of the reward function for the placement represented by the training input representation.

[0141] The reward prediction neural network can be, for example, a fully connected neural network that receives the encoder output and processes the encoder output to generate a reward prediction. When the encoder neural network has the architecture described above with reference to Figure 2 The encoder output can be a concatenation of a netlist graph embedding and a metadata embedding.

[0142] For example, the system can train the encoder neural network and the reward prediction neural network to optimize an objective function (e.g., mean squared error loss) that measures the error between the target value of the reward function and the predicted value of the reward function for the training input representation for a given training representation.

[0143] Then, the system trains a policy neural network through reinforcement learning to generate a score distribution that leads to placements that maximize the reward function, or, as will be described below, to generate a modified reward function that also takes into account overlapping macro nodes (step 306). The system can use any of various reinforcement learning techniques to train the node placement neural network.

[0144] For example, the system can be trained using policy gradient techniques (e.g., REINFORCE or Proximal Policy Optimization (PPO)). In these cases, when the neural network includes a value prediction neural network, the value prediction generated by the value neural network can be used to calculate a baseline value that modifies the reward function value when calculating the gradient of the reinforcement learning loss function.

[0145] While training the policy neural network via reinforcement learning, the system can keep the values of the parameters of the encoder neural network fixed to the values determined by training on supervised training data, or can fine-tune the values as part of the training of the policy neural network.

[0146] In particular, in some cases, while training the policy neural network via reinforcement learning for a given chip on a given netlist, the system can use the placement neural network to place the macro nodes in the given netlist one by one as described above. After the macro nodes have been placed, the system can place the standard cell nodes as described above to determine the final placement. Then, the system can calculate a reward function or a modified reward function for the final placement, for example, by calculating the required quantities described above, and use the reward value, the macro node placement, and the score distribution generated by the placement neural network to train the placement neural network via reinforcement learning. Thus, although the placement neural network is only used to place macro nodes, the reward value is calculated only after the standard cell nodes have also been placed, ensuring that the placement neural network generates macro node placements that still allow for high-quality placement of the standard cell nodes.

[0147] In some other cases, while training the policy neural network via reinforcement learning for a given chip on a given netlist, the system can use the placement neural network to place the macro nodes in the given netlist one by one until a termination criterion is met. After the termination criterion has been met, the system uses a default placer to place any remaining macro nodes and standard cell nodes to determine the final placement.

[0148] This is described in more detail below with reference to Figure 4 and Figure 5 For a more detailed description.

[0149] The system receives new netlist data (step 308).

[0150] In some implementations, the system uses a trained node placement neural network to generate an integrated circuit placement for the new netlist data, i.e., by placing the corresponding nodes in the new netlist data at each of multiple time steps using the score distribution generated by the trained node placement neural network (step 310). That is, the system generates a placement for the new netlist data without further training the node placement neural network at all.

[0151] That is, by training an encoder neural network via supervised learning and then training a policy neural network via reinforcement learning, the system trains a node placement neural network to generalize to new netlists without any additional training.

[0152] In some other implementations, to further improve the quality of the placement generated for new netlists, the system first fine-tunes the trained node placement neural network on new netlist data via reinforcement learning (step 312), and then uses the fine-tuned node placement neural network to generate an integrated circuit placement for the new netlist data (step 314), as described above. The system can use the same reinforcement learning techniques as described above during fine-tuning, and depending on the implementation, can keep the parameter values of the encoder neural network fixed or update the parameter values of the encoder neural network during this fine-tuning.

[0153] Figure 4 is a flowchart of an example process 400 for performing training steps while training a node placement neural network via reinforcement learning. For convenience, process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, a suitably programmed placement generation system (e.g., Figure 1 placement generation system 100) can perform process 400.

[0154] The system can repeatedly perform process 400 on different batches of training netlists to train the node placement neural network, i.e., to repeatedly update the values of the parameters of the node placement neural network. As described above, when the node placement neural network includes both an encoder neural network and a policy neural network, in some implementations, the system updates the values of the parameters of both the encoder neural network and the policy neural network, while in other implementations, the system only updates the values of the policy neural network while keeping the encoder neural network fixed.

[0155] The system obtains a batch that includes one or more training netlists (step 402).

[0156] For each training netlist, the system identifies a subset of the macro nodes in the training netlist to be placed by the node placement neural network (step 404).

[0157] In some implementations, the system does not employ curriculum learning and uses the node placement neural network to select all of the macro nodes to be placed in each training netlist at all iterations of process 400.

[0158] In some other implementations, the system employs curriculum learning and adjusts the size of the node subset across different iterations of process 400.

[0159] That is, the system determines the size of the subset of nodes for each training netlist based on how many iterations (i.e., how many training steps) the process 400 has performed during the reinforcement learning training.

[0160] More specifically, the system can use a schedule (function) to determine the size, which maps data characterizing the current training step (i.e., the current iteration of process 400) to a portion (a fraction or percentage) of the macro nodes in each training netlist that should be included in the subset. Generally, the schedule can be any suitable function that increases the size of this portion as the training progresses. For example, the schedule can be a non-decreasing function that maps the first training step to a predetermined portion less than all of the macro nodes in the training netlist, and maps a later training step (e.g., a training step with a specified index or occurring at a specified time during the training) to a portion corresponding to all of the macro nodes in the training netlist.

[0161] As a specific example, the learning schedule can be defined as a function that maps a curriculum learning progress ratio in [0, 1] to a fraction of the macro nodes to be included in the subset to be placed by the neural network. The curriculum learning progress ratio can be, for example, the ratio of the index of the current training step to the total number of training steps, or the ratio of the currently elapsed training time to the total amount of allotted training time. An example of such a function is (exp(a * x) - 1) / (exp(a) - 1), where a is a parameter that determines the shape of the function and x is the learning progress ratio. For example, if a is between -2 and -6, the function can have a concave shape, while if a is between 2 and 6, the function can have a convex shape.

[0162] Once the system determines how many macro nodes are in the subset, the system can select which macro nodes to add to the subset in any suitable way. For example, the system can randomly select macro nodes from the netlist until the total number of selected nodes reaches the identified size of the subset. As another example, the system can identify a grouping that divides the macro nodes into multiple hierarchical groups, and select macro nodes such that the subset includes at least one macro node from each hierarchical group, e.g., by randomly selecting a macro node from each group and then randomly selecting macro nodes until the subset reaches the identified size.

[0163] Then, the system can generate a corresponding placement for each training netlist.

[0164] Specifically, for each training netlist, the system generates a partial placement (step 406) by placing the macro nodes in the training netlist according to the macro node order of the training netlist using a node placement neural network until a termination criterion is met.

[0165] More specifically, the termination criterion is satisfied when (i) each macro node in the identified subset has been placed or (ii) the system determines that placing a particular macro node in the subset will cause the placement to enter an infeasible state.

[0166] That is, the system places the macro nodes in the subset one by one according to the macro node order until the termination criterion is satisfied. Thus, in the case of branch (i) where the termination criterion is satisfied, the partial placement includes the placement for each macro node in the subset, or in the case of branch (ii) where the termination criterion is satisfied, the partial placement includes the placement for each macro node in the macro node order up to and including the particular macro node.

[0167] In some implementations, the system receives the macro node order and netlist data as input.

[0168] In some other implementations, the system can generate the macro node order from the netlist data.

[0169] As an example, the system can sort the macro nodes according to size—for example, in decreasing size—and use topological classification to break connections. By placing larger macros first, the system reduces the chance that later macros will have no feasible placement. Topological classification can help the policy network learn to place connected nodes close to each other.

[0170] As another example, the system can process the input derived from the netlist data through a macro node order prediction machine learning model, which is configured to process the input derived from the netlist data to generate an output that defines the macro node order.

[0171] As yet another example, the node placement neural network can be further configured to generate a probability distribution over the macro nodes. Then, the system can dynamically generate the macro node order by selecting, for each particular time step among multiple time steps, the macro node to be placed at the next time step after the particular time step based on the probability distribution over the macro nodes. For example, the system can select the unplaced macro node with the highest probability.

[0172] The placement of macro nodes and determination of whether the termination criterion has been satisfied are described in more detail below with reference to Figure 5 More specifically, the placement of macro nodes and determination of whether the termination criterion has been satisfied are described in more detail below with reference to

[0173] For each training netlist, once the termination criterion has been satisfied, the system uses the above-described default placer to generate a placement for each training netlist starting from the partial placement (step 408).

[0174] In particular, for each training netlist, the system uses a default placer to place (i) any remaining macro nodes in the training netlist that have not been placed in the partial placement for the training netlist and (ii) the standard cell nodes in the training netlist (step 410). More specifically, when the identified iteration for the current training step is not a proper subset, the system uses the default placer to place (i) any remaining macro nodes in the subset, i.e., the macro nodes after the specific macro nodes that meet the criteria, and (ii) the standard cell nodes in the training netlist.

[0175] Thus, for at least some training steps, the system not only places the standard cells, but also uses the default placer to place one or more remaining macro nodes.

[0176] For each training netlist, the system calculates the reward value of the reward function for the placement for the training netlist (step 410).

[0177] Generally, as described above, the reward function measures the quality of the placement of the corresponding node netlist. Examples of components of the reward function are described above with reference to Figure 3 Describe the example components of the reward function.

[0178] In some implementations, the system modifies the above reward function to include an additional term that penalizes the neural network for generating an output that results in an overlapping macro node in the placement (i.e., after the placement is completed using the default placer).

[0179] In particular, the placement can have overlapping nodes because the default placer is required to place nodes starting from an infeasible state or from a state that the system determines will cause an infeasible state. Additionally, placing macro nodes can be a difficult task for the default placer, and thus, even starting from a non-infeasible state, the default placer may generate a placement with overlaps, i.e., may generate an illegal (i.e., with overlaps) placement because, given the complexity of the task, the placer cannot "find" a legal placement.

[0180] To address this issue and to discourage the neural network from generating output that results in a placement state that causes overlaps in the final placement, the system can include in the loss function a term that measures the degree of overlap of the macro nodes in the final placement.

[0181] As a specific example, the term can be based on the ratio of the total macro overlap area to the total macro area, where the total macro overlap area is the area on the surface of the chip covered by two or more macro nodes, and the total macro area is the area on the surface of the chip covered by at least one macro node. For example, the term can be the negative of the product of the ratio and a weight value, which is a hyperparameter used to control the influence of the overlap weight in the reward function.

[0182] The system trains a node placement neural network (step 412) using reinforcement learning with a reward value for the training netlist in a batch.

[0183] Figure 5 FIG. 500 is a flow chart of an example process 500 for placing macro nodes during training a node placement neural network by reinforcement learning. For convenience, process 500 will be described as being performed by a system of one or more computers located in one or more locations. For example, a placement generation system (e.g., Figure 1 placement generation system 100) can perform process 500.

[0184] The system can execute process 500 at each time step in a sequence of time steps to place a corresponding macro node at that time step until a termination criterion is determined to be satisfied.

[0185] The system generates an input representation from the netlist data that at least characterizes (i) the respective positions on the surface of the chip of any macro nodes in the macro node order that are before a particular macro node to be placed at a given time step, and (ii) the particular macro node to be placed at the given time step (step 502). Optionally, the input representation can also include other information about nodes in the netlist, netlist metadata, or both. An example of an input representation was described above with reference to Figure 2 FIG.

[0186] The system processes the input representation using the node placement neural network (step 504). The node placement neural network is configured to process the input representation according to the current values of the network parameters to generate a score distribution over multiple positions on the surface of the integrated circuit chip.

[0187] The system uses the score distribution to assign the macro node to be placed at a particular time step to a position among the multiple positions on the surface of the chip (step 506). In some implementations, the system uses the score distribution to directly select a position. In some other implementations and as described above, the system can modify the score distribution based on the density of the currently placed nodes (i.e., by setting the score of any position having a density value that meets a threshold to zero), and then select a position from the modified score distribution.

[0188] In some implementations, the system can use additional information to further modify the score distribution.

[0189] Specifically, as described above, in some implementations, the neural network is trained on multiple different placements of multiple different netlists for multiple different chips. This may require the neural network to generate fractional distributions on chip surfaces of different sizes. That is, when the multiple positions are grid squares in an N×M grid covering the surface of an integrated circuit chip, different chips can have different N values and M values. To address this, the system can configure the neural network to generate fractions on a fixed-size maxN × maxM grid. When the N value of the current chip is less than maxN, the system can set the fractions of the additional rows to zero. Similarly, when the M value of the current chip is less than maxM, the system can set the fractions of the additional columns to zero.

[0190] After assigning a particular macro node, the system determines whether assigning the particular macro node to a particular location would cause the placement to enter an infeasible state (step 508). That is, the system determines whether the (partial) placement would enter an infeasible state due to assigning the particular macro node to the particular location.

[0191] The system can make the determination in any of a variety of ways.

[0192] As an example, when the system directly uses the fractional distribution to place macro nodes, the system can determine that assigning a particular macro node to a particular location would cause the placement to enter an infeasible state if, after the macro node is placed at the particular location, the macro node overlaps with another macro node that was placed at an earlier time step.

[0193] As another example, when the system uses a modified fractional distribution to place macro nodes (and thus prevent a particular macro node from overlapping with any other macro node), the system can determine that assigning a particular macro node to a particular location would cause the placement to enter an infeasible state if, after the macro node is placed at the particular location, the next macro node in the macro node sequence cannot be placed without overlapping with another already-placed macro node. That is, the system can determine that there is not enough remaining area on the surface of the chip to accommodate the next macro node without overlapping with another macro node, based on the size of the next macro node.

[0194] In response to determining that assigning a particular macro node to a particular location would cause the placement to enter an infeasible state, the system determines whether a termination criterion is met after placing the particular node (step 510).

[0195] In response to determining that assigning a particular macro node to a particular location would not cause the placement to enter an infeasible state, and if the macro node placed at the time step is not the last macro node in the subset, the system determines that the termination criterion is not met and continues with another iteration of process 500 (step 512).

[0196] If the macro node placed at a time step is the last macro node in the subset, the system can determine that the termination criterion is met regardless of whether assigning a particular macro node to a particular location would cause the placement to enter an infeasible state.

[0197] This specification uses the term "configured" in connection with systems and computer program components. For a system of one or more computers configured to perform particular operations or actions, it means that software, firmware, hardware, or a combination thereof has been installed on the system that, in operation, causes the system to perform the operations or actions. For one or more computer programs configured to perform particular operations or actions, it means that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform the operations or actions.

[0198] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware (including the structures disclosed in this specification and structural equivalents thereof), or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. A computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver apparatus for execution by the data processing apparatus.

[0199] The term "data processing apparatus" refers to data processing hardware and encompasses all kinds of devices, apparatuses, and machines for processing data, including, by way of example, programmable processors, computers, or multiple processors or computers. The apparatus may also be, or further include, special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). In addition to hardware, the apparatus may optionally include code that creates an execution environment for the computer programs, such as code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0200] A computer program (which may also be referred to as or described as a program, software, software application, app, module, software module, script, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The program can, but need not, correspond to a file in a file system. The program can be stored in a part of a file that holds other programs or data (such as one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (such as files that store one or more modules, subroutines, or portions of code). The computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communication network.

[0201] In this specification, the term "database" is used broadly to refer to any collection of data: the data need not be structured in any particular way, or structured at all, and can be stored on a storage device in one or more locations. Thus, for example, an indexed database can include multiple collections of data, each of which can be organized and accessed differently.

[0202] Similarly, in this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Typically, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and run on the same one or more computers.

[0203] The processes and logical flows described in this specification can be performed by one or more programmable computers that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by special-purpose logic circuitry (such as an FPGA or ASIC), or by a combination of special-purpose logic circuitry and one or more programmed computers.

[0204] A computer suitable for executing a computer program can be based on a general or special purpose microprocessor or both, or any other kind of central processing unit. Generally, the central processing unit will receive instructions and data from a read only memory or a random access memory or both. The elements of a computer are the central processing unit for executing or carrying out instructions and one or more memory devices for storing the instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or be operatively coupled to receive data from, or transfer data to, one or more mass storage devices or both. However, a computer need not have such devices. In addition, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive, to name just a few.

[0205] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, by way of example, including semiconductor memory devices, such as, EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks.

[0206] To provide for interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide for interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input received from the user may be in any form, including sound, speech, or tactile input. In addition, a computer may interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a web page in response to a request received from a web browser on the user's device. Further, a computer may interact with the user by sending text messages or other forms of messages to a personal device, such as a smart phone running a messaging application, and receiving responsive messages from the user in return.

[0207] A data processing device for implementing a machine learning model may also include, for example, a dedicated hardware accelerator unit for processing the general and computationally intensive parts of machine learning training or production (i.e., inference, workload).

[0208] A machine learning model can be implemented and deployed using a machine learning framework (e.g., the TensorFlow framework or the Jax framework).

[0209] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a backend component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a frontend component (e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification), or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include a local area network (LAN) and a wide area network (WAN), such as the Internet.

[0210] The computing system can include a client and a server. The client and the server are typically located far apart from each other and typically interact through a communication network. The relationship between the client and the server is generated by computer programs that run on the respective computers and have a client - server relationship with each other. In some embodiments, the server transmits data (e.g., an HTML page) to a user device, e.g., for the purpose of displaying data to a user interacting with the device acting as the client and receiving user input from it. Data generated at the user device, e.g., the result of a user interaction, can be received at the server from the device.

[0211] Although this specification contains many specific implementation details, these details should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Certain features that are described in the context of separate embodiments in this specification can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub - combination in multiple embodiments. Moreover, although features may be described as acting in certain combinations and even initially claimed as such, in some cases one or more features from a claimed combination can be deleted from the combination, and the claimed combination can cover a sub - combination or a variant of the sub - combination.

[0212] Similarly, although the operations are depicted in the drawings and recited in the claims in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in an ordered sequence, or that all of the illustrated operations be performed, to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0213] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the actions recited in the claims may be performed in a different order and still achieve the desired result. As one example, the processes depicted in the figures do not necessarily require the particular order shown or an ordered sequence to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous.

Claims

1. A method performed by one or more computers, the method comprising: Training a node placement neural network through reinforcement learning, wherein the node placement neural network is configured to: receive an input representation at each of a plurality of time steps, the input representation including data representing a current state of placement of a node netlist on a surface of an integrated circuit chip as of the time step, and process the input representation to generate a score distribution over a plurality of locations on the surface of the integrated circuit chip, and wherein the training includes, at each of a plurality of training steps: Obtaining a batch including respective netlist data for each of one or more training netlists, wherein each training netlist corresponds to a respective training integrated circuit chip and specifies connectivity between a plurality of nodes on the corresponding training integrated circuit chip, the plurality of nodes each corresponding to one or more of a plurality of integrated circuit components of the corresponding integrated circuit chip, and wherein the plurality of nodes includes macro nodes representing macro components and standard cell nodes representing standard cell components; For each training netlist, identifying a subset of the macro nodes specified in the training netlist to be placed using the node placement neural network; For each of the training netlists in the training netlists: Generating a partial placement by placing the identified subset of macro nodes in the training netlist in the order of the macro nodes of the training netlist using the node placement neural network until a termination criterion is met; Generating a placement by placing the standard cell nodes in the training netlist and any remaining macro nodes in the training netlist that have not been placed in the partial placement for the training netlist using a default placer; Generating a reward function value of a reward function that measures the quality of the placement; and Training the node placement neural network through reinforcement learning using the reward function value of the training netlist.

2. The method according to claim 1, wherein, Generating a partial placement by placing the identified subset of macro nodes in the training netlist in the order of the macro nodes of the training netlist using the node placement neural network until a termination criterion is met includes, at each of a plurality of specific time steps: Generating an input representation from the netlist data of the training netlist, the input representation characterizing the current state of placement of the training netlist as of the specific time step; Using the node placement neural network to process the input representation to generate a score distribution over a plurality of locations on the surface of the integrated circuit chip; Using the score distribution to assign the macro node to be placed at the specific time step to a specific location among the plurality of locations; And Determining whether the termination criterion is met after assigning the macro node to the specific location.

3. The method according to claim 1 or claim 2, wherein The termination criterion is met when each macro node in the identified subset has been placed using the node placement neural network.

4. The method according to any one of claims 1 to 3, wherein Generating a partial placement by using the node placement neural network to place the macro nodes in the identified subset of the training netlist in corresponding positions according to the macro node order of the training netlist until a termination criterion is met includes: After placing each macro node, determining whether the partial placement will enter an infeasible state, where the infeasible state occurs when two macro nodes overlap on the surface of the training integrated circuit chip; and In response to determining that the partial placement will enter an infeasible state, determining that the termination criterion is met.

5. The method according to claim 4, wherein, After placing each macro node, determining whether the partial placement will enter an infeasible state includes: When the next macro node in the macro node order cannot be placed without overlapping with another already placed macro node in the partial placement, determining that the partial placement will enter the infeasible state.

6. The method according to any one of the preceding claims, wherein, The default placer is an analytical placer.

7. The method according to any one of the preceding claims, wherein, The reward function includes a term for measuring the degree of overlap of the macro nodes in the placement.

8. The method according to any one of the preceding claims, wherein, For each training step and for each training netlist, the identified subset includes all of the macro nodes in the training netlist.

9. The method according to any one of claims 1 to 7, wherein For each training netlist, identifying the subset of the macro nodes in the training netlist that are to be placed using the node placement neural network includes: Determining, according to a schedule, the portion of the macro nodes in the training netlist to be included in the subset, the schedule mapping data characterizing the training step to the portion of the macro nodes in the training netlist of each training netlist in the batch for the training step that is to be included in the subset.

10. The method according to claim 9, wherein, The schedule maps a first training step to a predetermined portion less than all of the macro nodes in the training netlist, and maps later training steps to a portion corresponding to all of the macro nodes in the training netlist.

11. The method according to any one of the preceding claims, wherein, When placing both the macro nodes and the standard cells in a given netlist, the default placer generates a placement that includes an overlap between two or more of the macro nodes.

12. The method according to any one of the preceding claims, further comprising: After training the node placement neural network by reinforcement learning: Receiving new netlist data; Fine-tuning the trained node placement neural network on the new netlist data by reinforcement learning; And Using the fine-tuned node placement neural network to generate an integrated circuit placement for the new netlist data, including placing the corresponding nodes in the new netlist data at each of a plurality of time steps using a score distribution generated by the fine-tuned node placement neural network.

13. The method according to any one of claims 1 to 11, further comprising: After training the node placement neural network by reinforcement learning: Receiving new netlist data; And Using the node placement neural network to generate an integrated circuit placement for the new netlist data, including placing the corresponding nodes in the new netlist data at each of a plurality of time steps using a score distribution generated by the node placement neural network.

14. The method according to any one of the preceding claims, wherein: The node placement neural network includes: an encoder neural network configured to receive the input representation and process the input representation to generate an encoded representation, and a policy neural network configured to process the encoded representation to generate the score distribution.

15. The method according to claim 14, further comprising: pre-training the encoder neural network by supervised learning before training the node placement neural network by reinforcement learning.

16. A system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the corresponding method according to any one of claims 1 to 15.

17. One or more computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the corresponding method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Generating integrated circuit floorplans using neural networks

    US20200175216A1