Three-dimensional fixed contour integrated circuit layout planning method and system based on reinforcement learning

CN116894419BActive Publication Date: 2026-09-29HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310876181.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-18
Publication Date
2026-09-29
Estimated Expiration
2043-07-18

AI Technical Summary

Technical Problem

然而,布图规划问题中的约束和目标需要用近似的可微函数表示,只能生成约束条件松弛后的解并不能生成严格的合法解,且其性能受超参数影响较大,可靠性不如模拟退火算法

Benefits of technology

[0073](1)首次对三维集成电路布图规划问题的局部搜索框架进行强化学习建模,在模块布图阶段提出了一个基于强化学习的搜索框架,通过智能体对解空间的探索,使其获得有效的启发式特性,相比模拟退火算法具有更快的收敛速度,相比分析类方法更具稳定性和可靠性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116894419B_ABST
    Figure CN116894419B_ABST
Patent Text Reader

Abstract

The application discloses a three-dimensional fixed wheel profile integrated circuit layout planning method and system based on reinforcement learning, and the method comprises the following steps: S1, acquiring a module and net list information, and calculating a module area; S2, dividing the circuit according to the module area and the net list information, optimizing the division result, distributing the module to different layers, and determining the number of TSVs of each layer; S3, adopting a sequence pair as a representation method of layout planning, representing a local search process of three-dimensional layout planning of the module through MDP, constructing a module layout planning model based on reinforcement learning, and outputting a layout solution of the module; and S4, performing TSV layout on the basis of the module layout, and outputting a final layout solution. The application can effectively complete three-dimensional integrated circuit layout planning, and obtain a more optimal scheme in terms of line length.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of three-dimensional integrated circuit design and machine learning technology, and particularly relates to a three-dimensional fixed contour integrated circuit layout planning method and system based on reinforcement learning. Background Technology

[0002] In recent years, with the rapid development of very large-scale integrated circuits (VLSI), the transistor integration density in chips has been continuously increasing, and chip performance has been greatly enhanced. However, the process size is now gradually approaching its physical limits, making it difficult to significantly increase transistor integration density by reducing process size, posing a serious challenge to Moore's Law. Three-dimensional integrated circuits are considered one of the most promising solutions for continuing Moore's Law. They stack multiple layers of chips vertically and utilize TSVs (Through-Screen Arrays).

[0003] Through-Silicon Via (TSV) technology enables vertical connections and communication between different chip layers. Compared to two-dimensional integrated circuits, three-dimensional integrated circuits offer advantages such as lower power consumption, higher integration density, and smaller chip size.

[0004] Layout planning is the first and most crucial step in the physical design phase of integrated circuits. In the 3D integrated circuit layout planning stage, the circuit needs to be partitioned, and circuit modules and TSVs are abstracted into geometric shapes and placed on multi-layer chips. The layers and specific locations of the modules and TSVs are determined, and the chip area, interconnect lengths, and the number of TSVs are optimized under the constraint that modules and TSVs do not overlap. The integrated circuit layout planning stage provides feedback data from the earlier planning stage, guiding adjustments to the system architecture. Furthermore, the results of layout planning have a significant impact on subsequent steps such as placement and routing.

[0005] Currently, there are two main types of methods for solving graph planning problems: simulated annealing and analytical methods. Simulated annealing uses graph planning representations to encode circuits, constructs a search space, randomly perturbs the graph solution to generate neighborhood graph solutions, and then determines the search direction through simulated annealing. This type of method is suitable for small-scale circuits, but suffers from slow convergence when dealing with large-scale circuits. Analytical methods express constraints (soft modules, module overlap, fixed profiles, etc.) mathematically, establishing a constrained optimization problem, and optimizing the objective function to obtain the graph solution. This type of method has a short running time and can handle complex constraints. However, the constraints and objective in graph planning problems need to be represented by approximately differentiable functions, and can only generate solutions after constraint relaxation, not strictly valid solutions. Furthermore, its performance is significantly affected by hyperparameters, and its reliability is lower than that of simulated annealing. Reinforcement learning, with its self-learning and generalization capabilities, has been widely applied in combinatorial optimization in recent years. Its agents learn through interaction with the environment, accumulating experience to obtain the maximum reward. Compared to simulated annealing, the reinforcement learning-based search framework offers faster convergence and greater stability and reliability compared to analytical methods. Based on this, this invention, for the first time, models the local search framework for the 3D integrated circuit layout planning problem using reinforcement learning, proposing a two-stage layout planning method. This method sequentially places modules and TSVs (Transformer Components). In the module layout stage, a reinforcement learning-based search framework is used to search for module layout solutions. In the TSV layout stage, based on the module layout solutions, the relative positions of the modules are maintained, and the simulated annealing algorithm is used to complete the TSV layout, outputting the final layout solution. Summary of the Invention

[0006] To overcome the shortcomings of the existing technology, the present invention provides a method and system for layout planning of three-dimensional fixed contour integrated circuits based on reinforcement learning.

[0007] The present invention adopts the following technical solution:

[0008] A reinforcement learning-based method for planning the layout of a three-dimensional fixed-contour integrated circuit includes the following steps:

[0009] S1, Data Preprocessing: Obtain module and netlist information, and calculate module area;

[0010] S2, Layer Assignment: Based on the module area and netlist information, the circuit is divided, and the division result is optimized to assign the module to different layers and determine the number of TSVs (Through-Silicon Vias) in each layer.

[0011] S3, Module Layout: Sequence pairs are used as the representation method for layout planning. The local search process of the module 3D layout planning is represented by MDP (Markov Decision Process). A module layout planning model based on reinforcement learning is constructed, and the layout solution of the module is output.

[0012] S4, TSV Layout: Based on the module layout, perform TSV layout and output the final layout solution.

[0013] Preferably, step S2 specifically includes:

[0014] S21. Based on the module area and netlist information, the circuit is divided into K parts using the hypergraph partitioning tool KaHyPar, where K equals the number of layers in the three-dimensional integrated circuit.

[0015] S22 uses simulated annealing to optimize the partitioning result, balancing the total area of ​​each layer's modules and the TSV. The perturbation performed by the simulated annealing algorithm is to swap modules from two different layers. The objective function is as follows:

[0016]

[0017] Where N TSV Indicates the total number of TSVs, K represents the number of layers, and A represents the number of layers. i This represents the total area of ​​the i-th layer module and the TSV. This represents the average area of ​​each module and TSV.

[0018] Preferably, step S3 specifically includes:

[0019] S31, This stage completes the module layout planning, and the objective function is defined as follows:

[0020]

[0021] This represents a feasible module layout diagram, where η is the weighting coefficient. A(·) is the area function, and W(·) is the line length function; where the area function A(·) is:

[0022] A = E W +E H ·λ+C1·max{E W E H ·λ}+C2·max{W,H·λ}

[0023] W and H represent the width and height of the current layout, respectively, E W =max{W-W0,0},E H =max{H-H0,0}, where W0 and H0 are the width and height of a fixed profile, and E W EH To account for the fixed profile constraint, the portion whose width and height exceed the fixed profile, λ is the ratio of the fixed profile's width to its length, and C1 and C2 are constants, with C1 = 1. The formulas for calculating W0 and H0 are as follows:

[0024]

[0025] A represents the total area of ​​all modules, K represents the number of layers, γ represents the blank area fraction, and λ represents the aspect ratio of the fixed contour. W(·) is the line length function, using the half-perimeter model (HPWL) to evaluate the line length of the module layout:

[0026]

[0027] m represents the total number of networks. Representing network n respectively i The maximum x-coordinate, minimum x-coordinate, maximum y-coordinate, and minimum y-coordinate of the module.

[0028] S32 uses sequence pairs as a representation method for module layout planning. Each layer of modules uses sequence pairs [Γ1,Γ2] to represent the relative positional relationship between modules.

[0029] S33, the local search process of the module 3D layout planning is represented by MDP:

[0030] State space: Each state s is a complete layout diagram. A complete layout diagram includes a set of sequence pairs [Γ1,Γ2] for each layer, as well as the orientation sequence of the modules (horizontal and vertical are represented by 0 and 1).

[0031] Action space: Neighborhood states are generated by defining minor perturbations (all perturbations occur at the same layer). The defined perturbations are as follows:

[0032] • Randomly swap two modules in Γ1.

[0033] • Randomly swap two modules in Γ2.

[0034] • Simultaneously, randomly swap two modules between Γ1 and Γ2.

[0035] • Randomly select a module and place it in any other arbitrary position.

[0036] • Randomly select a module and rotate it 90°.

[0037] A set of neighborhood states is sampled, and the agent chooses to accept moving to one of the states or refuse to move. Therefore, the action space is the number of sampled neighborhood states plus one.

[0038] State transition: Given a layout diagram (state), any of the above perturbations will cause the original state to transition to another definite state.

[0039] Reward: The cost reduction (Δε) at each step is taken as the reward R, and the initial cost is taken as the baseline. The reward R is normalized to [0,1] by dividing the reduced cost by the initial cost.

[0040]

[0041] S34, Extract state features, including the current state cost ε and the cost ε of sampling the neighboring states. ′ The lowest cost ε in this round so far * The current average cost of this round The average cost from the lowest cost in this round to the current state. The lowest cost ε′ in the sampling neighborhood states * The number of modules affected by the disturbance (num) per The area of ​​the module affected by the disturbance per The distance d between the current state and the previous n-step states n Search progress t.

[0042] Cost characteristics: Let the initial cost be ε0, and all cost-related characteristics are normalized to [-1,1] by min(1,ε / ε0-1).

[0043] Layout characteristics: The module affected by the disturbance is defined as the module whose coordinates change before and after the disturbance. The number and area of ​​the modules affected by the disturbance are normalized by dividing by the total number of modules and the total area, and the range is [0,1].

[0044] State distance feature: Calculate the edit distance between the current state sequence pair Γ1, Γ2 and the corresponding sequence pair Γ1′, Γ2′ of the previous states. (The total cost of the edit operations required to convert sequence pair Γ1, Γ2 into sequence pair Γ1′, Γ2′, including insertion, modification, and deletion) and the true matching distance d towards the sequence. ori (The number of times that the elements at corresponding positions in the two sequences are different). d ori This can be achieved by dividing by the total number of modules, num. block Normalized to [0,1]. The distance between the current state (current search step number t) and the states of the previous n steps is:

[0045]

[0046] Search progress feature: The ratio of the current number of search steps to the maximum number of steps per round.

[0047] S35, Construct a modular 3D layout planning model based on DQN. The hidden layers of the DQN neural network are activated by the ReLU function. Features are concatenated into the input, and the dimension of the input is the dimension of the features. The output of the neural network is the value prediction of the action, and the dimension of the output is 1. Define the action value function Q. π (s t ,a t To observe the expected reward that can be achieved after performing action a in a state s under policy π.

[0048]

[0049] γ is the discount rate, t is the time step, and Q is the optimal action value function. * (s,a)=max π Q(s,a) is approximated by a neural network to form the optimal action value function Q(s,a; θ)≈Q. * (s,a), where θ are the parameters of the neural network. The loss function for each iteration i of the Q-network is:

[0050] L(θ i ) = E s,a~ρ(·) [(y i -Q(s,a;θ) 2 ] = E (·) [(r+γmax a′ Q(s′,a′;θ i-1 )-Q(s,a;θ)) 2 ]

[0051] Where ρ(·) is the probability distribution of state s and action a, and the model is updated through the backpropagation algorithm;

[0052] S36, Randomly generate a dataset as the training set.

[0053] S37 performs M rounds of training and outputs a layout diagram of the modules.

[0054] Preferably, step S4 specifically includes:

[0055] S41, the network with modules distributed in different layers is divided into subnets containing TSVs according to the layer. The subnet includes the modules, TSVs in this layer and the TSVs of the previous layer.

[0056] S42: Number the TSVs and randomly insert each TSV from each layer into the module layout sequence pair output in S3 (TSVs are squares with no orientation). Simulated annealing is used to layout the TSVs. The perturbation involves randomly selecting a layer and one TSV from that layer, then randomly placing it in any other position. The perturbation does not change the relative positions of the modules in the sequence pair. The objective function is:

[0057] ε(fp)=A(fp)+ηW(fp)

[0058] fp represents the feasible module and TSV layout diagram, and η represents the weighting coefficients. A(·) is the area function, and W(·) is the line length function; where the area function A(·) is:

[0059] A = E W +E H ·λ+C1·max{E W E H ·λ}+C2·max{W,H·λ}

[0060] W and H represent the width and height of the current module and TSV layout, respectively, E W =max{W-W0,0},E H =max{H-H0,0}, where W0 and H0 are the width and height of a fixed profile, and E W E H To account for the fixed profile constraint, the portion whose width and height exceed the fixed profile, λ is the ratio of the fixed profile's width to its length, and C1 and C2 are constants, with C1 = 1. The formulas for calculating W0 and H0 are as follows:

[0061]

[0062] A represents the total area of ​​all modules and TSVs, K represents the number of layers, γ represents the blank area fraction, and λ represents the aspect ratio of the fixed profile.

[0063] W(·) is the line length function, which uses the half-perimeter length model (HPWL) to evaluate the line length of the module and TSV layout. The line length is the sum of the line lengths of all subnets.

[0064]

[0065] m represents the total number of networks. Representing subnet n respectively i The maximum x-coordinate, minimum x-coordinate, maximum y-coordinate, and minimum y-coordinate of the module, TSV, and TSV of the previous layer at the mapping position in this layer.

[0066] S43 outputs the final layout diagram, including sequence pairs of modules and TSVs in each layer, orientation sequence of modules, and the layer and coordinates of the modules and TSVs.

[0067] This invention also discloses a three-dimensional fixed-contour integrated circuit layout planning system based on reinforcement learning, used to execute the above-described method, which includes the following modules:

[0068] Data preprocessing module: Acquires module and netlist information, and calculates module area;

[0069] Layer allocation module: Based on the module area and netlist information, the circuit is divided, and the division result is optimized to allocate the module to different layers and determine the number of TSVs in each layer.

[0070] Layout Solution Output Module: Using sequence pairs as the representation method for layout planning, the local search process of the module's 3D layout planning is represented by MDP, a module layout planning model based on reinforcement learning is constructed, and the layout solution of the module is output.

[0071] TSV Layout Module: Performs TSV layout based on the module layout and outputs the final layout solution.

[0072] Compared with the prior art, the present invention has the following beneficial effects:

[0073] (1) For the first time, reinforcement learning modeling is used to model the local search framework for the three-dimensional integrated circuit layout planning problem. A reinforcement learning-based search framework is proposed in the module layout stage. Through the exploration of the solution space by the agent, it can obtain effective heuristic characteristics. Compared with the simulated annealing algorithm, it has a faster convergence speed and is more stable and reliable than the analysis method.

[0074] (2) Design state distance features to calculate the differences between different states during the search process, thereby evaluating the size of the search area in the current search stage, avoiding getting trapped in local optima, and guiding the agent to better balance exploration and utilization;

[0075] (3) A two-stage layout planning method is adopted to sequentially lay out the modules and TSVs. Compared with the method of collaborative layout of modules and TSVs, the present invention reduces the solution space and accelerates the convergence of the algorithm. Attached Figure Description

[0076] Figure 1 This is a flowchart of a preferred embodiment of the present invention, which describes a three-dimensional fixed-contour integrated circuit layout planning method based on reinforcement learning.

[0077] Figure 2 This is a layout diagram of the six modules in Example 1.

[0078] Figure 3 This is a block diagram of a three-dimensional fixed contour integrated circuit layout planning system based on reinforcement learning, according to a preferred embodiment of the present invention. Detailed Implementation

[0079] The present invention will now be described in detail with reference to the accompanying drawings and preferred embodiments.

[0080] Example 1

[0081] like Figure 1As shown in the figure, this embodiment presents a three-dimensional fixed contour integrated circuit layout planning method based on reinforcement learning, which includes the following steps:

[0082] S1, Data Preprocessing: Obtain module and netlist information from the n100 dataset in the GSRC benchmark. The n100 dataset contains 100 modules and 551 nets (excluding I / O Pads and the nets where I / O Pads are located). Calculate the module area based on the module's length and width information.

[0083] S2, Layer Allocation: Based on the module area and netlist information, the KaHyPar hypergraph partitioning tool is used to partition the circuit, and the simulated annealing algorithm is used to optimize the partitioning results, assigning the modules to different layers and determining the number of TSVs in each layer.

[0084] Step S2 is as follows:

[0085] S21. Based on the module area and netlist information, the circuit is divided into K parts using the hypergraph partitioning tool KaHyPar, where K equals the number of layers in the three-dimensional integrated circuit.

[0086] S22 uses simulated annealing to optimize the partitioning result, balancing the total area of ​​each layer's modules and the TSV. The perturbation performed by the simulated annealing algorithm is to swap modules from two different layers. The objective function is as follows:

[0087]

[0088] Where, N TSV Indicates the total number of TSVs, K represents the number of layers, and A represents the number of layers. i This represents the total area of ​​the i-th layer module and the TSV. This represents the average area of ​​each module and TSV. In this embodiment, K = 4, meaning the circuit is divided into 4 layers.

[0089] S3, Module Layout: Sequence pairs are used as the representation method for layout planning. The local search process of the module 3D layout planning is represented by MDP. A module layout planning model based on reinforcement learning is constructed, and the layout solution of the module is output.

[0090] Step S3 is as follows:

[0091] S31, This stage completes the module layout planning, and the objective function is defined as follows:

[0092]

[0093] This represents a feasible module layout diagram, where η is the weighting coefficient. A(·) is the area function, and W(·) is the line length function; where the area function A(·) is:

[0094] A = E W +E H ·λ+C1·max{E W E H ·λ}+C2·max{W,H·λ}

[0095] W and H represent the width and height of the current layout, respectively, E W =max{W-W0,0},E H =max{H-H0,0}, where W0 and H0 are the width and height of a fixed profile, and E W E H To account for the fixed profile constraint, the portion whose width and height exceed the fixed profile, λ is the ratio of the fixed profile's width to its length, and C1 and C2 are constants, with C1 = 1. The formulas for calculating W0 and H0 are as follows:

[0096]

[0097] A represents the total area of ​​all modules, K=4 represents the number of layers, γ=20% represents the blank area fraction, and λ=1 represents the aspect ratio of the fixed outline.

[0098] W(·) is the line length function, which uses the half-perimeter length model (HPWL) to evaluate the line length of the module layout:

[0099]

[0100] m represents the total number of networks. Representing network n respectively i The maximum x-coordinate, minimum x-coordinate, maximum y-coordinate, and minimum y-coordinate of the module.

[0101] S32 uses sequence pairs as a representation method for module layout planning. Each layer of modules is represented by sequence pairs [Γ1,Γ2] to indicate the relative positional relationships between modules. For example... Figure 2 The layout of the six modules shown can be represented by [431625,635412]. The information (length, width) of the six modules are 1(4,6), 2(3,7), 3(3,3), 4(2,3), 5(4,3), and 6(6,4).

[0102] S33, the local search process of the module 3D layout planning is represented by MDP:

[0103] State space: Each state s is a complete layout diagram. A complete layout diagram includes a set of sequence pairs [Γ1,Γ2] for each layer, as well as the orientation sequence of the modules (horizontal and vertical are represented by 0 and 1).

[0104] Action space: Neighborhood states are generated by defining minor perturbations (all perturbations occur at the same layer). The defined perturbations are as follows:

[0105] • Randomly swap two modules in Γ1.

[0106] • Randomly swap two modules in Γ2.

[0107] • Simultaneously, randomly swap two modules between Γ1 and Γ2.

[0108] • Randomly select a module and place it in any other arbitrary position.

[0109] • Randomly select a module and rotate it 90°.

[0110] A set of neighborhood states is sampled, and the agent chooses to accept or reject one of the states. Therefore, the action space is the number of sampled neighborhood states plus one.

[0111] State transition: Given a layout diagram (state), any of the above perturbations will cause the original state to transition to another definite state.

[0112] Reward: The cost reduction (Δε) at each step is used as the reward, and the initial cost is used as the baseline. The reward is normalized to [0,1] by dividing the reduced cost by the initial cost.

[0113]

[0114] S34, extract state features, including the current state cost ε, the cost of sampling the neighboring states ε′, and the current lowest cost ε in this round. * The current average cost of this round The average cost from the lowest cost in this round to the current state. The lowest cost ε′ in the sampling neighborhood states * The number of modules affected by the disturbance (num) per The area of ​​the module affected by the disturbance per The distance d between the current state and the previous n-step states n Search progress t.

[0115] Cost characteristics: Let the initial cost be λ0, and all cost-related characteristics are normalized to [-1,1] by min(1,ε / ε0-1).

[0116] Layout characteristics: The module affected by the disturbance is defined as the module whose coordinates change before and after the disturbance. The number and area of ​​the modules affected by the disturbance are normalized by dividing by the total number of modules and the total area, and the range is [0,1].

[0117] State distance feature: Calculate the edit distance between the current state sequence pair Γ1, Γ2 and the corresponding sequence pair Γ1′, Γ2′ of the previous states. (The total cost of the edit operations required to convert sequence pair Γ1, Γ2 into sequence pair Γ1′, Γ2′, including insertion, modification, and deletion) and the true matching distance d towards the sequence. ori (The number of times that the elements at corresponding positions in the two sequences are different). d ori This can be achieved by dividing by the total number of modules, num. block Normalized to [0,1]. The distance between the current state (current search step number t) and the states of the previous n steps is...

[0118]

[0119] Search progress feature: The ratio of the current number of search steps to the maximum number of steps per round.

[0120] The state distance feature in this embodiment includes the distance d between the current state and the state with the lowest cost in this round. ε* The distance d between the current state and the previous 10 states. 10 The distance d between the current state and the states of the previous 100 steps. 100 The distance d between the current state and the states of the previous 1000 steps. 1000 .

[0121]

[0122] S35, Construct a modular 3D layout planning model based on DQN. The hidden layers of the DQN neural network are activated by the ReLU function. Features are concatenated into the input, and the dimension of the input is the dimension of the features. The output of the neural network is the value prediction of the action, and the dimension of the output is 1. Define the action value function Q. π (s t ,a t To observe the expected reward that can be achieved after performing action a in a state s under policy π:

[0123]

[0124] γ is the discount rate, t is the time step, and Q is the optimal action value function. * (s,a)=max π Q(s,a) is approximated by a neural network to form the optimal action value function Q(s,a; θ)≈Q. * (s,a), where θ are the parameters of the neural network. The loss function for each iteration i of the Q-network is:

[0125] L(θ i ) = E s,a~ρ(·) [(yi -Q(s,a;θ) 2 ] = E (·) [(r+γmax a′ Q(s′,a′;θ i-1 )-Q(s,a;θ)) 2 ]

[0126] Where ρ(·) is the probability distribution of state s and action a, and the model is updated through the backpropagation algorithm;

[0127] S36. Randomly generate a dataset of 100 modules and 500 networks as the training set. The length and width of the modules are between 10 and 100 μm. In the network table, 250 networks contain 2 modules and 250 networks contain 3 modules.

[0128] S37 performs M=200 rounds of training and outputs the module layout diagram.

[0129] Step S37 is as follows:

[0130] S371, Initialize the experience pool D with a capacity of N = 20000; randomly initialize the weights θ of the network Q; set the greedy policy ∈ = 0.05; set the termination condition: if no better solution is found in the last w = 1000 steps, then end the current search.

[0131] S372, performs M=200 rounds of iteration, randomly generates the initial state s0 in each round, and executes a maximum of T=50000 steps in each round;

[0132] S373, for the m-th iteration, the state at time t is s t Each step generates a random number u in the range [0,1]; if u≥∈, an action a is randomly selected. t , if u<∈, Reaching the next state s t+1 Storage (s) t ,a t ,r t ,s t+1 ) to experience pool D.

[0133] S374. If the termination condition is met or a T-step action has been performed, end the current search round; otherwise, continue the search.

[0134] S375, if the storage capacity n>N, select a set (s) from the experience pool D. t ,a t ,r t ,s t+1 If s t+1 If y is the final state, then t =r t If st+1 If it is not the final state, then Calculate the loss L and update θ using the backpropagation algorithm.

[0135] S4, TSV Layout: Using the simulated annealing algorithm, a TSV layout is performed based on the module layout, and the final layout solution is output.

[0136] Step S4 is as follows:

[0137] S41, TSV adopts a via-first encapsulation method. TSVs are distributed in layers 1, 2, and 3. The network with modules distributed in different layers is divided into subnets containing TSVs according to the layer. The subnet includes the modules, TSVs in this layer, and TSVs in the layer above.

[0138] S42, TSVs are squares with sides of 3μm. The TSVs are numbered, and each TSV from each layer is randomly inserted into the module layout sequence output from S3 (TSVs are squares, with no orientation). Simulated annealing is used to layout the TSVs. The perturbation involves randomly selecting a layer and one TSV from that layer, then randomly placing it in any other position. The perturbation does not change the relative positions of the modules in the sequence pair. The objective function is:

[0139] ε(fp)=A(fp)+ηW(fp)

[0140] fp represents the feasible module and TSV layout diagram, and η represents the weighting coefficients. A(·) is the area function, and W(·) is the line length function; where the area function A(·) is:

[0141] A = E W +E H ·λ+C1·max{E W E H ·λ}+C2·max{W,H·λ}

[0142] W and H represent the width and height of the current module and TSV layout, respectively, E W =max{W-W0,0},E H =max{H-H0,0}, where W0 and H0 are the width and height of a fixed profile, and E W E H To account for the fixed profile constraint, the portion whose width and height exceed the fixed profile, λ is the ratio of the fixed profile's width to its length, and C1 and C2 are constants, with C1 = 1. The formulas for calculating W0 and H0 are as follows:

[0143]

[0144] A represents the total area of ​​all modules and TSVs, K = 4, γ = 20% represents the blank area fraction, and λ = 1 represents the aspect ratio of the fixed profile.

[0145] W(·) is the line length function, which uses the half-perimeter length model (HPWL) to evaluate the line length of the module and TSV layout. The line length is the sum of the line lengths of all subnets.

[0146]

[0147] m represents the total number of networks. Representing subnet n respectively i The maximum x-coordinate, minimum x-coordinate, maximum y-coordinate, and minimum y-coordinate of the module, TSV, and TSV of the previous layer at the mapping position in this layer.

[0148] S43 outputs the final layout diagram, including sequence pairs of modules and TSVs in each layer, orientation sequence of modules, and the layer and coordinates of the modules and TSVs.

[0149] Example 2

[0150] like Figure 3 As shown in the figure, this embodiment of a reinforcement learning-based three-dimensional fixed contour integrated circuit layout planning system includes the following modules:

[0151] Data preprocessing module: Acquires module and netlist information, and calculates module area;

[0152] Layer allocation module: Based on the module area and netlist information, the circuit is divided, and the division result is optimized to allocate the module to different layers and determine the number of TSVs in each layer.

[0153] Layout Solution Output Module: Using sequence pairs as the representation method for layout planning, the local search process of the module's 3D layout planning is represented by MDP, a module layout planning model based on reinforcement learning is constructed, and the layout solution of the module is output.

[0154] TSV Layout Module: Performs TSV layout based on the module layout and outputs the final layout solution.

[0155] Other aspects of this embodiment can be found in Embodiment 1.

[0156] In summary, this invention belongs to the field of 3D integrated circuit design and machine learning technology, specifically relating to a reinforcement learning-based method and system for 3D fixed-contour integrated circuit layout planning. The method includes the following steps: S1, Data preprocessing: acquiring module and netlist information, and calculating module area; S2, Layer allocation: based on module area and netlist information, using the KaHyPar hypergraph partitioning tool to partition the circuit, and using simulated annealing algorithm to optimize the partitioning results, allocating modules to different layers and determining the number of TSVs (Through-Silicon Vias) in each layer; S3, Module layout: using sequence pairs as the representation method for layout planning, representing the local search process of 3D module layout planning through MDP (Markov Decision Process), constructing a reinforcement learning-based module layout planning model, and outputting the module layout solution; S4, TSV layout: using simulated annealing algorithm, performing TSV layout based on the module layout, and outputting the final layout solution. This invention can effectively complete 3D integrated circuit layout planning and obtain a scheme with better line length.

[0157] The above description is merely a detailed explanation of preferred embodiments and principles of the present invention. For those skilled in the art, there may be changes in specific implementation methods based on the ideas provided by the present invention, and these changes should also be considered within the scope of protection of the present invention.

Claims

1. A three-dimensional fixed-contour integrated circuit layout planning method based on reinforcement learning, characterized in that, Includes the following steps: S1: Obtain module and netlist information, and calculate module area; S2, based on the module area and netlist information, divide the circuit, optimize the division results, allocate the modules to different layers and determine the number of TSVs in each layer; S3 uses sequence pairs as the representation method for graph planning, and represents the local search process of module 3D graph planning through MDP, constructs a module graph planning model based on reinforcement learning, and outputs the graph solution of the module. Step S3 is as follows: S31, Complete the module layout planning. The objective function is defined as follows: This represents a feasible module layout diagram. These are the weighting coefficients; It is an area function. Let be a function of line length; where is a function of area. : and These represent the width and height of the current layout, respectively. , , , To fix the width and height of the outline, , To account for fixed contour constraints, the portion whose width and height exceed the fixed contour, To fix the ratio of the width to the length of the profile, , It is a constant. , , The calculation formula is as follows: This represents the total area of ​​all modules. Indicates the number of floors. Represents the fraction of the blank area. The aspect ratio of a fixed profile; As a line length function, the half-perimeter model is used to evaluate the line length of the module layout: Indicates the total number of networks. , , , They represent networks respectively. The largest of the middle modules Coordinates, minimum Coordinates, Maximum Coordinates, minimum coordinate; S32 uses sequence pairs as a representation method for module layout planning. Each layer of modules is represented by sequence pairs. The term "pair" indicates the relative positional relationship between modules; S33 represents the local search process of the module 3D layout planning using MDP; S34, Extract state features, including the current state cost. The cost of sampling neighborhood states The lowest cost in this round so far The current average cost of this round The average cost from the lowest cost in this round to the current state. The lowest cost in sampling neighborhood states The number of modules affected by the disturbance Module area affected by disturbance The current state is different from the previous state. Distance of step state Search progress ; S35, Construct a modular 3D layout planning model based on DQN. The hidden layers of the DQN neural network are activated by the ReLU function. Features are concatenated into the input, and the dimension of the input is the dimension of the features. The output of the neural network is the value prediction of the action, and the dimension of the output is 1. Define the action value function. To observe in the strategy Next state Execute action Expected rewards to be achieved in the future: For discount rate, For each time step, the optimal action value function Using neural networks to approximate the optimal action value function ,in These are the parameters of the neural network; Each iteration of the network Loss function: in It's about the state. and actions The probability distribution is used to update the model through the backpropagation algorithm; S36, Randomly generate a dataset as the training set; S37, proceed Round training, outputting a layout diagram of the modules; S4, based on the module layout, performs TSV layout and outputs the final layout solution.

2. The method for planning the layout of a three-dimensional fixed contour integrated circuit based on reinforcement learning according to claim 1, characterized in that: Step S2 is as follows: S21, based on the module area and netlist information, the circuit is partitioned using the hypergraph partitioning tool KaHyPar. Part of, among which Equal to the number of layers in a three-dimensional integrated circuit; S22, the simulated annealing algorithm is used to optimize the partitioning result, balancing the total area of ​​each layer of modules and the TSV; the perturbation performed by the simulated annealing algorithm is to swap modules from two different layers, and the objective function is as follows: in, This represents the total number of TSVs. Indicates the number of floors. Indicates the first The total area of ​​the layer module and TSV. This represents the average area of ​​each module and TSV.

3. The method for planning the layout of a three-dimensional fixed contour integrated circuit based on reinforcement learning according to claim 1, characterized in that: In step S33, the local search process of the module 3D layout planning is represented by MDP as follows: State space: each state Each is a complete layout diagram; a complete layout diagram includes a set of sequence pairs for each layer. And the orientation sequence of the modules, with horizontal and vertical directions represented by 0 and 1; Action space: The neighborhood state is generated by defining a perturbation as follows: Random in The two modules are switched in the middle; Random in The two modules are switched in the middle; At the same time, randomly in The two modules are switched in the middle; Randomly select a module and place it in any other arbitrary position; Randomly select a module and rotate it 90°; Sample a set of neighborhood states, and the agent chooses to accept or reject one of the states. The action space is the number of sampled neighborhood states plus one. State transition: Given a layout diagram, any disturbance will cause the original state to transition to another definite state; Reward: Reduce the cost of each step As a reward and using the initial cost as a benchmark. The reward will be given by dividing the reduced cost by the initial cost. Normalization to , 。 4. The method for planning the layout of a three-dimensional fixed contour integrated circuit based on reinforcement learning according to claim 3, characterized in that: In step S34: Cost characteristics: Let the initial cost be... All cost-related features are through Normalization to Layout characteristics: Modules affected by disturbances are defined as those whose coordinates change before and after the disturbance. The number and area of ​​modules affected by disturbances are normalized by dividing by the total number of modules and the total area, respectively, within a certain range. ; State distance feature: Calculate the current state sequence pairs , Sequence pairs corresponding to previous states , Edit distance , True matching distance with orientation sequence The number of times corresponding elements in two sequences are different; sequence pair , Convert to sequence pairs , The total cost of all required editing operations, including insertion, modification, and deletion; , , It can be done by dividing by the total number of modules. Normalization to The current search steps are: The current state is different from the previous state. The distance between the states is: Search progress feature: The ratio of the current number of search steps to the maximum number of steps per round.

5. The method for planning the layout of a three-dimensional fixed contour integrated circuit based on reinforcement learning according to claim 4, characterized in that: Step S4 is as follows: S41, the network with modules distributed in different layers is divided into subnets containing TSVs according to the layers. The subnets include the modules, TSVs in this layer and the TSVs of the previous layer. S42, number the TSVs, and randomly insert each TSV from each layer into the module layout sequence pair output in S3. Simulated annealing is used to layout the TSVs. The perturbation involves randomly selecting a layer and a TSV from that layer, then randomly placing it in any other position. The perturbation does not change the relative positions of the modules in the sequence pair. The objective function is: A feasible module and TSV layout diagram. These are the weighting coefficients; It is an area function. Let be a function of line length; where is a function of area. : and These represent the width and height of the current module and the TSV layout, respectively. , , , To fix the width and height of the outline, , To account for fixed contour constraints, the portion whose width and height exceed the fixed contour, To fix the ratio of the width to the length of the profile, , It is a constant. , , The calculation formula is as follows: This represents the total area of ​​all modules and TSVs. Indicates the number of floors. Represents the fraction of the blank area. The aspect ratio of a fixed profile; The line length function uses a half-perimeter model to evaluate the line length of modules and TSV layouts. The line length is the sum of the line lengths of all subnets. Indicates the total number of networks. , , , Representing subnets The maximum mapping position of the module, TSV, and TSV of the previous layer in this layer. Coordinates, minimum Coordinates, Maximum Coordinates, minimum coordinate; S43 outputs the final layout diagram, including sequence pairs of modules and TSVs at each layer, orientation sequence of modules, and the layer and coordinates of the modules and TSVs.

6. A three-dimensional fixed-contour integrated circuit layout planning system based on reinforcement learning, used to execute the method according to any one of claims 1-5, characterized in that, The system includes the following modules: Data preprocessing module: Acquires module and netlist information, and calculates module area; Layer allocation module: Based on the module area and netlist information, the circuit is divided, and the division results are optimized to allocate the modules to different layers and determine the number of TSVs in each layer. Layout Solution Output Module: Using sequence pairs as the representation method for layout planning, the local search process of the module's 3D layout planning is represented by MDP, a module layout planning model based on reinforcement learning is constructed, and the layout solution of the module is output. TSV Layout Module: Performs TSV layout based on the module layout and outputs the final layout solution.