PCB dummy pad intelligent layout optimization method based on deep reinforcement learning
Through the intelligent layout optimization method based on deep reinforcement learning, the problem of inefficient optimization of PCB dummy pad layout is solved, and the copper clad uniformity and production cost are improved in the electroplating process.
Patent Information
- Application Number
- CN202510046898.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art relies on manual experience in the optimization of PCB dummy pad layout, which is inefficient and costly, and cannot meet the needs of large-scale production, resulting in poor uniformity of electroplating copper thickness, affecting the performance and reliability of PCB.
Using an intelligent layout optimization method based on deep reinforcement learning, we decompose the whole PCB into a daughter board, extract joint features, build an adaptive scale feature module and an instant reward network, optimize the layout of the dummy pads, and improve the copper clad uniformity of the electroplating process.
It improves the efficiency of PCB design optimization, improves the uniformity of copper cladding in the electroplating process, reduces the cost of PCB production, shortens the R&D cycle, and achieves a more efficient production process.
Smart Images

Figure CN119990047A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of PCB electroplating copper thickness optimization, and in particular to the technical field of PCB dummy pad layout optimization based on deep reinforcement learning. Background Art
[0002] With the rapid development of the electronics industry, the application of printed circuit boards (PCBs) in electronic products is becoming more and more extensive, and the requirements for the uniformity of the thickness of the entire PCB electroplating copper are also increasing. The uniformity of the thickness of the entire PCB electroplating copper directly affects the performance and reliability of the PCB, such as conductivity, corrosion resistance, mechanical strength, etc. The traditional electroplating copper uniformity optimization method mainly relies on the experience and experiments of engineers, which is inefficient and costly and cannot meet the needs of large-scale production. An industrial-grade PCB electroplating layout usually takes several months of manual time from design to optimization, and PCB layout automation can greatly improve the efficiency of industrial production.
[0003] Reasonable placement of dummy pads in the idle area of the PCB can effectively improve the unevenness of current density during the electroplating process, thereby ensuring uniform thickness of the electroplated copper layer of the entire PCB board. In fact, dummy pads do not assume actual circuit functions, but play a role in balancing current distribution during the electroplating process. Before production, engineers pre-allocate dummy pad resources so that they can effectively guide current without interfering with actual circuit functions. However, PCB dummy pad layout relies heavily on the experience of engineers and lacks mature automation solutions, which severely limits the efficiency and consistency of design.
[0004] like Figure 2As shown in the figure, the PCB electroplating tank includes a cathode, an anode and an electrolytic tank filled with electroplating solution. The complete PCB contains through-holes, sub-board boundaries and dummy pads. In the electroplating tank, the cathode is placed on the PCB to be plated, and the metal copper ions are transferred from the electrolyte to the inner wall of the through-hole of the PCB to be plated by applying current in the electrolyte. Although the direction of the current is fixed, due to the uneven distribution of the current, the deposition thickness of the copper ions in the through-hole varies. It is necessary to accurately arrange dummy pads in the idle area of the PCB to improve and balance the local current density. Due to the large number of optional dummy pad placement positions in the PCB, the search space of the entire PCB is very large, and it is difficult for manual methods to efficiently find a suitable solution. In recent years, Deep Reinforcement Learning (DRL) has gradually become a technology that has attracted much attention due to its advantages in dynamic optimization tasks. DRL combines the feature generalization ability of deep neural networks and the optimization ability of reinforcement learning, and can handle complex nonlinear problems. The complexity of PCB design stems from its scale (a full-page PCB contains hundreds of sub-boards), the layout accuracy of the pads, and the high computational cost (multi-physics simulation takes hours, sometimes even more than a day). DRL needs to be reward-driven to find the optimal layout strategy through trial and error. Finite Element Analysis (FEA) can provide high-precision results in predicting the thickness of electroplated copper. However, FEA requires a large amount of meshing, complex multi-physics coupling calculations, and nonlinear solutions, which significantly increase the computational complexity. Long-term dynamic simulation and high-precision calculations further increase the consumption of computing resources and time. These all limit the application capabilities of FEA in real-time design optimization. The existing technology still requires a lot of manpower and material resources to optimize the layout of PCB virtual pads, making the production cost of PCB high and the design cycle long. Summary of the invention
[0005] In view of the above problems, the purpose of the present invention is to propose a PCB virtual pad intelligent layout optimization method based on deep reinforcement learning. It effectively solves the problems of lengthy calculation time and low design efficiency of multi-physics field simulation, improves the PCB design optimization efficiency and copper coating uniformity of the electroplating process, reduces PCB production costs, and shortens the R&D cycle.
[0006] The method comprises the following steps:
[0007] S1. Decompose the entire PCB and generate several PCB sub-board layouts;
[0008] S2, extracting joint features of PCB sub-boards;
[0009] S3, construct an adaptive scale feature module to extract the joint feature vector of each sub-board;
[0010] S4. Build and train the immediate reward network;
[0011] S5. Construct and optimize the virtual pad layout optimization model;
[0012] S6, calculating an immediate reward signal;
[0013] S7, confirming the placement position of the dummy pad;
[0014] S8. Optimize PCB dummy pad layout based on deep reinforcement learning loop.
[0015] Furthermore, the decomposition of the entire PCB to generate a plurality of PCB sub-board layouts is specifically as follows: according to the boundaries of the PCB sub-boards, the entire PCB is decomposed into a plurality of PCB sub-boards, and according to the actual physical size of the PCB sub-boards and the distribution of through holes, the PCB sub-boards are divided into a plurality of grid units of fixed sizes.
[0016] Furthermore, the joint feature includes a spatial feature M and a structural feature C; the spatial feature The 3 is the number of channels, the H and W are the height and width of the PCB sub-board layout, respectively. Represents the real number domain; the structural features The T is the maximum number of dummy pads arranged in the daughter board.
[0017] Furthermore, the adaptive scale feature module passes through the spatial branch and the structural branch from the input layer to the output layer and then passes through the aggregation module; the spatial branch passes through the pyramid pooling module and two linear transformation layers from the input layer to the output layer; the structural branch passes through the external attention mechanism module and two linear transformation layers from the input layer to the output layer;
[0018] The pyramid pooling module includes a convolution layer and a maximum pooling layer from the input layer to the output layer; the external attention mechanism module includes a linear layer and a normalization layer from the input layer to the output layer; the extraction of the joint feature vector of each sub-board is specifically as follows: the joint feature is input into the adaptive scale feature module to obtain the joint feature vector h l .
[0019] Furthermore, the instant reward network is specifically established by introducing a multi-layer perceptron based on the adaptive scale feature module, wherein the multi-layer perceptron includes two linear transformation layers; the calculation formula of the instant reward network is: out =MLP(h l ); said h out It represents the copper thickness result predicted by the instant reward network, and the MLP() represents the multi-layer perceptron function.
[0020] Furthermore, the virtual pad layout optimization model includes two branches with the same structure. The construction of each branch is as follows: a multi-layer perceptron is introduced on the basis of the adaptive scale feature module; the branches are respectively set as a strategy model and a value model; the calculation formula of the strategy model is: The softmax() represents a normalization function, the Represents the probability distribution matrix of the action space, is the current state S t The calculation formula of the value model is: Said Represents the expected value of the discounted cumulative reward.
[0021] Furthermore, the dummy pad layout optimization model uses a proximal strategy optimization algorithm to optimize the strategy model and the value model; a clipping loss function is used when optimizing the strategy model, specifically by: To achieve the L PF (θ) represents the optimized strategy model, The parameter difference between the policy model and the optimized policy model is limited to (1-∈, 1+ε), the value of ε is 0.2, θ represents the trainable parameter of the policy model, and R t (θ) is the probability ratio between the policy model and the optimized policy model, and the A t represents the advantage function, the A t =r t -V t , the r t Represents the immediate reward signal; when optimizing the value model, the clipping loss function is defined by the mean square error, specifically by: To achieve represents the trainable parameters of the value model, Represents the optimized value model.
[0022] Furthermore, the instant reward signal r t The calculation formula is: t =|f ReNet (s t )-τ|-|f ReNet (s t+1 )-τ|, wherein τ represents the target thickness value of the copper plating of the entire PCB; ReNet () indicates an instant reward network.
[0023] Further, the determination of the placement position of the dummy pad is specifically as follows: in each time step, the dummy pad layout optimization module selects the probability distribution matrix of the action space by sampling to obtain the dummy pad placement position at ; Place the dummy pad at position a t After that, the layout state changes from the current state S t Transition to the next state S t+1 .
[0024] Further, the optimization of the PCB virtual pad layout based on the reinforcement learning cycle is specifically as follows: the virtual pad layout optimization module transmits the current state S t and the next state S t+1 To the immediate reward module, the immediate reward module transmits the immediate reward signal r t to a value model, which is based on the immediate reward signal r t Calculate the expected value of discounted cumulative reward V t And update the parameters of the strategy model. When t=T, the optimization iteration ends and the optimized PCB virtual pad layout is obtained.
[0025] The beneficial effects of the method of the present invention are:
[0026] During the PCB electroplating process, the spatial distribution of current density and the specific layout of through holes on the sub-board will significantly affect the thickness distribution of the copper layer. The electroplating process involves multi-physical field coupling. Only inputting the local information of each sub-board into the prediction model separately cannot fully reflect the complexity and global correlation of the overall electroplating process. To solve this problem, the present invention proposes to use spatial features and structural features as the joint feature representation of PCB sub-boards.
[0027] The present invention overcomes the limitations of the prior art, decomposes the entire PCB into multiple sub-board areas to reduce the complexity of the optimization problem, integrates the Markov decision process and the method based on reinforcement learning, confirms the placement position of the dummy pad through a PCB dummy pad layout optimization model based on deep reinforcement learning, and optimizes and adjusts the layout of the dummy pad using the predicted electroplating copper thickness, thereby improving the PCB design optimization efficiency and the copper coating uniformity of the electroplating process, reducing the PCB production cost, shortening the research and development cycle, and showing a wide range of use value and application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 This is a main program flow chart in an embodiment of the present invention;
[0029] Figure 2 This is a schematic diagram of an electrolyte model in an embodiment of the present invention;
[0030] Figure 3 This is a schematic diagram of a PCB layout in an embodiment of the present invention;
[0031] Figure 4(a) is a schematic diagram of the training loss of the instant reward network training set according to an embodiment of the present invention, where the horizontal axis is the number of training times and the vertical axis is the loss function value;
[0032] (b) is a schematic diagram of the accuracy of the instant reward network of an embodiment of the present invention at a precision threshold of 0.5 μm, where the horizontal axis is the number of training times and the vertical axis is the prediction accuracy;
[0033] (c) is a schematic diagram of the accuracy of the instant reward network of an embodiment of the present invention at a 1 μm accuracy threshold, where the horizontal axis is the number of training times and the vertical axis is the prediction accuracy;
[0034] Figure 5 Schematic diagram showing the comparison of uniform thickness results of electroplated copper based on the placement strategy of the dummy pad layout optimization model and the traditional constraint-based random optimization model in the embodiment of the present invention. The red dots represent the predicted results of electroplated copper thickness in different sub-boards under the placement strategy of the dummy pad layout optimization model, and the blue dots represent the predicted results of the constraint-based random optimization model. The horizontal axis is the PCB sub-board number, and the vertical axis is the thickness of the PCB sub-board copper. DETAILED DESCRIPTION
[0035] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0036] This embodiment provides a PCB dummy pad intelligent layout optimization method based on deep reinforcement learning. The flow chart of the method is as follows: Figure 1 , the method comprises the following steps:
[0037] S1. Decompose the entire PCB and generate several PCB sub-board layouts.
[0038] The relevant operations of step S1 are introduced with a specific example:
[0039] The whole PCB contains hundreds of sub-boards. According to the boundaries of the PCB sub-boards, the whole PCB is decomposed into several PCB sub-boards. Dozens of virtual pads need to be accurately placed in the layout of each sub-board to reduce the unevenness of the copper layer thickness during the electroplating process. According to the actual physical size of the PCB sub-board and the distribution of through holes, a gridding strategy is adopted to divide the two-dimensional canvas into multiple fixed-size grid units. The gridding strategy is specifically: converting the vector map into a pixel map. The number of rows and columns of the grid unit division has a direct impact on the computational complexity of the optimization problem and the accuracy and quality of the layout. Therefore, the maximum number of rows and columns of the grid unit is limited to 128, which can strike a balance between computational efficiency and optimization effect. In this embodiment, the grid division problem is modeled as a knapsack optimization problem. Based on the distribution characteristics of components in each sub-board and the influence of different row and column combinations on space utilization, the advantages and disadvantages of each division scheme are calculated and sorted, and the optimal grid division strategy is selected to support the subsequent optimization process. After backpack optimization, this embodiment adopts a grid division scheme of 100 rows and 107 columns to ensure the efficiency and quality of the layout. Therefore, it can be concluded that the size of the grid cells generated for each PCB daughter board depends on the actual physical size of the daughter board and the distribution of through holes.
[0040] S2. Extract joint features of PCB sub-boards.
[0041] The relevant operations of step S2 are introduced with a specific example:
[0042] The joint features include spatial features M and structural features C.
[0043] like Figure 3 As shown, the spatial feature M represents a multi-channel binary image matrix of the PCB sub-board layout, The 3 represents the number of channels, including the plated through hole layer, the boundary layer and the dummy pad layer. Each feature occupies a channel separately. The H and W represent the height and width of the PCB sub-board layout respectively. Represents the field of real numbers.
[0044] The structural feature C represents a sequence consisting of a set of vectors, each vector describing key information of a dummy pad. T is the maximum number of dummy pads arranged in the daughter board. In the set of vectors, each vector C t =(x t ,y t ,b), where t represents the time step, t∈{1,2,...,T}, and x t and t is the physical coordinate of the virtual pad in the whole PCB, and b∈{0,1} is a Boolean type. When b=0, b represents a linear virtual pad, and when b=1, b represents an arc virtual pad.
[0045] S3. Encode the sub-boards through the adaptive scale feature module and extract the joint feature vector of each sub-board.
[0046] The relevant operations of step S3 are introduced with a specific example:
[0047] The adaptive scale feature module passes through the spatial branch and the structural branch from the input layer to the output layer and then passes through the aggregation module; the spatial branch passes through the pyramid pooling module and two linear transformation layers from the input layer to the output layer; the structural branch passes through the external attention mechanism module and two linear transformation layers from the input layer to the output layer; the pyramid pooling module includes a convolution layer and a maximum pooling layer from the input layer to the output layer, the convolution layer includes a 3×3 convolution layer, a 5×5 convolution layer and an 8×8 convolution layer arranged in parallel, and the maximum pooling layer includes a 1×1 maximum pooling layer, a 2×2 maximum pooling layer and a 4×4 maximum pooling layer arranged in parallel; the external attention mechanism module includes a linear layer (Linear) and a normalization layer (Norm) from the input layer to the output layer, and a memory unit E is added between the linear layer (Linear) and the normalization layer (Norm) k , add memory unit E after the normalization layer (Norm) v .
[0048] The spatial branch uses spatial pyramid pooling to embed features of the multi-channel binary image of the PCB sub-board layout, extracts the spatial features of the PCB sub-board through the convolution Conv module of 3×3 convolution layer, 5×5 convolution layer and 8×8 convolution layer, and then uses the pyramid pooling SPPNet operation to convert spatial features of any size into hidden feature vectors of fixed length through 1×1 maximum pooling layer, 2×2 maximum pooling layer and 4×4 maximum pooling layer, integrates local geometric information, and obtains the geometric information h of the PCB sub-board. m , the h m =SPPNet(CNN(M)), The SPPNet() represents a pyramid pooling function, the CNN represents a convolutional layer function, and d represents a dimension. In the pyramid pooling operation, the original spatial features are divided into multiple scale regions for pooling, and finally uniformly converted into feature vectors with a dimension of d=128, thereby ensuring the consistency of feature representation of inputs of different sizes.
[0049] The structure branch uses an external attention mechanism to extract structural features and weightedly integrates the attention weights between the structural features and the external memory unit to extract global information to obtain the layout features h of the PCB sub-board. c , the h c =AE v , The (α) i,j Represents the i-th element in the structural feature and the memory unit E k Similarity weight between rows j; memory unit and It is a learnable parameter independent of the structural features and belongs to the linear transformation layer. It is used to store the global semantic information of the entire data set. Its initialization and update are achieved through one-dimensional convolution operations. The U represents the hyperparameter and the Norm() represents the normalization function. The memory unit is generated by two stacked one-dimensional convolution layers, in which the weight parameters of the one-dimensional convolution layer are responsible for encoding and optimizing the global information.
[0050] The geometric information h m and layout features h c After two linear layers, the geometric information h is concatted by the aggregation module m and layout features h c Splice to form a joint feature vector h l , The h l The calculation formula is: l =concat(linear(h m ),linear(h c )), the linear() represents the linear layer function.
[0051] S4. Build and train an immediate reward network.
[0052] The relevant operations of step S4 are introduced with a specific example:
[0053] The immediate reward network ReNet aims to learn a mapping function f ReNet (S t )=h out , predict the copper thickness distribution generated by different PCB sub-board layouts during the electroplating process. The instant reward network introduces a multi-layer perceptron MLP containing two linear layers based on the adaptive scale feature module to perform nonlinear transformation and inference prediction;
[0054] The calculation formula of the instant reward network is: out =MLP(h l ); said h out represents the copper thickness result predicted by the instant reward network, which is transmitted as a reward signal to the subsequent dummy pad layout optimization model optimized by deep reinforcement learning (DRL) for dynamically adjusting the decision strategy and optimizing the PCB design layout to achieve uniformity of copper thickness and improvement of design performance; the MLP() represents a multi-layer perceptron function;
[0055] The instant reward network randomly selects a full-page PCB as a test set, and the remaining data is used as a training set for training. The instant reward network uses the root mean square error as the loss function. The loss function of the instant reward network is expressed as: The n is the number of training set samples, and the h true is the true value of the training set sample, represents the trainable parameters of the immediate reward network, the |||| 2 Represents the square of the norm.
[0056] like Figure 4 As shown in (a), the training loss of the training set gradually decreases with the increase of training times. In order to further evaluate the prediction performance of the immediate reward network, an intuitive evaluation criterion is introduced: Figure 4 As shown in (b), 0.5 μm is used as the accuracy threshold to measure the acceptability of the prediction results. The final prediction accuracy on the test set reaches 97.4%, showing the high accuracy of the model in practical applications. Figure 4 (c) shows that when using 1μm as a relatively loose accuracy standard, the fully trained immediate reward network achieves 100% accuracy on the test set.
[0057] The training process uses small-batch stochastic gradient descent SGD optimization, and the model performance is evaluated through 5-fold cross-validation. The fluctuation of the loss function value gradually decreases and eventually stabilizes, indicating that the training is completed.
[0058] S5. Build and optimize the dummy pad layout optimization model optimized by deep reinforcement learning;
[0059] The relevant operations of step S5 are introduced with a specific example:
[0060] In the dummy pad layout optimization module, a deep reinforcement learning optimized dummy pad layout optimization model ReNet is constructed. The dummy pad layout optimization model includes two branches with the same structure, and each branch is constructed as follows: a multi-layer perceptron is introduced on the basis of the adaptive scale feature module; and the branches are respectively set as a strategy model and a value model.
[0061] The strategy model uses a multi-layer perceptron to l Perform nonlinear transformation and use the normalized exponential function softmax() to output the probability distribution of the action space. The specific formula is: P a =softmax(MLP(h l )), Represents the probability distribution matrix of the action space, is the current state S tThe output of the value model is a scalar value, which is used to estimate the value of the current state. The specific formula is: Said Represents the expected value of the discounted cumulative reward, which represents the placement position a taken by the virtual pad layout optimization model according to the strategy model in the current state t The weighted sum of future cumulative rewards that can be obtained. This value will be used as feedback to update the parameters of the policy model to improve the performance of the model in subsequent decisions.
[0062] The current state t =(M t ,C t ) is defined as a tuple consisting of the spatial features M of the PCB sub-board at the current time step t and structural features C t The action space corresponds to the two-dimensional grid of the sub-board, with a size of H×W. Each point in the grid represents a potential action position, that is, the placement position of the dummy pad. Due to the large action space, according to the empirical constraints, the placement area of the dummy pad is limited to the range of 6 grids outside the boundary of the sub-board, which effectively reduces the size of the action space and reduces the complexity of the optimization problem. The policy model generates a corresponding probability distribution matrix of the action space for each possible placement position. represents the probability of choosing different actions at each time step t.
[0063] The virtual pad layout model uses the proximal policy optimization algorithm PPO to optimize the trainable parameters of the policy model and the value model; the clipping loss function is used when optimizing the policy model, specifically through: The L PF (θ) represents the optimized strategy model, The parameter difference between the strategy model and the optimized strategy model is limited to (1-∈, 1+∈), the value of ∈ is 0.2, and the A t =r t -V t , the A t represents the advantage function, the r t represents the immediate reward function, θ represents the trainable parameters of the policy model, and R t (θ) is the probability ratio between the old policy and the current policy; the clipping loss function of the value model is defined by the mean square error: Said Represents the trainable parameters of the value model.
[0064] S6. Calculate the immediate reward signal.
[0065] The relevant operations of step S6 are introduced with a specific example:
[0066] Instant Reward Signal t It is estimated by calculating the discounted cumulative reward, which is based on the delayed reward at the final time step and weighted by a predetermined discount factor for future rewards. The specific calculation formula is: t =γ T-t r T , the r T is the final reward at the end of the round, and γ is the discount factor. If rewards are only given after time step t, it may lead to delayed reward problems. Since the search space is very large and the rewards obtained during the exploration process may be sparse, this will affect the convergence of the algorithm. In order to address this problem, combined with the feedback provided by the dummy pad layout optimization module, this paper designs an appropriate immediate reward function for the dummy pad layout optimization module to reduce the adverse effects of reward sparsity on model training. The specific formula is: t =|f ReNet (s t )-τ|-|f ReNet (s t+1 )-τ|, wherein τ represents the target thickness value of the copper plating of the entire PCB; ReNet () indicates an instant reward network.
[0067] The immediate reward function can be understood as the value of the electroplated copper thickness closer to the target thickness value τ after adding the dummy pads. This reward mechanism decomposes the contribution of all dummy pads so that the placement behavior of each dummy pad can have a clear impact on the final goal. This design effectively solves the sparse reward problem and guides the policy model to explore and converge more effectively.
[0068] In this problem, the discount factor is set to γ = 1, and the reward of each step is completely equivalent to the total reward in calculation. The following is the specific proof process:
[0069]
[0070]
[0071] The c represents a constant.
[0072] S7. Confirm the placement position of the dummy pad.
[0073] The relevant operations of step S7 are introduced with a specific example:
[0074] The specific method of confirming the placement position of the dummy pad is as follows: in the current time step, the dummy pad layout optimization module samples the probability distribution matrix P of the action space at Select and get the virtual pad placement position a t , ensuring that the dummy pad layout optimization model can fully utilize every possible position during the layout generation process; the dummy pad is placed at position a t After that, the layout state changes from the current state S t Transition to the next state S in a deterministic manner t+1 ;
[0075] S8. Iteratively optimize the PCB virtual pad layout.
[0076] The relevant operations of step S8 are introduced with a specific example:
[0077] The optimization of PCB virtual pad layout based on reinforcement learning loop is specifically as follows: the virtual pad layout optimization module transmits the current state S t and the next state S t+1 To the immediate reward module, the immediate reward module transmits the immediate reward signal r t to a value model, which is based on the immediate reward signal r t Calculate the expected value of discounted cumulative reward V t And update the parameters of the strategy model. When t=T, the optimization iteration ends and the optimized PCB virtual pad layout is obtained.
[0078] In the study of electroplated copper uniformity, the optimization goal is to achieve the best result in the final placement state. t The predicted electroplated copper thickness h out =f ReNet (s t ) as close as possible to the preset target thickness value τ. To evaluate the effect of the final layout, the absolute error |h out -τ| is used as a global indicator to evaluate the effect of the final layout. By defining the reward signal r t And back-propagate it to the action selection process at each time step to guide the optimization of the policy model. t The specific definition of is:
[0079]
[0080] Among them, f ReNet (S t ) represents the instant reward network in state s t The predicted electroplating copper thickness value reflects the impact of the final dummy pad layout on the electroplating process.
[0081] In order to explicitly verify the superiority of the proposed dummy pad placement optimization model based on deep reinforcement learning optimization, we quantitatively compare it with the traditional constraint-based stochastic optimization model, e.g. Figure 5 As shown in the figure, the experimental goal is to make the electroplated copper thickness as close to the target value of 10.5μm as possible. Figure 5 In the figure, red represents the predicted results of the electroplated copper thickness in different sub-boards of the placement strategy of the dummy pad layout optimization model, and blue represents the predicted results of the constraint-based automated optimization model. In the test of 320 sub-boards, the red is closer to the target value of 10.5μm than the blue, showing a significant optimization effect. Especially in the sub-board area with low current density, the dummy pad layout optimization model can make full use of the diversion effect of the dummy pad and significantly improve the electroplated copper thickness, while the constraint-based random optimization model failed to achieve the same effect. It is worth mentioning that even though the strategy model knew nothing about the multi-physics simulation knowledge of the dummy pad placement problem in the initial stage, it still successfully discovered an efficient and elegant placement strategy through exploration and learning. This further demonstrates the potential and robustness of the dummy pad layout optimization model in solving complex problems.
Claims
1. A PCB dummy pad intelligent layout optimization method based on deep reinforcement learning, characterized in that: The optimization method comprises the following steps: S1. Decompose the entire PCB and generate several PCB sub-board layouts; S2, extracting joint features of PCB sub-boards; S3, construct an adaptive scale feature module to extract the joint feature vector of each sub-board; S4. Build and train the immediate reward network; S5. Construct and optimize the virtual pad layout optimization model; S6, calculating an immediate reward signal; S7, confirming the placement position of the dummy pad; S8. Optimize PCB dummy pad layout based on deep reinforcement learning loop.
2. According to the method for optimizing the intelligent layout of PCB virtual pads based on deep reinforcement learning in claim 1, it is characterized in that: The decomposition of the entire PCB to generate a plurality of PCB sub-board layouts is specifically as follows: according to the boundaries of the PCB sub-boards, the entire PCB is decomposed into a plurality of PCB sub-boards, and according to the actual physical size of the PCB sub-boards and the distribution of through holes, the PCB sub-boards are divided into a plurality of grid units of fixed sizes.
3. The PCB virtual pad intelligent layout optimization based on deep reinforcement learning according to claim 1, characterized in that: The joint feature includes a spatial feature M and a structural feature C; the spatial feature The 3 is the number of channels, the H and W are the height and width of the PCB sub-board layout, respectively. Represents the real number domain; the structural features The T is the maximum number of dummy pads arranged in the daughter board.
4. The PCB virtual pad intelligent layout optimization method based on deep reinforcement learning according to claim 3 is characterized in that: The adaptive scale feature module passes through the spatial branch and the structural branch from the input layer to the output layer and then passes through the aggregation module; the spatial branch passes through the pyramid pooling module and two linear transformation layers from the input layer to the output layer; the structural branch passes through the external attention mechanism module and two linear transformation layers from the input layer to the output layer; the pyramid pooling module includes a convolution layer and a maximum pooling layer from the input layer to the output layer; the external attention mechanism module includes a linear layer and a normalization layer from the input layer to the output layer; the extraction of the joint feature vector of each sub-board is specifically as follows: the joint feature is input into the adaptive scale feature module to obtain the joint feature vector h l .
5. The PCB virtual pad intelligent layout optimization method based on deep reinforcement learning according to claim 4 is characterized in that: The specific steps of establishing the instant reward network are as follows: a multi-layer perceptron is introduced on the basis of the adaptive scale feature module, wherein the multi-layer perceptron includes two linear transformation layers; the calculation formula of the instant reward network is: out =MLP(h l ); said h out It represents the copper thickness result predicted by the instant reward network, and the MLP() represents the multi-layer perceptron function.
6. The PCB virtual pad intelligent layout optimization method based on deep reinforcement learning according to claim 5 is characterized in that: The dummy pad layout optimization model includes two branches with the same structure. The construction of each branch is as follows: a multi-layer perceptron is introduced on the basis of the adaptive scale feature module; the branches are respectively set as a strategy model and a value model; the calculation formula of the strategy model is: The softmax() represents a normalization function, the Represents the probability distribution matrix of the action space, is the current state S t The calculation formula of the value model is: Said Represents the expected value of the discounted cumulative reward.
7. The PCB virtual pad intelligent layout optimization method based on deep reinforcement learning according to claim 6 is characterized in that: The virtual pad layout optimization model adopts the proximal strategy optimization algorithm to optimize the strategy model and the value model; the clipping loss function is used when optimizing the strategy model, specifically through: To achieve the L PF (θ) represents the optimized strategy model, The parameter difference between the policy model and the optimized policy model is limited to (1-∈, 1+∈), the value of ∈ is 0.2, θ represents the trainable parameter of the policy model, and R t (θ) is the probability ratio between the policy model and the optimized policy model, and the A t represents the advantage function, the A t =r t -V t , the r t Represents the immediate reward signal; when optimizing the value model, the clipping loss function is defined by the mean square error, specifically by: To achieve represents the trainable parameters of the value model, Represents the optimized value model.
8. The PCB virtual pad intelligent layout optimization method based on deep reinforcement learning according to claim 7 is characterized in that: The immediate reward signal r t The calculation formula is: t =|f ReNet (s t )-τ|-|f ReNet (s t+1 )-τ|, wherein τ represents the target thickness value of the copper plating of the entire PCB; ReNet () indicates an instant reward network.
9. The PCB virtual pad intelligent layout optimization method based on deep reinforcement learning according to claim 8 is characterized in that: The specific method of confirming the placement position of the dummy pad is as follows: in each time step, the dummy pad layout optimization module selects the probability distribution matrix of the action space by sampling to obtain the dummy pad placement position a t ; Place the dummy pad at position a t After that, the layout state changes from the current state S t Transition to the next state S t+1 .
10. The PCB virtual pad intelligent layout optimization method based on deep reinforcement learning according to claim 9 is characterized in that: The optimization of PCB virtual pad layout based on reinforcement learning loop is specifically as follows: the virtual pad layout optimization module transmits the current state S t and the next state S t+1 To the immediate reward module, the immediate reward module transmits the immediate reward signal r t to a value model, which is based on the immediate reward signal r t Calculate the expected value of discounted cumulative reward V t And update the parameters of the strategy model. When t=T, the optimization iteration ends and the optimized PCB virtual pad layout is obtained.