PCB module automatic layout optimization method and system based on deep reinforcement learning
By using deep reinforcement learning methods, combined with CNN and GNN networks, the layout of PCB modules is optimized, solving the problem of low efficiency in traditional layout methods. This achieves high efficiency, multi-objective optimization, and manufacturing process compatibility, thereby improving the performance and reliability of the circuit board.
Patent Information
- Application Number
- CN202511118850.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-21
AI Technical Summary
Existing PCB layout methods are inefficient, fail to guarantee optimal layout, are prone to errors, and struggle to balance electrical performance, heat dissipation performance, and manufacturing process constraints, especially performing poorly in high-density circuit designs.
A deep reinforcement learning-based approach is employed to optimize PCB module layout through data preprocessing, hybrid situational encoding, reinforcement learning optimization, and a closed-loop feedback mechanism. Specific steps include data cleaning, feature extraction, data augmentation, hybrid situational encoding, multi-objective fitness evaluation, and closed-loop feedback. A multi-factor weighted priority experience replay mechanism is used to optimize the layout, combining CNN and GNN network architectures.
Significantly improve layout optimization efficiency, enhance electrical and thermodynamic performance, ensure manufacturing process compliance, reduce reliance on large-scale annotation data, enhance cross-scenario adaptability and resource utilization, and achieve seamless integration with the EDA ecosystem.
Smart Images

Figure CN120995966A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electronic design automation technology, and in particular to a method and system for automatic layout optimization of PCB modules based on deep reinforcement learning. Background Technology
[0002] With the continuous miniaturization and high performance of electronic devices, the design and manufacturing technology of printed circuit boards (PCBs), as carriers of electronic components, is becoming increasingly important. The layout of PCB modules directly affects the performance, reliability, and manufacturing cost of the circuit board. In modern electronic devices, PCBs not only need to meet the ever-increasing functional requirements but also need to achieve optimal electrical and heat dissipation performance within a limited space. Traditional PCB layout methods often rely on manual experience, which is not only inefficient but also difficult to cope with the increasingly complex circuit design needs. In recent years, with the rapid development of artificial intelligence technology, technologies such as deep learning and reinforcement learning have been gradually introduced into the field of PCB design, providing new possibilities for achieving automated and intelligent PCB layout. Through deep reinforcement learning algorithms, the layout of PCB modules can be automatically learned and optimized, thereby improving design efficiency and circuit board performance.
[0003] However, existing PCB placement algorithms still have some problems in practical applications. First, traditional placement methods often require a significant investment of time and manpower, and it is difficult to guarantee optimal placement. Second, as the complexity of circuit boards and the number of components increase, manual placement is prone to errors, leading to a decline in circuit board performance. Furthermore, existing automated placement algorithms often struggle to balance electrical performance, thermal performance, and manufacturing process constraints when handling large-scale, high-density PCB layouts. These problems severely limit the efficiency and quality of PCB design. Therefore, developing an automated PCB module placement optimization method based on deep reinforcement learning is of great significance for improving the performance and reliability of electronic devices. Summary of the Invention
[0004] This invention addresses the shortcomings of existing technologies by providing a method and system for automatic layout optimization of PCB modules based on deep reinforcement learning.
[0005] To achieve the above-mentioned objectives, the technical solution adopted by the present invention is as follows:
[0006] An automatic layout optimization method for PCB modules based on deep reinforcement learning includes the following steps:
[0007] Step S1: Data Preprocessing
[0008] Cleaning PCB design data includes removing duplicate values, handling missing values, and removing outliers;
[0009] Extract the dimensional features, electrical characteristics, and position coordinates of the components;
[0010] Perform data augmentation operations on PCB images, including random rotation, flipping, or translation;
[0011] Step S2: Hybrid Situation Coding
[0012] Features of the preprocessed PCB raster image Input the CNN branch and output the global feature f img ;
[0013] The component interconnection diagram G=(V,E) and node characteristics are defined. Inputting a GNN branch, the output node embeddings are processed through L layers of graph convolution.
[0014] Features are concatenated and calculated in the fusion layer:
[0015]
[0016] Where || represents the concatenation operation, W f Here, f is the weight matrix of the fusion layer, V is the set of all component nodes, and f is the weight matrix of the fusion layer. img The extracted global feature vector of the PCB image, b f H is the bias vector of the fusion layer. (L) This is the output of the Lth layer graph convolution.
[0017] Step S3: Reinforcement Learning Optimization
[0018] Define action space A, which includes operations for adjusting the position of components:
[0019] The move operation moves the selected component a predetermined distance in a specified direction: a move (i,Δx,Δy);
[0020] Swap operation, swapping the positions of two components: a swap (i,j);
[0021] Rotation operation, rotate the selected component: a rotate (i,θ).
[0022] Experience is sampled using a multi-factor weighted priority experience replay mechanism (MF-PER), with the following sample priority:
[0023] p i =(|δ i |+ε) α ·(|ΔR i |+ε′) β
[0024] Where, δi α and β control the sensitivity of each term; ε and ε′ are stability factors to prevent them from being zero.
[0025] Step S4: Multi-objective fitness assessment
[0026] Calculate the Tchebycheff aggregate function:
[0027]
[0028] Among them, R i It is the ideal value of the i-th sub-objective; w i It is a dynamic weight, and the update formula is:
[0029]
[0030] Where β is the weighted focusing intensity parameter;
[0031] Step S5: Closed-loop feedback
[0032] Based on the electrical performance, thermal distribution, and DFM indicators of the actual PCB test, adjust the model parameters and iteratively execute S1-S4.
[0033] Furthermore, the data augmentation in step S1 includes:
[0034] Random rotation: rotation angle θ∈[-15°,15°], coordinate transformation formula is:
[0035]
[0036] To flip horizontally or vertically, use the following formula:
[0037] (x′,y′)=(ex,y)
[0038] (x′,y′)=(x,hy)
[0039] Random translation, the formula is:
[0040] (x′,y′)=(x+Δx,y+Δy)
[0041] Where x′, y′ are the original coordinates; x, y are the new coordinates; w, h are the image width and height; Δx, Δy are the horizontal and vertical translation amounts.
[0042] Furthermore, the graph convolution calculation in the GNN branch of step S2 satisfies:
[0043] H (l+1) =σ(D -1 / 2 (A+I)D -1 / 2 H (l) W (l) ),H(0) =X node
[0044] Where: a is the adjacency matrix; D is the degree matrix; I is the identity matrix; and σ is the ReLU activation function.
[0045] Furthermore, the Z-score method is used to handle missing and outlier values in step S1:
[0046]
[0047] Where, x i Here, μ is the data point, σ is the mean of the data, and σ is the standard deviation of the data.
[0048] This invention also discloses an automatic PCB module layout optimization system based on the above method, characterized in that it includes:
[0049] Data preprocessing module: Cleans PCB design data, including removing duplicate values, handling missing values and outliers; extracts the dimensional features, electrical characteristics and position coordinates of components; and performs data augmentation operations on PCB images, including random rotation, flipping or translation.
[0050] Hybrid situation coding module: Receives PCB grid image features and component interconnection diagram data from the data preprocessing module, processes the image features through a CNN branch to obtain global features, processes the component interconnection diagram and node features through a GNN branch to obtain node embeddings, and fuses the global features and node embeddings to obtain fused features;
[0051] Reinforcement learning optimization module: Defines the action space based on the fused features, and uses a multi-factor weighted priority experience replay mechanism to sample experience data and generate layout optimization actions;
[0052] Multi-objective fitness evaluation module: Evaluates layout schemes based on multiple objective functions and calculates fitness values using the Tchebycheff aggregation function;
[0053] Closed-loop feedback module: Based on the measured PCB electrical performance, thermal distribution and manufacturability indicators, adjust the model parameters of the reinforcement learning optimization module and trigger iterative optimization.
[0054] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described method.
[0055] The present invention also discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method.
[0056] Compared with the prior art, the advantages of the present invention are as follows:
[0057] 1. Optimize efficiency and improve performance
[0058] Significantly shorten the layout optimization cycle: By integrating a deep reinforcement learning framework, end-to-end rapid processing is achieved from initial layout generation to multi-objective optimization;
[0059] Significantly improves the ability to search for the global optimum: By combining a hybrid neural network architecture with a multi-factor experience replay mechanism, it effectively avoids the trap of local optima;
[0060] Enhanced adaptability to complex scenarios: Demonstrates stronger robustness to complex PCB design scenarios such as high-density interconnects and high-speed signals.
[0061] 2. Breakthrough in multi-objective collaborative optimization
[0062] Precise control of electrical performance: synchronously optimizes signal integrity, power integrity and electromagnetic compatibility indicators;
[0063] Thermodynamic performance is improved in a balanced way: automatically balancing hotspot distribution and heat dissipation path planning;
[0064] Deep compatibility with manufacturing processes: The embedded DFM rule engine ensures that the layout scheme conforms to the constraints of mounting / soldering processes;
[0065] Dynamic weight focus mechanism: Automatically strengthens the optimization of the weakest link based on real-time performance deviation.
[0066] 3. Improvements in resource consumption and generalization ability
[0067] Reduce data dependency: Reduce the need for large-scale labeled data through topological feature extraction and data augmentation strategies;
[0068] Reduce computational resource overhead: Optimize the number of neural network parameters and training process to improve hardware resource utilization;
[0069] Enhance cross-scenario generalization: Establish a mapping relationship between circuit topology features and layout rules to adapt to various PCB design specifications.
[0070] 4. Enhanced project feasibility
[0071] Seamless integration with the EDA ecosystem: compatible with the input / output interface specifications of mainstream design tools;
[0072] Closed-loop design verification system: Iterative upgrades of layout solutions are achieved through feedback from real-world testing.
[0073] Design rule self-learning mechanism: Automatically absorbs feedback data from the manufacturing end to continuously optimize the constraint model.
[0074] 5. The value of technological integration and innovation
[0075] Breaking through the limitations of traditional algorithms: Integrating the spatial relationship modeling of graph neural networks with the visual feature extraction advantages of CNNs;
[0076] Reconstructing the reinforcement learning paradigm: Innovating a multi-objective reward-driven mechanism to improve convergence efficiency;
[0077] Building cross-disciplinary technology bridges: Deeply embedding multi-objective optimization theory into electronic design automation processes. Attached Figure Description
[0078] Figure 1 This is a flowchart of the automatic layout optimization method for PCB modules based on deep reinforcement learning, according to an embodiment of the present invention. Detailed Implementation
[0079] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and examples.
[0080] like Figure 1 As shown, an automatic layout optimization method for PCB modules based on deep reinforcement learning includes the following steps:
[0081] (1) Data cleaning is an important step in data preprocessing, removing noise and outliers from the data. The first step is to remove duplicate values. This is done by comparing the unique identifiers or specific attribute values of data records to delete duplicate records. The formula is as follows:
[0082]
[0083] Where D is the original dataset, and D′ is the dataset after removing duplicate values.
[0084] (2) Next, missing values need to be processed. For numerical data, the median or mean can be used to fill missing values; for categorical data, the mode can be used to fill missing values.
[0085] (3) For outlier handling, outliers are detected, removed, or corrected using the z-score method. The z-score of each data point is calculated; if the absolute value of the z-score is greater than a threshold, it is considered an outlier. The formula is:
[0086]
[0087] Where, x i Here, μ is the data point, σ is the mean of the data, and σ is the standard deviation of the data.
[0088] (4) Feature extraction is the process of converting the characteristics of components into an input form that the model can understand. It extracts the dimensional features of the components, such as length and width, electrical characteristics of current and voltage, and coordinate positions on the circuit board, from the PCB layout data.
[0089] (5) Data augmentation increases the diversity of training data by transforming the data, thereby reducing the risk of overfitting.
[0090] (6) Randomly rotate the PCB image, typically within the range of [-15°, 15°], using the following formula:
[0091]
[0092] Where θ is the rotation angle.
[0093] (7) To flip the PCB image horizontally or vertically, the formula is:
[0094] (x′,y′)=(wx,y)
[0095] (x′,y′)=(x,hy)
[0096] Where w and h are the width and height of the image, respectively.
[0097] (8) Randomly shift the PCB image. The shift range is usually no more than 20% of the image. The formula is:
[0098] (x′,y′)=(x+Δx,y+Δy)
[0099] Where Δx and Δy are randomly selected translation amounts.
[0100] (9) Select a suitable model architecture, instantiate the selected model architecture, and set the parameters of each layer.
[0101] (10) The process of adjusting model parameters through algorithm optimization so that the model can minimize the loss function on the training data. In the PCB module layout task, the loss function can be defined as the difference between the predicted layout and the actual layout, using the mean squared error loss function:
[0102]
[0103] Among them, y i y is the actual layout position of the i-th sample. ′ i N is the layout location predicted by the model, and N is the number of samples.
[0104] (11) Use stochastic gradient descent or its variants to update model parameters. The parameter update direction is adaptively adjusted during training. The parameter update formula is:
[0105] m t =β1m t-1 +(1-β1)g t
[0106]
[0107] Among them, g t It is the gradient of the t-th iteration, m t and v t These are the first-order moment estimate and the second-order moment estimate of the gradient, θ. t These are the model parameters, α is the learning rate, β1 and β2 are the decay coefficients of RMSprop, and ∈ is a very small number used to prevent division by zero.
[0108] (12) Input the preprocessed PCB design data into the trained model, including the characteristics and layout requirements of the components.
[0109] (13) Based on the input data, the model generates an initial layout scheme and outputs the position coordinates of each component. The model's output can be represented as:
[0110] P = {(x1,y1),(x2,y2),…,(x...} n ,y n )}
[0111] Among them, (x i ,y i ) is the position coordinate of the i-th component, and n is the total number of components.
[0112] (14) Further optimize the generated layout scheme by continuously adjusting the positions of components through reinforcement learning algorithms:
[0113] 1. Define the state space S to represent the current state of the PCB layout, including the current position coordinates of all components:
[0114] (x1,y1),(x2,y2),…,(x n ,y n )
[0115] 2. Define motion space A, including operations for adjusting the position of components:
[0116] The move operation moves the selected component a predetermined distance in a specified direction: a move (i,Δx,Δy).
[0117] Swap operation, swapping the positions of two components: a swap (i,j).
[0118] Rotation operation, rotate the selected component: a rotate (i,θ).
[0119] Actions can be represented as vectors: a = [a type (i,j)] m×n , where a type Indicates the action type, and i and j represent the component indices of the operation.
[0120] (15) Continue until the preset optimization goal or maximum number of iterations is reached. The optimized layout scheme can be expressed as:
[0121] P′={(x′1,y′1),(x′2,y′2),…,(x′ n ,y′ n )}
[0122] Among them, (x′ i ,y′ i The optimized position coordinates of the i-th component.
[0123] (16) By defining the reward function R(s,a), we can initially evaluate the degree of improvement of each action on the layout performance.
[0124] Line length reward R l :h j (s)=-∑ i ∑ k C i (s)d k,j (i), where d k,j (i) is the Manhattan distance between components k and j.
[0125] Heat distribution reward R t :R t (s)=-σ(T), where σ(T) is the standard deviation of the temperature distribution.
[0126] Density reward R p :R p (s) = -max(D(x,y)), where D(x,y) is the component density at position (x,y).
[0127] Manufacturing constraint reward R m :R m (s)=-∑i max (0,v(i)-v max ), where v(i) is the degree of violation of manufacturing constraints.
[0128] Comprehensive reward function: R(s,a)=w1·R t +w2·R t +w3·R ρ +w4·R m , where w k The weights are for each item.
[0129] (17) The novel deep network architecture and multi-factor experience replay mechanism proposed in this invention can more effectively capture the spatial distribution and connection topology of PCB components and accelerate the convergence of multi-objective optimization.
[0130] (18) A novel deep network architecture proposed in this invention is adopted, which integrates a hybrid situation encoder of convolutional neural network (CNN) and graph neural network (GNN). Local image features and device connection context information are extracted to enhance the representation capability of layout position prediction.
[0131] 1. Input: PCB spatial raster image features Component interconnection diagram structure G=(V,E) and node characteristics
[0132] 2. Network structure process:
[0133] a) CNN branch: Convolution + Pooling → Global Features
[0134] b) GNN branch: L-layer graph convolution → node embedding
[0135] c) Fusion layer, which combines average node embeddings with image features:
[0136]
[0137] Where || represents the concatenation operation, W f Here, f is the weight matrix of the fusion layer, V is the set of all component nodes, and f is the weight matrix of the fusion layer. img The extracted global feature vector of the PCB image, b f H is the bias vector of the fusion layer. (L) This is the output of the Lth layer graph convolution.
[0138] d) Coordinate Output Layer: A fully connected network predicts the placement coordinates (x, y) of each component. i ,y i ).
[0139] Mathematical form of graph convolutional layers:
[0140] H (l+1) =σ(D -1 / 2 (A+I)D -1 / 2 H (l) W(l) ),H (0) =X node
[0141] Where A is the adjacency matrix, D is the degree matrix, and α is the nonlinear activation function.
[0142] (19) This invention introduces a multi-factor weighted priority experience replay mechanism based on the classic TD error-driven priority experience replay (PER), defining the priority of each sample as a joint function of TD error and multi-objective reward changes, thereby improving the utilization efficiency of experience samples. Specific details are as follows:
[0143] 1. The multi-objective reward gain is defined as:
[0144]
[0145] Among them, R j Let w represent the evaluation function for the j-th sub-objective. j For weights.
[0146] 2. Sample priority is defined as:
[0147] p i =(|δ i |+ε) α ·(|ΔR i |+ε′) β
[0148] Where, δ i α and β control the sensitivity of each term; ε and ε′ are stability factors to prevent them from being zero.
[0149] 3. The sampling probability is:
[0150]
[0151] (20) The present invention proposes a multi-objective fitness aggregation strategy based on the Tchebycheff method to simultaneously measure the optimization effect of multiple indicators such as signal integrity, electromagnetic compatibility, thermal uniformity and manufacturing process compliance. The main steps are as follows:
[0152] 1. Definition of an ideal vector
[0153]
[0154] Among them, R i (P) represents the evaluation of the i-th sub-objective, and m represents the total number of sub-objectives.
[0155] 2. Tchebycheff aggregate function
[0156]
[0157] Where w = (w1, w2, ..., w m ) represents the normalized weighting coefficients, satisfying This represents the absolute deviation between the i-th indicator and its ideal value, reflecting the degree of "insufficiency" of the target.
[0158] 3. Dynamic weight adaptive update
[0159]
[0160] The β control focuses on the intensity of attention given to indicators with larger deviations. The summation of the beta powers of the deviations of all sub-targets is used for normalization.
[0161] (21) After model output evaluation, substitute the optimized layout scheme P′ into the above F. T (P;w), calculate the final fitness value, and use it as a comprehensive evaluation index of the layout quality.
[0162] (22) By manufacturing and testing actual circuit boards, key indicators such as signal error, electromagnetic radiation, temperature distribution, and DFM pass rate are obtained and fed back to the model, thereby achieving closed-loop iterative optimization.
[0163] In another embodiment of the present invention, an automatic PCB module layout optimization system is provided, which can be used to implement the above-described method, specifically including:
[0164] Data preprocessing module: Cleans PCB design data, including removing duplicate values, handling missing values and outliers; extracts the dimensional features, electrical characteristics and position coordinates of components; and performs data augmentation operations on PCB images, including random rotation, flipping or translation.
[0165] Hybrid situation coding module: Receives PCB grid image features and component interconnection diagram data from the data preprocessing module, processes the image features through a CNN branch to obtain global features, processes the component interconnection diagram and node features through a GNN branch to obtain node embeddings, and fuses the global features and node embeddings to obtain fused features;
[0166] Reinforcement learning optimization module: Defines the action space based on the fused features, and uses a multi-factor weighted priority experience replay mechanism to sample experience data and generate layout optimization actions;
[0167] Multi-objective fitness evaluation module: Evaluates layout schemes based on multiple objective functions and calculates fitness values using the Tchebycheff aggregation function;
[0168] Closed-loop feedback module: Based on the measured PCB electrical performance, thermal distribution and manufacturability indicators, adjust the model parameters of the reinforcement learning optimization module and trigger iterative optimization.
[0169] In another embodiment of the present invention, a terminal device is provided, comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to achieve a corresponding method flow or corresponding function. The processor described in this embodiment of the present invention can be used for the operation of the above-mentioned method.
[0170] In another embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory), which is a memory device in a terminal device used to store programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and extended storage media supported by the terminal device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor, which can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device.
[0171] One or more instructions stored in a computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the above-described methods in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by a processor.
[0172] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0173] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0174] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0175] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0176] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the implementation methods of the present invention, and should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of the present invention.
Claims
1. A method for automatic layout optimization of PCB modules based on deep reinforcement learning, characterized in that, Includes the following steps: Step S1: Data Preprocessing Cleaning PCB design data includes removing duplicate values, handling missing values, and removing outliers; Extract the dimensional features, electrical characteristics, and position coordinates of the components; Perform data augmentation operations on PCB images, including random rotation, flipping, or translation; Step S2: Hybrid Situation Coding Features of the preprocessed PCB raster image Input the CNN branch and output the global feature f img ; The component interconnection diagram G=(V,E) and node characteristics are defined. Inputting a GNN branch, the output node embeddings are processed through L layers of graph convolution. Features are concatenated and calculated in the fusion layer: Where || represents the concatenation operation, W f Here, f is the weight matrix of the fusion layer, V is the set of all component nodes, and f is the weight matrix of the fusion layer. img The extracted global feature vector of the PCB image, b f H is the bias vector of the fusion layer. (L) The output of the Lth layer graph convolution; Step S3: Reinforcement Learning Optimization Define action space A, which includes operations for adjusting the position of components: The move operation moves the selected component a predetermined distance in a specified direction: a move (i,Δx,Δy); Swap operation, swapping the positions of two components: a swap (i,j); Rotation operation, rotate the selected component: a rotate (i,θ); Experience is sampled using a multi-factor weighted priority experience replay mechanism (MF-PER), with the following sample priority: p i =(|δ i |+e) α ·(|ΔR i |+e′) β Where, δ i TD error; α and β control the sensitivity of each term; ε and ε′ are stability factors to prevent them from being zero; Step S4: Multi-objective fitness assessment Calculate the Tchebycheff aggregate function: Among them, R i It is the ideal value of the i-th sub-objective; w i It is a dynamic weight, and the update formula is: Where β is the weighted focusing intensity parameter; Step S5: Closed-loop feedback Based on the electrical performance, thermal distribution, and DFM indicators of the actual PCB test, adjust the model parameters and iteratively execute S1-S4.
2. The method according to claim 1, characterized in that, The data augmentation in step S1 includes: Random rotation: rotation angle θ∈[-15°,15°], coordinate transformation formula is: To flip horizontally or vertically, use the following formula: (x′,y′)=(wx,y) (x′,y′)=(x,hy) Random translation, the formula is: (x′,y′)=(x+Δx,y+Δy) Where x′, y′ are the original coordinates; x, y are the new coordinates; w, h are the image width and height; Δx, Δy are the horizontal and vertical translation amounts.
3. The method according to claim 1, characterized in that, In step S2, the graph convolution calculation in the GNN branch satisfies: H (l+1) =σ(D -1 / 2 (A+I)D -1 / 2 H (l) W (l) ),H (0) =X node Where: A is the adjacency matrix; D is the degree matrix; I is the identity matrix; and σ is the ReLU activation function.
4. The method according to claim 1, characterized in that, In step S1, missing and outlier values are handled using the Z-score method: Where, x i Here, μ is the data point, σ is the mean of the data, and σ is the standard deviation of the data.
5. A PCB module automatic layout optimization system based on the method of any one of claims 1 to 4, characterized in that, include: Data preprocessing module: Cleans PCB design data, including removing duplicate values, handling missing values and outliers; Extract the dimensional features, electrical characteristics, and position coordinates of components; and perform data augmentation operations on the PCB image, including random rotation, flipping, or translation; Hybrid situation coding module: Receives PCB grid image features and component interconnection diagram data from the data preprocessing module, processes the image features through a CNN branch to obtain global features, processes the component interconnection diagram and node features through a GNN branch to obtain node embeddings, and fuses the global features and node embeddings to obtain fused features; Reinforcement learning optimization module: Defines the action space based on the fused features, and uses a multi-factor weighted priority experience replay mechanism to sample experience data and generate layout optimization actions; Multi-objective fitness evaluation module: Evaluates layout schemes based on multiple objective functions and calculates fitness values using the Tchebycheff aggregation function; Closed-loop feedback module: Based on the measured PCB electrical performance, thermal distribution and manufacturability indicators, adjust the model parameters of the reinforcement learning optimization module and trigger iterative optimization.
6. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, It contains a computer program that, when executed by a processor, implements the method described in any one of claims 1 to 4.