Circuit board layout method and apparatus, electronic device, and storage medium

CN122197787APending Publication Date: 2026-06-12GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU SHIYUAN ELECTRONICS CO LTD
Filing Date
2024-12-12
Publication Date
2026-06-12

Smart Images

  • Figure CN122197787A_ABST
    Figure CN122197787A_ABST
Patent Text Reader

Abstract

The application relates to a circuit board layout method and device, electronic equipment and a storage medium. According to circuit device information, the application embodiment determines a current layout state of a sample circuit board. A reinforcement learning agent outputs a current action and a current Q value according to the current layout state, obtains a current reward value, and performs state transition to obtain a next layout state of the current layout state. The above process is repeated, so that the reinforcement learning agent continuously learns, the reinforcement learning agent is optimized through the current Q value and the current reward value, and therefore a layout result meeting a layout constraint condition and having high aesthetic degree is obtained, and the layout efficiency of the circuit board is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of circuit board layout design technology, and in particular to a circuit board layout method, apparatus, electronic device, and storage medium. Background Technology

[0002] A printed circuit board (PCB) is the support structure for electronic components and the carrier for the electrical interconnection of electronic components.

[0003] Currently, circuit board layout is mainly done manually. Specifically, based on the circuit schematic, engineers use layout design software to determine the appropriate placement and rotation angle for each component within a specified layout area.

[0004] However, with the increase in circuit board integration and the size of components, manual circuit board layout involves a lot of repetitive work and long adjustment time, resulting in low circuit board layout efficiency. Summary of the Invention

[0005] Based on this, the purpose of the present invention is to provide a circuit board layout method, apparatus, electronic device, and storage medium, which have the advantage of improving circuit board layout efficiency.

[0006] According to a first aspect of the embodiments of this application, a circuit board layout method is provided, comprising the following steps:

[0007] Obtain the circuit component information of the sample circuit board; the circuit component information includes the information of the already laid-out components and the information of the components to be laid out.

[0008] Based on the circuit component information, the current layout state of the sample circuit board is obtained; where the current layout state includes the current layout state mask and the current undirected graph;

[0009] The current layout state of the sample circuit board is input into the reinforcement learning agent to obtain the current action and the current Q value; based on the current action and the initial layout position of the device to be laid out, the next layout state of the current layout state is determined; based on the next layout state of the current layout state and the preset reward function, the current reward value is calculated; the current layout state, the next layout state of the current layout state, the current action, the current reward value, and the current Q value are stored as a set of sample data in the experience pool; wherein, the device to be laid out is a device selected from the devices to be laid out.

[0010] When the number of sample data sets in the experience pool is greater than or equal to a preset number, the reinforcement learning agent is trained based on the sample data in the experience pool to obtain a trained reinforcement learning agent.

[0011] Obtain the circuit component information of the circuit board to be laid out; determine the current layout state of the circuit board to be laid out based on the circuit component information of the circuit board to be laid out; input the current layout state of the circuit board to be laid out into the trained reinforcement learning agent to obtain the layout position of the component to be laid out in the circuit board to be laid out.

[0012] Based on the layout location, the devices to be laid out are placed.

[0013] According to a second aspect of the embodiments of this application, a circuit board layout apparatus is provided, comprising:

[0014] The circuit component information acquisition module is used to acquire the circuit component information of the sample circuit board; wherein, the circuit component information includes the information of the already laid-out components and the information of the components to be laid out.

[0015] The current layout state acquisition module is used to obtain the current layout state of the sample circuit board based on the circuit component information; wherein, the current layout state includes the current layout state mask and the current undirected graph;

[0016] The current action acquisition module is used to input the current layout state of the sample circuit board into the reinforcement learning agent to obtain the current action and the current Q value; determine the next layout state of the current layout state based on the current action and the initial layout position of the device to be laid out; calculate the current reward value based on the next layout state of the current layout state and the preset reward function; and store the current layout state, the next layout state of the current layout state, the current action, the current reward value, and the current Q value as a set of sample data into the experience pool; wherein, the device to be laid out is a device selected from the devices to be laid out.

[0017] The reinforcement learning agent training module is used to train the reinforcement learning agent based on the sample data in the experience pool when the number of sample data sets in the experience pool is greater than or equal to a preset number, so as to obtain a trained reinforcement learning agent.

[0018] The layout position acquisition module is used to obtain the circuit device information of the circuit board to be laid out; based on the circuit device information of the circuit board to be laid out, the current layout state of the circuit board to be laid out is determined, and the current layout state of the circuit board to be laid out is input into the trained reinforcement learning agent to obtain the layout position of the device to be laid out in the circuit board to be laid out.

[0019] The device placement module is used to place the devices to be placed according to their placement positions.

[0020] According to a third aspect of the embodiments of this application, an electronic device is provided, including: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed as a circuit board layout method as described above.

[0021] According to a fourth aspect of the embodiments of this application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the circuit board layout method as described in any of the above claims.

[0022] This embodiment of the application determines the current layout state of the sample circuit board based on circuit device information. The reinforcement learning agent outputs its current action and current Q-value based on the current layout state, obtains its current reward value, and performs a state transition to obtain the next layout state. This process is repeated, allowing the reinforcement learning agent to continuously learn and optimize itself using the current Q-value and current reward value. This results in a layout that meets layout constraints and has a high aesthetic appeal, improving the layout efficiency of the circuit board.

[0023] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit this application.

[0024] To better understand and implement this invention, the following detailed description is provided in conjunction with the accompanying drawings. Attached Figure Description

[0025] Figure 1 A schematic flowchart illustrating a circuit board layout method provided in one embodiment of this application;

[0026] Figure 2 This is a structural block diagram of a circuit board layout apparatus provided in one embodiment of this application;

[0027] Figure 3 This is a schematic block diagram of the structure of an electronic device provided in one embodiment of this application. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0029] It should be understood that the described embodiments are merely some, not all, of the embodiments in this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0030] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the embodiments of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0031] In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims. In the description of this application, it should be understood that the terms "first," "second," "third," etc., are used only to distinguish similar objects and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0032] Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0033] In the process of developing this invention, the inventors discovered that current circuit board layout is mainly done manually. When there are many PCB modules to be laid out on a circuit board and they are highly similar, manual layout often involves a lot of repetitive work. Furthermore, in order to meet the layout constraints and achieve the desired layout effect, adjustments to the components are required, which takes a long time and results in low circuit board layout efficiency.

[0034] To this end, this application models the circuit board layout problem as a decision-making process of a reinforcement learning agent. The reinforcement learning agent performs actions based on the layout state, obtains rewards, and continuously learns through state transitions to optimize its policy network, thereby obtaining a layout result that maximizes the layout objective (meets layout constraints and has high aesthetic appeal), thus improving the layout efficiency of the circuit board.

[0035] Please see Figure 1 This is a flowchart illustrating a circuit board layout method provided in one embodiment of this application. This application provides a circuit board layout method, including the following steps:

[0036] S10: Obtain the circuit component information of the sample circuit board; wherein, the circuit component information includes the information of the already laid-out components and the information of the components to be laid out.

[0037] The circuit board layout method of this application is implemented by a circuit board layout device, which includes, but is not limited to, computers, tablets, and mobile phones.

[0038] The information on placed devices includes their position and parameters. The position information includes their coordinates and rotation angle. The device parameters include their length, width, device type, pin net number, pin type, and pin coordinates.

[0039] The device information to be laid out includes the device parameters of the device to be laid out, including the length, width, device type, pin net number, and pin type of the device to be laid out.

[0040] In this embodiment of the application, the user can import the circuit schematic in JSON file format through PCB layout software. After receiving the circuit schematic imported by the user, the circuit component information of the sample circuit board is extracted from the circuit schematic.

[0041] S20: Obtain the current layout state of the sample circuit board based on the circuit device information; wherein, the current layout state includes the current layout state mask and the current undirected graph.

[0042] The current layout state mask includes a laid-out state mask and a pending layout state mask. The laid-out state mask represents the location information, device type, and pin connection relationship of the already laid-out devices. The pending layout state mask represents the device type and pin connection relationship of the devices to be laid out.

[0043] An undirected graph consists of nodes with object characteristics and edges representing the relationships between nodes. In circuit device placement, undirected graphs can be used to represent the various devices (nodes) in a circuit and the electrical connections (edges) between them. The current undirected graph is used to represent the overall information of a sample circuit board.

[0044] In this embodiment of the application, there are multiple devices to be laid out on the sample circuit board, and each device needs to be laid out individually. Therefore, any one device is selected from the devices to be laid out as the current device to be laid out, and the current layout state is determined based on the information of the current device to be laid out and the information of the devices already laid out.

[0045] S30: Input the current layout state of the sample circuit board into the reinforcement learning agent to obtain the current action and the current Q value; determine the next layout state of the current layout state based on the current action and the initial layout position of the device to be laid out; calculate the current reward value based on the next layout state of the current layout state and the reward function; store the current layout state, the next layout state of the current layout state, the current action, the current reward value, and the current Q value as a set of sample data into the experience pool; wherein, the device to be laid out is a device selected from the devices to be laid out.

[0046] In this process, reinforcement learning agents interact with the environment, obtain the current layout state from the environment, perform a certain action based on the current layout state, and learn problem-solving strategies by obtaining corresponding rewards based on the actions performed. By exploring the environment and adjusting their strategies based on the reward feedback obtained, the strategies can maximize future cumulative rewards.

[0047] The current action is the action selected from the action space, which represents a set of actions that a reinforcement learning agent can take. It can be a set of all positional information for placing the device to be laid out.

[0048] The initial layout position includes the initial position coordinates and the initial rotation angle.

[0049] The Q-value refers to the expected reward of taking a specific action in a given state, and is used to evaluate the merits of all actions in that state. The current Q-value is the expected reward of taking the current action in the current layout state.

[0050] The reward function is used to calculate the reward value for reinforcing the agent to perform a certain action in a certain state.

[0051] The experience pool is a container used to store the experiences generated by the reinforcement learning agent during its interaction with the environment. These experiences are usually in the form of a quadruple (s, a, r, s'), where s represents the current state, a represents the action taken in the current state, r represents the reward value obtained after performing the action, and s' represents the new state transitioned to after performing the action.

[0052] In this embodiment, a reinforcement learning agent is used to lay out each device to be laid out on a sample circuit board. Specifically, a blank mask of size 256*256 is obtained, and the initial position coordinates of the device to be laid out are set as the center point coordinates of the blank mask. The initial rotation angle of the device to be laid out is set to 0 degrees. Let the layout state at the initial time t0 be s0. The first device to be laid out is laid out. Based on the layout state s0, the reinforcement learning agent outputs the current action a0 and the current Q value q0, and obtains the feedback reward value r0 and the layout state s1 at the next time t1. Here, a0 represents the position information of the first device to be laid out, and the layout state s1 is determined based on the current action a0, the initial position coordinates of the first device to be laid out, and the initial rotation angle. To lay out the second component, the reinforcement learning agent outputs the current action a1 and the current Q value q1 based on the layout state s1, and obtains the feedback reward value r1 and the layout state s2 at the next time step t2. Here, a1 represents the position information of the second component to be laid out. The layout state s2 is determined based on the current action a1, the initial position coordinates of the second component to be laid out, and the initial rotation angle. This process is repeated for each component to be laid out, resulting in a layout state sT and a reward value rT.

[0053] S50: When the number of sample data sets in the experience pool is greater than or equal to the preset number, the reinforcement learning agent is trained based on the sample data in the experience pool to obtain the trained reinforcement learning agent.

[0054] The preset quantity can be manually set according to actual needs.

[0055] The reinforcement learning agent includes a policy network and a value network, which are neural network models. The policy network is used to output actions based on the state, and the value network is used to output Q-values ​​based on the state and actions.

[0056] In this embodiment, the network weight parameters of the policy network and the value network are trained based on sample data in the experience pool to obtain a trained reinforcement learning agent. Specifically, the reinforcement learning agent is based on the Deep Reinforcement Algorithm (SAC) framework. Through the loss function, Q-value, and reward value provided by the SAC framework, the network weight parameters of the policy network and the value network can be continuously optimized to obtain a trained policy network and value network.

[0057] S60: Obtain the circuit device information of the circuit board to be laid out; determine the current layout state of the circuit board to be laid out based on the circuit device information of the circuit board to be laid out, input the current layout state of the circuit board to be laid out into the trained reinforcement learning agent, and obtain the layout position of the device to be laid out in the circuit board to be laid out.

[0058] In this embodiment, after obtaining the trained reinforcement learning agent, device placement can be performed on the circuit board to be laid out. The process of obtaining the layout state of the circuit board to be laid out is the same as that of the sample circuit board, and will not be described again here. The current layout state of the circuit board to be laid out is input to the trained reinforcement learning agent, which outputs corresponding actions to update the current layout state, thereby obtaining the placement positions of the devices to be laid out one by one.

[0059] S70: Layout the devices to be laid out according to their layout positions.

[0060] The layout position includes the position coordinates and rotation angle.

[0061] In this embodiment of the application, the device to be laid out is placed at the corresponding position coordinates, and the rotation angle of the device to be laid out is set to the corresponding rotation angle.

[0062] According to the embodiments of this application, the current layout state of the sample circuit board is determined based on the circuit device information. The reinforcement learning agent outputs the current action and the current Q-value based on the current layout state, obtains the current reward value, and performs a state transition to obtain the next layout state. By repeating the above process, the reinforcement learning agent continuously learns and optimizes itself using the current Q-value and the current reward value, thereby obtaining a layout result that meets the layout constraints and has a high aesthetic appeal, thus improving the layout efficiency of the circuit board.

[0063] In one embodiment, the information of the placed devices includes the position information and device parameters of the placed devices; the information of the devices to be placed includes the device parameters of the devices to be placed; the current placement state mask includes the placed state mask and the placement state mask. Step S30 includes steps S31 to S33, as follows:

[0064] S31: Construct the layout state mask based on the position information and parameters of the already laid-out devices.

[0065] The layout status mask includes the pin relationship mask of the layout devices and the device type mask of the layout devices.

[0066] In this embodiment, the layout area of ​​the sample circuit board is mapped to a high-fine-grained mask of size 256*256, and the layout area includes all laid-out devices. The layable area in the high-fine-grained mask is marked as 0, and for each device, it is marked with 1 in the high-fine-grained mask based on its position information (position coordinates and rotation angle) and shape information (length and width). In the high-fine-grained mask, pins are marked using pin type numbers; specifically, pins can be numbered according to their order of appearance to obtain a pin relationship mask for the laid-out devices. For each device, a device type mask is obtained by marking it using device type numbers.

[0067] S32: Construct the state mask for the device to be laid out based on the initial placement position and device parameters.

[0068] The state mask to be laid out includes the pin relationship mask of the device to be laid out and the device type mask of the device to be laid out.

[0069] In this embodiment, the device to be laid out is placed at the center of a 256*256 blank mask, and the initial rotation angle of the device is set to 0 degrees. The layable area in the blank mask is marked as 0. For the device to be laid out, it is marked with 1 in the blank mask based on its position information (position coordinates and rotation angle) and shape information (length and width). In the blank mask, pins are marked with pin type numbers; specifically, pins can be numbered according to their order of appearance to obtain the pin relationship mask of the device to be laid out. Devices are marked with device type numbers to obtain the device type mask of the device to be laid out.

[0070] S33: Construct the current undirected graph based on the position information and parameters of the already placed devices, the initial placement position and parameters of the device to be placed.

[0071] In this embodiment, the device layout state (including laid-out state, current pending layout state, and unlaid-out state), device type, device length, width, position coordinates (the position coordinates of unlaid-out devices are set to 0), rotation angle, pin network number, and pin position coordinates are used as node information, and the pin connection relationship of the currently laid-out devices is used as the edge information of the undirected graph to obtain the current undirected graph.

[0072] To avoid introducing device type information into the order relationship, the layout status, angle, component type, and pin type in the node information are all encoded using one-hot encoding. For example, the laid-out status is represented as [1, 0, 0], the current pending layout status is represented as [0, 1, 0], and the unlaid-out status is represented as [0, 0, 1]. Meanwhile, other node information is normalized.

[0073] This application embodiment, by preprocessing the information of already placed devices and devices to be placed, can automatically and quickly obtain the pin relationship mask, device type mask, pin relationship mask, device type mask, and current undirected graph of the devices already placed. These masks and the current undirected graph can reflect the state information of each placement stage, thereby improving the learning effect of the reinforcement learning agent.

[0074] In one embodiment, the reinforcement learning agent includes a policy network and a value network. The policy network includes a first convolutional network, a second convolutional network, a first fully connected layer, and a first output layer; the value network includes a third convolutional network, a fourth convolutional network, a second fully connected layer, and a second output layer. Step S40, which inputs the current layout state into the reinforcement learning agent to obtain the current action and the current Q-value, includes steps S41 to S42, as follows:

[0075] S41: Input the already laid-out state mask and the state mask to be laid out into the first convolutional network to obtain the first feature vector; input the current undirected graph into the second convolutional network to obtain the second feature vector; input the first feature vector and the second feature vector into the first fully connected layer to obtain the third feature vector; input the third feature vector into the first output layer to obtain the current action.

[0076] The system comprises two convolutional neural networks (CNNs): the first and third CNNs, each with three convolutional layers; the second and fourth CNNs, each with three convolutional layers; the first and second fully connected layers, which are multilayer perceptrons with two fully connected layers, where the hidden layers employ a non-linear rectified linear units (ReLU) activation function; and the first and second output layers, which use a softmax function.

[0077] In this embodiment, the pin relationship mask of the already laid-out device, the device type mask of the already laid-out device, the pin relationship mask of the device to be laid out, and the device type mask of the device to be laid out are input into a first convolutional network for feature extraction to obtain a first feature vector. The current undirected graph is input into a second convolutional network for feature extraction to obtain a second feature vector. The first feature vector and the second feature vector are concatenated and input into a first fully connected layer for dimensionality reduction to obtain a third feature vector. The third feature vector is input into a first output layer to obtain an action probability distribution. Based on the action probability distribution, the current action is obtained. S42: The already laid-out state mask and the state mask to be laid out are input into a third convolutional network to obtain a fourth feature vector; the current undirected graph is input into a fourth convolutional network to obtain a fifth feature vector; the fourth feature vector and the fifth feature vector are input into a second fully connected layer to obtain a sixth feature vector; the sixth feature vector is input into a second output layer to obtain the current Q value.

[0078] In this embodiment, the pin relationship mask of the already placed device, the device type mask of the already placed device, the pin relationship mask of the device to be placed, and the device type mask of the device to be placed are input into the third convolutional network for feature extraction to obtain the fourth feature vector. The current undirected graph is input into the fourth convolutional network for feature extraction to obtain the fifth feature vector. The fourth and fifth feature vectors are concatenated and input into the second fully connected layer for dimensionality reduction to obtain the sixth feature vector. The sixth feature vector is input into the second output layer to obtain the current Q value.

[0079] This application embodiment, by setting specific network structures for the policy network and value network, can automatically and quickly obtain the current action and the current Q value based on the current layout state mask and the current undirected graph.

[0080] In one embodiment, step S41, which inputs the third feature vector to the first output layer to obtain the current action, includes steps S411 to S412, as follows:

[0081] S411: Input the third feature vector into the first output layer to obtain the action probability distribution.

[0082] In this embodiment, the first output layer typically uses a softmax function to convert the third feature vector output by the first fully connected layer into a probability distribution. Specifically, the softmax function can convert any real-valued vector into a probability distribution vector, where the value of each element is between 0 and 1, and the sum of the values ​​of all elements is 1. This facilitates the subsequent reinforcement learning agent in selecting actions based on the action probability distribution output by the first output layer. S412: Based on the action probability distribution, select the corresponding action from a preset action space and take the corresponding action as the current action; wherein, the preset action space includes moving upward by a preset distance, moving downward by a preset distance, moving left by a preset distance, moving right by a preset distance, and rotating by a 90-degree angle.

[0083] The preset distance can be manually set according to actual needs.

[0084] In this embodiment, the action probability distribution indicates the probability of each action in a preset action space. The reinforcement learning agent selects the action with the highest probability from the preset action space as the current action.

[0085] The action space preset in this embodiment only includes five actions: moving up a preset distance, moving down a preset distance, moving left a preset distance, moving right a preset distance, and rotating by 90 degrees. Compared with the traditional method that treats each position information in the layout area as an action, this significantly reduces the action space and greatly improves the training efficiency of the reinforcement learning agent.

[0086] In one embodiment, step S40, which determines the next layout state based on the current action and the initial layout position of the device to be laid out, includes steps S43 to S45, as follows:

[0087] S43: Based on the current action and the initial placement position of the device to be placed, obtain the next placement position of the device to be placed.

[0088] In this embodiment of the application, taking the current action as moving upward by a preset distance as an example, the device to be laid out is moved upward by a preset distance from the center point of the mask to be laid out, so as to obtain the next layout position of the device to be laid out.

[0089] S44: When the next placement position of the current device to be placed satisfies the preset placement constraints, update the already placed state mask and the current undirected graph according to the next placement position of the current device to be placed, and obtain the updated already placed state mask and the updated current undirected graph; obtain the device information of the next device to be placed of the current device to be placed; construct a new device to be placed state mask according to the device information of the next device to be placed of the current device to be placed; use the new device to be placed state mask as the updated device to be placed state mask; obtain the next placement state of the current placement state according to the updated already placed state mask, the updated current undirected graph and the updated device to be placed state mask.

[0090] Among them, the preset layout constraints are hard layout constraints, which include, but are not limited to, that devices cannot overlap, that the spacing between devices must meet a safe distance, and that the size of the heat dissipation area of ​​the devices must meet a certain area threshold.

[0091] In this embodiment, after obtaining the next placement position of the current device to be placed, it is necessary to check whether the next placement position of the current device to be placed meets the preset placement constraints. If so, the placement of the current device to be placed is completed, and the next placement position of the current device to be placed is updated in the placement state mask and the corresponding node information of the undirected graph. At the same time, the device information of the next device to be placed is read and its position is initialized, and then updated in the placement state mask, wherein the initialized position is the center point position of the blank mask.

[0092] S45: When the next placement position of the current device to be placed does not meet the preset placement constraints, update the placement state mask and the current undirected graph according to the next placement position of the current device to be placed, and obtain the updated placement state mask and the updated current undirected graph; obtain the next placement state of the current placement state according to the updated placement state mask, the updated current undirected graph and the already placed state mask.

[0093] In this embodiment, when the next placement position of the current device to be placed does not meet the preset placement constraints, the next placement position of the current device to be placed is only updated in the state mask to be placed and the corresponding node information of the undirected graph, so that the reinforcement learning agent can continue to perform the next action on the current device to be placed.

[0094] The embodiments of this application realize the transition of layout state through the above process. At the same time, integrating the hard constraints of layout into the layout state transition can effectively reduce the complexity of the reward function design of the subsequent reinforcement learning agent.

[0095] In one embodiment, the preset reward function includes a first reward function, a second reward function, a third reward function, a fourth reward function, a fifth reward function, and a sixth reward function. Step S40, which calculates the current reward value based on the next layout state of the current layout state and the preset reward function, includes steps S401 to S407, as follows:

[0096] S401: Obtain the first reward value based on the position of each pin of the device to be laid out, the position of other pins in the same network as the pins in the device to be laid out, and the first reward function.

[0097] The layout constraints in the circuit board device placement process include hard layout constraints and soft layout constraints. Soft layout constraints include, but are not limited to, shorter distances to network pins and alignment with network components. Since the hard layout constraints have already been reflected in the layout state transition, the design of the reward function only needs to consider the soft layout constraints.

[0098] In this embodiment, the expression for the first reward function is as follows:

[0099]

[0100] Where n is the total number of pins of the device to be laid out, x1 and y1 are the x and y coordinates of the pins to be laid out, and x2 and y2 are the x and y coordinates of other pins that are located on the same network as the pins of the device to be laid out.

[0101] S402: Obtain the second reward value based on the overlapping area between the current position of the device to be laid out and the already laid-out devices, the total area of ​​the current device to be laid out, and the second reward function.

[0102] In order to prevent the reinforcement learning agent from minimizing the distance between network pins and causing device overlap, an overlap reward function (second reward function) is set up to guide the reinforcement learning agent to complete the device layout.

[0103] In this embodiment of the application, the expression for the second reward function is as follows:

[0104]

[0105] Among them, s t s0 represents the overlap area between the current location of the device to be laid out and the already laid-out devices, and s0 represents the total area of ​​the device to be laid out.

[0106] S403: Obtain the third reward value based on the alignment degree between the current device to be laid out and other devices in the same network as the current device to be laid out, and the third reward function.

[0107] The third reward function is used to measure the alignment between the current device to be laid out and other devices in the same network.

[0108] In this embodiment, the device to be laid out is aligned with other devices in the same network, and the third reward function r align The given third reward value is 0. The current device to be placed is not aligned with other devices in the same network, and the third reward function r... align The given third reward value is a negative real number.

[0109] S404: Obtain the fourth reward value based on the current position of the device to be laid out, the layout area boundary of the circuit board, and the fourth reward function.

[0110] When a device's position exceeds the circuit board layout boundary, it is considered to have hit a wall, and a wall-hitting bonus (fourth bonus value) is given to reset the device's position to its position before hitting the wall.

[0111] In this embodiment of the application, when the position of the device to be laid out exceeds the boundary of the layout area of ​​the circuit board, the fourth reward function r bump The output fourth reward value is a negative real number. The fourth reward function r is applied when the current position of the device to be placed does not exceed the boundary of the board's placement area. bump The fourth reward value output is 0.

[0112] S405: Obtain the fifth reward value based on the rotation angle of the current device to be laid out and the fifth reward function.

[0113] In this embodiment, the rotation angle of the device to be laid out is obtained. The rotation angle of the device to be laid out has changed continuously by 360 degrees compared to the initial rotation angle. The fifth reward function r turn The fifth reward value is a negative real number. The rotation angle of the current placement device has not changed continuously by 360 degrees compared to the initial rotation angle. The fifth reward function r... turn The fifth reward value output is 0.

[0114] S406: Obtain the sixth reward value based on the current position of the device to be laid out, the historical best position, and the sixth reward function.

[0115] In order to encourage reinforcement learning agents to learn from the layout experience of experts (historical best positions), an expert knowledge reward function (sixth reward function) is designed.

[0116] In this embodiment, the expression for the sixth reward function is as follows:

[0117] r exper t = -(d0 - d t )

[0118] Where d0 represents the half-perimeter of the initial placement position of the device to be placed from the historical best position, d t This indicates the half-circumference of the next placement position of the device to be placed from the historical best position.

[0119] S407: Obtain the current reward value based on the first reward value, second reward value, third reward value, fourth reward value, fifth reward value, and sixth reward value.

[0120] In this embodiment, the first reward value, second reward value, third reward value, fourth reward value, fifth reward value, and sixth reward value are weighted and summed to obtain the current reward value. Specifically, the expression for the current reward value is as follows:

[0121] r=w1(n)r dist +w2r overlap +w3r align +w4r bump +w5r turn +w6r expert

[0122] Where w1 represents the first weighting coefficient, w2 represents the second weighting coefficient, w3 represents the third weighting coefficient, w4 represents the fourth weighting coefficient, w5 represents the fifth weighting coefficient, and w6 represents the sixth weighting coefficient.

[0123] Since the layout order of components affects the optimization of the first reward function, the layout order is considered in the first weight coefficient w1 of the first reward function. Meanwhile, since the preset distances corresponding to each action in the action space are consistent, while the lengths and widths of different components are inconsistent, the component sizes are considered in the second weight coefficient w2 of the second reward function.

[0124] This application's embodiments set a reward function based on layout soft constraints, and use the reward value given by the reward function to train the reinforcement learning agent. Furthermore, by introducing an expert knowledge reward function, the learning performance of the reinforcement learning agent can be improved.

[0125] The following are embodiments of the apparatus described in this application, which can be used to execute the methods described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the methods described in the embodiments of this application.

[0126] Please see Figure 2 This diagram illustrates the structure of a circuit board layout apparatus provided in an embodiment of this application. The circuit board layout apparatus 7 provided in this embodiment includes:

[0127] The circuit component information acquisition module 71 is used to acquire the circuit component information of the sample circuit board; wherein, the circuit component information includes the information of the already laid-out components and the information of the components to be laid out.

[0128] The current layout state acquisition module 72 is used to obtain the current layout state of the sample circuit board based on the circuit device information; wherein, the current layout state includes the current layout state mask and the current undirected graph;

[0129] The current action acquisition module 73 is used to input the current layout state of the sample circuit board into the reinforcement learning agent to obtain the current action and the current Q value; determine the next layout state of the current layout state based on the current action and the initial layout position of the device to be laid out; calculate the current reward value based on the next layout state of the current layout state and the preset reward function; and store the current layout state, the next layout state of the current layout state, the current action, the current reward value, and the current Q value as a set of sample data into the experience pool; wherein, the device to be laid out is a device selected from the devices to be laid out.

[0130] The reinforcement learning agent training module 74 is used to train the reinforcement learning agent based on the sample data in the experience pool when the number of sample data groups in the experience pool is greater than or equal to a preset number, so as to obtain a trained reinforcement learning agent.

[0131] The layout position acquisition module 75 is used to acquire the circuit device information of the circuit board to be laid out; based on the circuit device information of the circuit board to be laid out, the current layout state of the circuit board to be laid out is determined, and the current layout state of the circuit board to be laid out is input to the trained reinforcement learning agent to obtain the layout position of the device to be laid out in the circuit board to be laid out.

[0132] The device placement module 76 is used to place the devices to be placed according to their placement positions.

[0133] According to the embodiments of this application, the current layout state of the sample circuit board is determined based on the circuit device information. The reinforcement learning agent outputs the current action and the current Q-value based on the current layout state, obtains the current reward value, and performs a state transition to obtain the next layout state. By repeating the above process, the reinforcement learning agent continuously learns and optimizes itself using the current Q-value and the current reward value, thereby obtaining a layout result that meets the layout constraints and has a high aesthetic appeal, thus improving the layout efficiency of the circuit board.

[0134] The following are embodiments of the device described in this application, which can be used to execute the methods described in the embodiments of this application. For details not disclosed in the embodiments of the device described in this application, please refer to the methods described in the embodiments of this application.

[0135] Please see Figure 3 This application also provides an electronic device 300, which may specifically be a computer, mobile phone, tablet computer, circuit board layout device, etc. In an exemplary embodiment of this application, the electronic device 300 is a circuit board layout device, which may include: at least one processor 301, at least one memory 302, at least one display, at least one network interface 303, user interface 304, and at least one communication bus 305.

[0136] The user interface 304 is primarily used to provide an input interface for the user and to acquire user input data. Optionally, the user interface may also include a standard wired interface or a wireless interface.

[0137] The network interface 303 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0138] The communication bus 305 is used to enable communication between these components.

[0139] The processor 301 may include one or more processing cores. The processor connects to various parts of the electronic device using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory, and by calling data stored in memory. Optionally, the processor may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also be implemented as a separate chip without being integrated into the processor.

[0140] The memory 302 may include random access memory (RAM) or read-only memory. Optionally, the memory may include a non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, code, code sets, or instruction sets. The memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), instructions for implementing the various method embodiments described above, etc.; the data storage area may store data involved in the various method embodiments described above, etc. The memory may also optionally be at least one storage device located remotely from the aforementioned processor. Figure 3 As shown, a memory, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and operating applications.

[0141] The processor can be used to call the application program of the circuit board layout method stored in the memory and specifically execute the method steps of the above-described embodiments. For the specific execution process, please refer to the detailed description shown in the embodiments, which will not be repeated here.

[0142] This application also provides a computer-readable storage medium storing a computer program thereon, the instructions of which are adapted to be loaded by a processor and executed by the method steps of the embodiments shown above. For details of the execution process, please refer to the specific descriptions shown in the embodiments, which will not be repeated here. The device containing the storage medium can be an electronic device such as a personal computer, laptop computer, smartphone, or tablet computer.

[0143] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative, wherein the components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0144] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0145] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function selected in one or more boxes.

[0146] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function selected in one or more boxes.

[0147] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0148] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0149] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape storage, disk storage, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0150] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0151] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A circuit board layout method, characterized in that, The method includes the following steps: Obtain the circuit component information of the sample circuit board; wherein, the circuit component information includes the information of the already laid-out components and the information of the components to be laid out; Based on the circuit device information, the current layout state of the sample circuit board is obtained; wherein, the current layout state includes the current layout state mask and the current undirected graph; The current layout state of the sample circuit board is input into the reinforcement learning agent to obtain the current action and the current Q value; based on the current action and the initial layout position of the device to be laid out, the next layout state of the current layout state is determined; based on the next layout state of the current layout state and a preset reward function, the current reward value is calculated; the current layout state, the next layout state of the current layout state, the current action, the current reward value, and the current Q value are stored as a set of sample data in the experience pool; wherein, the device to be laid out is a device selected from the devices to be laid out. When the number of sample data sets in the experience pool is greater than or equal to a preset number, the reinforcement learning agent is trained based on the sample data in the experience pool to obtain a trained reinforcement learning agent. Obtain the circuit component information of the circuit board to be laid out; determine the current layout state of the circuit board to be laid out based on the circuit component information of the circuit board to be laid out, input the current layout state of the circuit board to be laid out into the trained reinforcement learning agent, and obtain the layout position of the components to be laid out in the circuit board to be laid out. The devices to be laid out are arranged according to the layout positions.

2. The circuit board layout method according to claim 1, characterized in that: The information on the already placed devices includes the position information and device parameters of the already placed devices; the information on the devices to be placed includes the device parameters of the devices to be placed; the current placement state mask includes the already placed state mask and the placement state mask. The step of obtaining the current layout state of the sample circuit board based on the circuit device information includes: Based on the position information and parameters of the deployed devices, the deployed state mask is constructed. Based on the initial placement position and device parameters of the device to be placed, construct the state mask to be placed; Based on the position information and parameters of the already placed devices, the initial placement position and parameters of the device to be placed, the current undirected graph is constructed.

3. The circuit board layout method according to claim 2, characterized in that: The reinforcement learning agent includes a policy network and a value network; the policy network includes a first convolutional network, a second convolutional network, a first fully connected layer, and a first output layer; the value network includes a third convolutional network, a fourth convolutional network, a second fully connected layer, and a second output layer. The step of inputting the current layout state into the reinforcement learning agent to obtain the current action and the current Q value includes: The already laid-out state mask and the state mask to be laid out are input into the first convolutional network to obtain the first feature vector; The current undirected graph is input into the second convolutional network to obtain a second feature vector; the first feature vector and the second feature vector are input into the first fully connected layer to obtain a third feature vector; the third feature vector is input into the first output layer to obtain the current action; The laid-out state mask and the state mask to be laid out are input into the third convolutional network to obtain the fourth feature vector; the current undirected graph is input into the fourth convolutional network to obtain the fifth feature vector; the fourth feature vector and the fifth feature vector are input into the second fully connected layer to obtain the sixth feature vector; the sixth feature vector is input into the second output layer to obtain the current Q value.

4. The circuit board layout method according to claim 3, characterized in that: The step of inputting the third feature vector into the first output layer to obtain the current action includes: The third feature vector is input into the first output layer to obtain the action probability distribution; According to the action probability distribution, a corresponding action is selected from a preset action space and the corresponding action is taken as the current action; wherein, the preset action space includes moving up a preset distance, moving down a preset distance, moving left a preset distance, moving right a preset distance, and rotating by 90 degrees.

5. The circuit board layout method according to claim 1, characterized in that: The step of determining the next layout state of the current layout state based on the current action and the initial layout position of the device to be laid out includes: Based on the current action and the initial layout position of the device to be laid out, the next layout position of the device to be laid out is obtained; When the next placement position of the current device to be placed satisfies the preset placement constraints, the already placed state mask and the current undirected graph are updated according to the next placement position of the current device to be placed, to obtain the updated already placed state mask and the updated current undirected graph; the device information of the next device to be placed is obtained; a new placement state mask is constructed according to the device information of the next device to be placed; the new placement state mask is used as the updated placement state mask; the next placement state of the current placement state is obtained according to the updated already placed state mask, the updated current undirected graph, and the updated placement state mask. When the next placement position of the current device to be placed does not meet the preset placement constraints, the placement state mask and the current undirected graph are updated according to the next placement position of the current device to be placed to obtain the updated placement state mask and the updated current undirected graph; the next placement state of the current placement state is obtained according to the updated placement state mask, the updated current undirected graph and the already placed state mask.

6. The circuit board layout method according to any one of claims 1 to 5, characterized in that: The preset reward functions include a first reward function, a second reward function, a third reward function, a fourth reward function, a fifth reward function, and a sixth reward function; The step of calculating the current reward value based on the next layout state of the current layout state and a preset reward function includes: A first reward value is obtained based on the position of each pin of the device to be laid out, the position of other pins in the same network as the pins of the device to be laid out, and the first reward function. The second reward value is obtained based on the overlapping area between the current position of the device to be laid out and the already laid-out devices, the total area of ​​the current device to be laid out, and the second reward function; A third reward value is obtained based on the alignment degree between the current device to be laid out and other devices in the same network as the current device to be laid out, and the third reward function; The fourth reward value is obtained based on the position of the device to be laid out, the boundary of the layout area of ​​the circuit board, and the fourth reward function; The fifth reward value is obtained based on the rotation angle of the current device to be laid out and the fifth reward function; The sixth reward value is obtained based on the current position of the device to be laid out, the historical best position, and the sixth reward function; The current reward value is obtained based on the first reward value, the second reward value, the third reward value, the fourth reward value, the fifth reward value, and the sixth reward value.

7. A circuit board layout device, characterized in that, include: A circuit component information acquisition module is used to acquire circuit component information of a sample circuit board; wherein, the circuit component information includes information on already laid-out components and information on components to be laid out. The current layout state acquisition module is used to obtain the current layout state of the sample circuit board based on the circuit device information; wherein, the current layout state includes a current layout state mask and a current undirected graph; The current action acquisition module is used to input the current layout state of the sample circuit board into the reinforcement learning agent to obtain the current action and the current Q value; determine the next layout state of the current layout state based on the current action and the initial layout position of the device to be laid out; calculate the current reward value based on the next layout state of the current layout state and a preset reward function; and store the current layout state, the next layout state of the current layout state, the current action, the current reward value, and the current Q value as a set of sample data into an experience pool; wherein, the device to be laid out is a device selected from the devices to be laid out. The reinforcement learning agent training module is used to train the reinforcement learning agent based on the sample data in the experience pool when the number of sample data sets in the experience pool is greater than or equal to a preset number, so as to obtain a trained reinforcement learning agent. The layout position acquisition module is used to acquire the circuit device information of the circuit board to be laid out; based on the circuit device information of the circuit board to be laid out, the current layout state of the circuit board to be laid out is determined, and the current layout state of the circuit board to be laid out is input to the trained reinforcement learning agent to obtain the layout position of the device to be laid out in the circuit board to be laid out. The device layout module is used to lay out the device to be laid out according to the layout position.

8. An electronic device, characterized in that, include: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the circuit board layout method as described in any one of claims 1 to 6.