Method, system and product for guiding EDA automatic wiring based on reinforcement learning global wiring

By introducing deep reinforcement learning and a global routing scheme into the A* algorithm, the problem of local optima in PCB routing of the A* algorithm is solved, and more efficient global routing optimization is achieved to meet complex PCB design rules.

CN121480433APending Publication Date: 2026-02-06WUHAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511506618.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing A* algorithms cannot effectively incorporate global information in PCB routing, leading to local optima problems and failing to meet the design rules and constraints of complex PCB routing.

Method used

A deep reinforcement learning algorithm is used for global routing, and its results are integrated into the A* search algorithm. The global routing scheme is optimized through a deep Q network, and the multi-pin problem is decomposed by combining the minimum spanning tree algorithm. An extended cost function is used to consider the global routing cost factor.

Benefits of technology

It effectively avoids local optima problems, improves the efficiency and accuracy of automatic PCB routing, and meets the design rules and constraints of complex PCB routing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121480433A_ABST
    Figure CN121480433A_ABST
Patent Text Reader

Abstract

The invention discloses a reinforcement learning-based global wiring guiding EDA automatic wiring method, system and product, and the method comprises the steps: firstly reading in and analyzing a PCB layout file based on a deep reinforcement learning wiring environment, dividing a wiring region into n layers of M * M global wiring grids by adopting a uniform wiring region division mode, and n is the number of PCB layers defined by the PCB layout file; meanwhile, reading in and analyzing a PCB layout file based on an A * wiring environment, and initializing a wiring grid; then based on a deep reinforcement learning wiring environment, decomposing a multi-pin wiring problem into several groups of double-pin problems, and then performing global wiring by using a deep Q network to obtain a final global wiring scheme; and finally, inputting the global wiring scheme into an A * algorithm PCB wiring device to realize automatic wiring. According to the method, before PCB automatic wiring is carried out, global wiring is carried out by using a deep reinforcement learning algorithm, and a result is integrated into an A * search algorithm, so that global information is considered in a local wiring stage, and a local optimum problem is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of integrated PCB routing technology, and relates to an EDA automatic routing method, system and product, specifically to an EDA automatic routing method, system and product based on reinforcement learning-guided global routing. Background Technology

[0002] With the rapid development of intelligent electronic devices, the pin density and signal complexity of printed circuit boards (PCBs) are increasing daily. PCB routing refers to the electrical connection of various devices and chips on the circuit board using wires, while meeting various design rules and constraints. Due to the complexity of modern PCBs, manual routing has become a time-consuming and labor-intensive task; therefore, automatic routing technology has gradually become a key step in PCB design.

[0003] Automatic routing in PCBs is often modeled as a maze-finding problem. The entire circuit is mapped as a maze, and the modules, standard cells, and input / output interfaces distributed within the chip core are interconnected according to logical relationships. This is equivalent to finding a path given the start and end points of the maze.

[0004] The A* algorithm is a heuristic search algorithm commonly used for path finding on a graphical plane. It builds upon breadth-first search by incorporating heuristic functions to increase search efficiency.

[0005] Breadth-first search (BFS) is a graph traversal algorithm that starts with a given vertex and explores all adjacent unvisited vertices in all directions, visiting them each time. Then, for each visited vertex, it searches for its unvisited neighbors. This process expands outwards layer by layer until all vertices in the graph have been visited.

[0006] In certain pathfinding scenarios, different types of movement have different movement costs, making Dijkstra's algorithm more suitable. Dijkstra's algorithm favors lower-cost paths rather than exploring all possible paths equally. In this algorithm, the actual movement costs from the starting point are used to prioritize the path, thus encouraging movement along certain routes by assigning lower costs and avoiding obstacles by assigning higher costs. Dijkstra's algorithm can find the shortest path well, but it spends a lot of time exploring useless directions. In contrast, the greedy best-first search algorithm uses the estimated distance to the target for prioritization, exploring the closest locations first, but may ultimately fail to find the shortest path.

[0007] The A* algorithm is a combined implementation of Dijkstra's algorithm and the greedy best-first search algorithm, using both the actual movement cost from the starting point and the estimated distance to the goal. Like Dijkstra's algorithm, A* guarantees finding the shortest path; however, its optimal search strategy eliminates most of the solution space processed by Dijkstra's algorithm, significantly improving Dijkstra's running time. The A* algorithm defines a cost function. .in, Indicates the distance from the starting point to the node. The actual cost of the optimal path, Indicates from node The actual cost of the optimal path to the destination. This represents the node The value of exploration The smaller the value, the more the A* algorithm tends to explore that node. It can be seen that when... When, the A* algorithm degenerates into Dijkstra's algorithm, when At that time, the A* algorithm degenerates into a greedy optimal finite search.

[0008] The existing A* algorithm has the following two main problems: (1) The simple cost function of the A* algorithm can solve the simple and abstract maze search problem. However, in complex PCB routing, multiple design rules and constraints need to be considered. A single cost function cannot meet the requirements and needs to be extended and improved to cope with the complexity of design rules and the challenges of practical applications.

[0009] (2) The A* algorithm selects paths based on heuristic evaluation of the current node, which leads it to tend to choose the path that appears optimal under the current circumstances. In other words, the A* algorithm can only solve local optima and cannot perceive global information. Therefore, preceding routing can affect subsequent routing. In complex PCB routing problems, the preceding routing may be the optimal path selected by the A* algorithm during its own routing process, but subsequent routing may fail to complete because the preceding routing occupies routing resources. Summary of the Invention

[0010] To address the aforementioned technical problems, this invention provides a reinforcement learning-based global routing-guided EDA automatic routing method, system, and product. It innovatively proposes using a deep reinforcement learning algorithm for global routing before performing automatic PCB routing, and integrating the results into the A* search algorithm. This allows global information to be considered during the local routing stage, avoiding local optima problems.

[0011] The technical solution adopted by the method of the present invention is: an automatic routing method for EDA based on reinforcement learning global routing guidance, comprising the following steps: Step 1: Based on the deep reinforcement learning routing environment, read and parse the PCB layout file, and use the uniform routing area division method to divide the routing area into an n-layer M×M global routing grid, where n is the number of PCB layers defined in the PCB layout file and M is a preset value. Simultaneously, the PCB layout file is read and parsed based on the A* routing environment to initialize the routing mesh; Step 2: Based on the deep reinforcement learning routing environment, the multi-pin routing problem is decomposed into several groups of two-pin problems, and then a deep Q-network is used for global routing. The deep reinforcement learning routing environment provides reward feedback and corresponding routing environment state changes to the deep Q-network, and updates the environment to provide the next round of rewards and states after the deep Q-network router gives an action. In the deep Q-network router, the routing scheme with the highest total reward value after routing all two-pin problems is taken as the final global routing scheme. Step 3: Input the global routing scheme into the A* algorithm PCB router to achieve automatic routing. The preceding routing path will be added to the A* routing environment and become a routed obstacle area in the subsequent routing process. After multiple iterations, the automatic routing task of all pins will be completed.

[0012] Preferably, in step 1, the routing grid is initialized with a square grid unit length of one-tenth of the smallest coordinate unit in the PCB layout file.

[0013] Preferably, in step 2, the minimum spanning tree algorithm is used to decompose the multi-pin wiring problem into several groups of two-pin problems; The minimum spanning tree algorithm uses each pin as a vertex and the connection between different pins as edges. It divides all vertices into connected and unconnected ones and constructs a traversal list. At the beginning, any vertex is added to the traversal list. In each iteration, all points that are connected to the vertices in the current list are traversed, the point with the shortest distance is selected, added to the traversal list, and the edge between the two vertices is added to the minimum spanning tree. Subsequently, the algorithm iteratively searches for nodes in the original connected graph until the minimum spanning tree is constructed.

[0014] Preferably, in step 2, the deep Q-network consists of a multilayer perceptron containing three fully connected layers, each with 32, 64, and 32 hidden units respectively; each layer is followed by a ReLU activation layer; the input vector of the deep Q-network is 12, corresponding to the 12 components in the state vector, and the output is 6, corresponding to the six actions in the action space.

[0015] Preferably, in step 2, the state s is a 12-dimensional vector. The first three components are the x, y, and z coordinates of the agent's current position in the environment. The fourth to sixth components encode the distance from the current position to the target pin position in the x, y, and z directions. The remaining six components encode the capacity information of all the edges that the agent will traverse. The action a , represented by integers from 0 to 5, corresponding to the direction of movement from the current state; The reward R is the selected action. a and the next state Functions: ; in, Indicate the next state The agent has reached the target pin location; each action that fails to reach the target pin reduces the cumulative reward. Furthermore, the maximum number of steps the agent can take to solve each pair of two-pin wiring problems is limited by the problem size. Therefore, when successfully routing a two-pin problem, the cumulative reward is always in... and Between, otherwise The cumulative reward for the global routing solution is obtained by summing up the cumulative rewards for each group of two-pin problems, which distinguishes whether the global routing problem has been successfully solved or whether no feasible solution can be found.

[0016] Preferably, in step 2, the global wiring using a deep Q-network involves using a deep Q-network to approximate the Q-function, taking the state s as input and outputting the action value function Q for all actions in that state. Specifically, for each two-pin problem, the environment provides state information to the network. Since there may be six actions, the agent of the deep Q-network evaluates all Q-values ​​for the next state. Then, based on the ε-greedy policy, an action is selected and executed; the agent's action causes a change in the environment state, the environment records the capacity information of the change, and the agent receives a reward; all state transitions, agent actions and rewards are stored in the experience buffer, data is sampled in the experience buffer, and finally the loss function is calculated to update the network parameters; the environment and agent continuously repeat the above loop for iteration.

[0017] Preferably, the Greedy strategy ; where random action means that the action is selected randomly, ε is a hyperparameter of the deep Q network representing the probability of selecting a random action, A represents the set of all possible actions, i.e., the action space, and Q(s,a) represents the action value function Q corresponding to taking action a in the current state s.

[0018] Preferably, in step 3, the cost function of the A* algorithm PCB router is: f(x)=g(x)+h(x), where g(x) refers to the actual cost from the starting point to the current position X, and h(x) refers to the estimated cost from the current position X to the end point T; g(x)=C layer +C wl +C bend +C global +C obstacle ,in The cost of routing on a specific layer; It consists of two parts: one part is the routing cost, which is the cost of each step of mesh routing; the other part is the design cost of via layer replacement routing. Cost of routing at corners; This incurs the cost of global routing. The obstacle cost of the grid where the current position X is located is used to determine whether the current position X is an obstacle area; all five costs are cumulative costs, that is, the cumulative value of the grid routing cost at each step from the starting point to the current position X; , where minD=min(|Xx-Tx|,|Xy-Ty| ), maxD=max(|Xx-Tx|,|Xy-Ty| ), Xx and Xy are the x and y coordinates of the current position X, respectively, and Tx and Ty are the x and y coordinates of the endpoint T, respectively.

[0019] The technical solution adopted by the system of the present invention is: an automatic routing system for global routing guidance (EDA) based on reinforcement learning, comprising: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, enable the one or more processors to implement the reinforcement learning-based global routing guidance EDA automatic routing method.

[0020] The technical solution adopted by the product of the present invention is: an automatic routing product based on reinforcement learning global routing guidance EDA, including computer program instructions, which, when the computer program instructions are run on a computer, cause the computer to execute the automatic routing method based on reinforcement learning global routing guidance EDA.

[0021] Compared with the prior art, the beneficial effects of the present invention include: (1) In this invention, before performing automatic PCB routing, a deep reinforcement learning algorithm is used to perform global routing, and then the global routing scheme is input into the A* algorithm PCB router, so that the actual cost function of the A* algorithm is increased by the global routing cost factor. Therefore, in the process of automatic PCB routing, global information can be combined to avoid getting trapped in local optima.

[0022] (2) In the case of PCB automatic routing, the present invention extends the actual cost function and estimated cost function of the A* algorithm according to the design rules of PCB automatic routing. The layered design, through-hole design, minimum corner, bypass obstacle avoidance and 135° routing corner are realized by layer routing cost factor, routing cost factor, routing corner cost factor, obstacle cost factor and estimated cost respectively. Attached Figure Description

[0023] The technical solutions of the present invention will be further illustrated below using embodiments and specific implementation methods. In addition, some accompanying drawings are used in the description of the technical solutions. Those skilled in the art can obtain other drawings and the intent of the present invention from these drawings without any creative effort.

[0024] Figure 1 This is a schematic diagram of the method according to an embodiment of the present invention; Figure 2 This invention presents a simulated two-layer (4×4×2) environment and a global routing scheme to solve the dual-pin problem. Detailed Implementation

[0025] To facilitate understanding and implementation of the present invention by those skilled in the art, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0026] Please see Figure 1 This embodiment provides an automatic routing method for EDA based on reinforcement learning-guided global routing, which includes the following steps: Step 1: Based on the deep reinforcement learning routing environment, read and parse the PCB layout file, and use a uniform routing area division method to divide the routing area into an n-layer M×M global routing grid, where n is the number of PCB layers defined in the PCB layout file; M can be selected in three sizes: 8, 16, and 32. When the number of pins to be routed in the PCB layout file is less than 64, M is selected as 8; when it is greater than or equal to 64 but less than 256, M is selected as 16; when it is greater than or equal to 256, M is selected as 32.

[0027] At the same time, the PCB layout file is read and parsed based on the A* routing environment, and the routing grid is initialized with one-tenth of the smallest coordinate unit of the PCB layout file as the square grid unit length. Step 2: Based on the deep reinforcement learning routing environment, the minimum spanning tree algorithm is used to decompose the multi-pin routing problem into several groups of two-pin problems. Then, a deep Q-network (DQN) is used for global routing. The deep reinforcement learning routing environment provides reward feedback and corresponding routing environment state changes to the deep Q-network, and waits for the deep Q-network router to give an action to update the environment to provide the reward and state for the next round. In the deep Q-network router, the routing scheme with the highest total reward value after routing all two-pin problems is taken as the final global routing scheme. In one implementation, the minimum spanning tree algorithm uses each pin as a vertex and the connections between different pins as edges. It categorizes all vertices into connected and disconnected, constructing a traversal list. Initially, any vertex is added to the traversal list. In each iteration, all points connected to vertices in the current list are traversed, the point with the shortest distance is selected, added to the traversal list, and the edge between the two vertices is added to the minimum spanning tree. Subsequently, the algorithm iteratively searches for nodes in the original connected graph until the minimum spanning tree is constructed.

[0028] In one implementation, the deep Q-network consists of a multi-layer perceptron (MLP) containing three fully connected layers, each with 32, 64, and 32 hidden units, respectively. Each layer is followed by a ReLU activation layer. The Q-network has an input vector size of 12, corresponding to the 12 components in the state vector, and an output size of 6, corresponding to the six actions in the action space.

[0029] The network's hyperparameters include batch size, learning rate, and network architecture, which are deep learning parameters. The target network update step size, discount factor γ, and greedy policy ε are also included. The network's hyperparameters are shown in Table 1 below: Table 1

[0030] In one implementation, the deep reinforcement learning wiring environment provides reward feedback and corresponding wiring environment state changes to the DQN algorithm, and updates the environment to provide rewards and states for the next round after the DQN router gives an action. Figure 2This diagram illustrates a simulated two-layer (4×4×2) environment and the global routing scheme used in this embodiment to address the dual-pin issue. Each layer contains 16 cells. By manually setting the bold black edges to have zero capacity, all vertical edges in the first layer have zero capacity, allowing only vertical routing. Similarly, all horizontal edges in the second layer have zero capacity, allowing only horizontal routing. Layers are connected via vias. This design significantly reduces routing complexity. White modules represent edge capacities of zero due to pre-routed lines or layout chip modules. The diagram provides a possible routing scheme under these edge capacity constraints.

[0031] In the Deep Q-Network algorithm, the agent, starting from a cell in this environment, has six possible actions. There are four actions in a plane: vertical (up / down) and horizontal (left / right). Between different planes, there are two actions: from the first layer to the second layer and from the second layer to the first layer. The DQN algorithm uses a deep neural network (called a Q-network in DQN) to approximate the Q-function. It takes a state `s` as input and outputs the action value function `Q` for all actions in that state. Specifically, for each two-pin problem, the environment provides state information to the network. Since there are potentially six actions, the agent evaluates all Q-values ​​for the next state. Then, based on an ε-greedy policy, an action is selected and executed. The agent's action causes a state transition in the environment, the environment records the capacity information of the change, and the agent receives a reward. All state transitions, agent actions, and rewards are stored in an experience buffer. Data is sampled in the experience buffer, and finally, the loss function is calculated to update the network parameters. The environment and agent continuously repeat the above loop for iteration. The detailed design is as follows: State: State s is defined as a 12-dimensional vector. The first three components are the x, y, z coordinates of the agent's current position in the environment. The fourth to sixth components encode the distances from the current position to the target pin position in the x, y, z directions. The remaining six components encode the capacity information of all edges that the agent will traverse.

[0032] Action: Action a Represented by integers from 0 to 5, corresponding to the direction of movement from the current state.

[0033] Reward: The reward R is defined as a function of the chosen action and the next state.

[0034] in, Indicate the next state The agent has already reached the target pin location; each action that fails to reach the target pin reduces the cumulative reward. This design allows the agent to learn the shortest path possible. Furthermore, the maximum number of steps the agent can take to solve each pair of two-pin wiring problems is limited by the size of the problem to be solved. , This can be adjusted; the default value is 50. Therefore, when successfully routing a two-pin problem, the cumulative reward is always... and Between, otherwise The cumulative reward for the global routing solution is obtained by summing up the cumulative rewards for each group of two-pin problems, which distinguishes whether the global routing problem has been successfully solved or whether no feasible solution can be found.

[0035] Greedy strategy ): When selecting an action, take Greedy strategy: .

[0036] Where random action means that the action is selected randomly, ε is a hyperparameter of the deep Q network representing the probability of selecting a random action, A represents the set of all possible actions, i.e., the action space, and Q(s,a) represents the action value function Q corresponding to taking action a in the current state s.

[0037] Step 3: Input the global routing scheme into the A* algorithm PCB router to achieve automatic routing. The preceding routing paths will be added to the A* routing environment and become the already routed obstacle areas in the subsequent routing process. After multiple iterations, the automatic routing task of all pins is finally completed. This adds a global routing cost factor to the actual cost function of the A* algorithm, thus combining global information during the PCB automatic routing process to avoid getting trapped in local optima.

[0038] In one implementation, the cost function of the A* algorithm PCB router can be defined as follows: ,in This refers to the actual cost of traveling from the starting point to the current position X. This refers to the estimated cost from the current position X to the destination T.

[0039] Based on the rules of PCB routing obstacle avoidance, via design, and layered design, the actual cost is designed. ,in The cost of routing on a specific layer. For example, the cost of routing a power / ground line on the power / ground layer is 1, while the cost of routing it on a normal signal layer is 100, and vice versa for normal signal lines. It consists of two parts: one part is the routing cost, with each step of mesh routing costing 1, and the other part is the design via layer replacement routing cost of 10. The cost for routing corners is 1.4 per corner. The global routing cost is 0 if the current pin's PCB routing grid can be mapped to the global routing grid corresponding to the DQN global routing scheme; otherwise, it is 2. Additionally... The obstacle cost is calculated for the mesh at the current position X. It is determined whether position X is an obstacle region: if it is a pad, the obstacle cost is 1000; if it is a via, the obstacle cost is 100; if it is a routed region, the obstacle cost is 50; and if it is a routeable region, the obstacle cost is 0. All five cost functions are cumulative costs, representing the cumulative mesh routing cost from the starting point to the current position X.

[0040] Based on the 135° routing angle rule for PCB routing, the estimated cost is designed. ,in , and These are the x and y coordinates of the current position X, respectively. and These are the x and y coordinates of the endpoint T, respectively.

[0041] It should be noted that this invention completes the reinforcement learning environment modeling for global routing. The DQN algorithm used by the DQN router can be replaced by other deep reinforcement learning algorithms such as DDQN, DeulingDQN, A3C, etc.

[0042] It should be understood that the embodiments described above are only some, not all, of the embodiments of the present invention. Furthermore, the technical features of the various embodiments or individual embodiments provided by the present invention can be arbitrarily combined to form feasible technical solutions. Such combinations are not constrained by the order of steps and / or structural composition patterns, but must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0043] It should be understood that the above description of the preferred embodiments is quite detailed, but it should not be considered as a limitation on the scope of protection of this invention. Those skilled in the art, under the guidance of this invention, can make substitutions or modifications without departing from the scope of protection of the claims of this invention, and all such substitutions or modifications fall within the scope of protection of this invention. The scope of protection of this invention should be determined by the appended claims.

Claims

1. A global routing guide EDA automatic routing method based on reinforcement learning, characterized in that, The method comprises the following steps: Step 1: reading and parsing the PCB layout file based on the deep reinforcement learning wiring environment, dividing the wiring area into n layers of MxM global wiring grids in a uniform wiring area division manner, n being the number of PCB layers defined by the PCB layout file, and M being a preset value; At the same time, based on the A* wiring environment, the PCB layout file is read and parsed, and the wiring grid is initialized; Step 2: based on the deep reinforcement learning wiring environment, the multi-pin wiring problem is decomposed into several groups of double-pin problems, and then the deep Q network is used for global wiring; wherein the deep reinforcement learning wiring environment provides reward feedback and corresponding state transformation of the wiring environment to the deep Q network, and updates the environment after the deep Q network wiring device gives an action to provide the next round of reward and state; In the deep Q network wiring device, the total reward value of the wiring scheme with the highest total reward value after all double-pin problems are wired is taken as the final global wiring scheme; Step 3: inputting the global wiring scheme into the A* algorithm PCB wiring device to realize automatic wiring, and adding the pre-wiring path into the A* wiring environment to become an already-wired obstacle area in the subsequent wiring process, and performing multiple iterations to finally complete the automatic wiring task of all pins.

2. The global routing guidance EDA automatic routing method based on reinforcement learning according to claim 1, characterized in that: In step 1, the wiring grid is initialized by taking one-tenth of the minimum unit of the coordinates of the PCB layout file as the square grid unit length.

3. The global routing guidance EDA auto-routing method based on reinforcement learning of claim 1, wherein: In step 2, the minimum spanning tree algorithm is used to decompose the multi-pin wiring problem into several groups of double-pin problems; The minimum spanning tree algorithm takes each pin as a vertex and the connection between different pins as an edge, divides all vertices into connected and unconnected, and constructs a traversal list, and initially adds any vertex to the traversal list; In each round of iteration, all points connected to the vertices in the current list are traversed, the shortest point is selected, added to the traversal list, and the edge between the two vertices is added to the minimum spanning tree; Then, the nodes in the original connected graph are iteratively searched until the construction of the minimum spanning tree is completed.

4. The global routing guidance EDA auto-routing method based on reinforcement learning of claim 1, wherein: In step 2, the deep Q network consists of a multilayer perceptron, including three fully connected layers, each layer being provided with 32, 64 and 32 hidden units respectively; each layer is provided with a ReLU activation layer after it; the input vector size of the deep Q network is 12, corresponding to 12 components in the state vector, and the output size is 6, corresponding to six actions in the action space.

5. The global routing guidance EDA auto-routing method based on reinforcement learning of claim 1, wherein: In step 2, the state s is a 12-dimensional vector, the first three components are the x, y and z coordinates of the current position of the agent in the environment, the fourth to sixth components encode the distances in x, y and z directions from the current position to the target pin position, and the remaining six components encode the capacity information of all edges to be traversed by the agent; The action a with an integer from 0 to 5, corresponding to the direction of movement from the current state; The reward R, as a function of the selected action a and the next state . ; wherein, represents the next state the current position of the agent in the maze has reached the target pin position; each action that does not move the agent to the target pin reduces the cumulative reward, while limiting the maximum number of steps the agent can take to solve each set of double pin placement problems according to the size of the problem to be solved Thus, when a double pin placement problem is successfully solved, the cumulative reward is always and otherwise The cumulative reward for each set of double pin placement problems is summed to obtain the cumulative reward for the global placement solution, which is used to distinguish whether the global placement problem is successfully solved or whether a feasible solution cannot be found.

6. The global routing guidance EDA auto-routing method based on reinforcement learning of claim 1, wherein: In step 2, the global routing using the deep Q network is to use the deep Q network to approximate the fitting of the Q function, input state s, and output the action value function Q of all actions in the state; specifically, for each two-pin problem, the environment provides the state information to the network; since there can be six actions, the deep Q network agent evaluates all Q values of the next state ( ), and then selects and executes an action based on the ε-greedy strategy; the agent action causes the state of the environment to change, and the environment records the changed capacity information, and the agent gets the reward; all state transitions, agent actions and rewards are stored in the experience buffer, data is sampled in the experience buffer, and finally the loss function is calculated to update the network parameters; the environment and the agent repeatedly iterate the above loop.

7. The global routing guidance EDA auto-routing method based on reinforcement learning of claim 6, wherein: The Greed policy ; wherein random action indicates that the selected action is selected in a random manner, ε is a deep Q network hyperparameter indicating the probability of selecting a random action, A represents a set of all possible actions, i.e. an action space, and Q(s, a) represents an action value function Q corresponding to the current state s and the action a.

8. The global routing guidance EDA auto-routing method based on reinforcement learning of any one of claims 1-7, wherein: In step 3, the cost function of the A* algorithm PCB wiring device is f(x) = g(x) + h(x), wherein g(x) refers to the actual cost from the starting point to the current position X, and h(x) refers to the estimated cost from the current position X to the terminal T; g(x) = C layer + C wl + C bend + C global + C obstacle wherein is the cost of routing on a particular layer; contains two parts, one part is the routing cost, the cost of routing each step of the grid, the other part is the design of the via layer routing cost; is the corner cost of routing; is the global routing cost; is the obstacle cost of the grid where the current position X is located, to determine whether the current position X is an obstacle region; All the above five costs are cumulative costs, that is, the cumulative value of the grid wiring cost of each step from the starting point to the current position X; where minD = min(|X.x - T.x|, |X.y - T.y|), maxD = max(|X.x - T.x|, |X.y - T.y|), X.x and X.y are the horizontal and vertical coordinates of the current position X, and T.x and T.y are the horizontal and vertical coordinates of the end point T, respectively.

9. A reinforcement learning based global routing guided EDA automatic routing system, characterized in that, It comprises: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the reinforcement learning based global routing guided EDA auto-routing method according to any one of claims 1 to 8.

10. A global routing guidance EDA automatic routing product based on reinforcement learning, comprising computer program instructions, characterized in that: The computer program instructions, when running on a computer, cause the computer to perform the reinforcement learning based global routing guided EDA auto-routing method according to any one of claims 1 to 8.