Surface mounting path optimization method based on deep reinforcement learning

Through deep reinforcement learning technology, self-attention mechanism and recurrent neural network are used to optimize the mounting path, which solves the problem that existing algorithms rely on manual design rules and achieves more efficient mounting path planning.

CN119997497APending Publication Date: 2025-05-13HARBIN INST OF TECH

Patent Information

Application Number
CN202510312728.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing mounting path optimization algorithm relies on heuristic rules of manual design, and the data feature extraction capability is weak, resulting in low production efficiency.

Method used

Using a method based on deep reinforcement learning, high-dimensional data feature extraction is performed through a self-attention mechanism and a combined mask encoder, combined with a recurrent neural network and an attention mechanism decoder outputs the number allocation results of the patch head-mount nodes, and the mounting path is determined through dynamic planning.

Benefits of technology

It significantly shortens the mounting path, improves the production efficiency of the patch machine, and reduces the dependence on manual design rules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119997497A_ABST
    Figure CN119997497A_ABST
Patent Text Reader

Abstract

The invention discloses a surface mounting path optimization method based on deep reinforcement learning, solves the problem that the production efficiency of a chip mounter is affected due to the fact that the data feature extraction capability is weak by adopting an artificial design rule in an existing mounting path optimization method, and belongs to the field of electric appliance technologies and electrical engineering. The method comprises the following steps: acquiring chip mounter parameters and circuit board production data, and constructing a chip mounting node candidate node set; taking the position of each node and whether the node is mounted as input, and performing high-dimensional data feature extraction by using an encoder based on a self-attention mechanism and a combined mask to obtain node embedding; the node embedding is used as input, and a decoder based on a recurrent neural network and an attention mechanism is used for outputting a chip head-mounting node serial number distribution result; an encoder and a decoder form a strategy network; and according to the output of the trained strategy network, a dynamic planning method is used to determine the sequence of mounting elements in each pick-and-mount period, and a final mounting path optimization result is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a surface mounting path optimization method based on deep reinforcement learning, and belongs to the field of electrical technology and electrical engineering. Background Art

[0002] With the improvement of computing power and the popularization of big data technology, artificial intelligence has become the core driving force for industrial development. Through technologies such as machine learning, computer vision, and natural language processing, AI can optimize production processes, predict equipment failures, improve energy efficiency, and make intelligent decisions. This technical background provides technical support for the implementation of intelligent manufacturing and promotes the transformation and upgrading of traditional industries towards efficiency, intelligence, and greenness. With the development trend of intelligence and networking, autonomous decision-making capabilities based on deep reinforcement learning (DRL) are increasingly valued in industrial production. In industrial production, surface mount equipment is the core of printed circuit board assembly production lines. Figure 1 The common structure of the placement machine on the current market is demonstrated, which adopts a gantry-type three-dimensional motion platform. In the transport carrier of the placement machine, multiple placement heads with dual functions of suction and blowing are evenly distributed. These placement heads are driven by the motor inside the carrier and can move up and down along the Z axis. The transport carrier moves along the X and Y axes in the horizontal plane driven by the lateral and longitudinal guides. Before the pick and place operation begins, the circuit board will be transported to the predetermined position and firmly fixed by the conveyor belt and the clamp. During the pick and place process, the feeder fixed on the machine base will automatically replenish and supply components. The feeder will push the components to be picked up to the pick point at the front end, waiting for the placement head to pick up and use them. Usually, the number of components that need to be mounted on the printed circuit board is much larger than the number of placement heads used to pick and place components. Therefore, it often takes multiple pick and place cycles to complete the placement of a circuit board, hereinafter referred to as the "pick and place cycle". Pick and place are the main processes in circuit board placement. How to optimize the pick and place process is the key to improving the efficiency of circuit board placement. Figure 2 Shown is a flow chart of the pick-and-place process of the placement machine.

[0003] Optimizing the pick-and-place process of the placement machine can significantly shorten the time consumed in surface mounting. This optimization problem includes two main sub-problems: component allocation problem and placement site selection path planning problem. At present, the component allocation problem can be solved by using the algorithms provided in steps one and two of the patent CN202310039384.7 based on hierarchical heuristics for optimizing the surface mounting process of in-line placement machines. The path planning problem within the pick-and-place cycle can usually be regarded as a small-scale traveling salesman problem, which can be solved by using step three of the patent CN202010388770.3 based on clustering for optimizing the pick-and-place path of LED placement machines.

[0004] DRL has significant advantages in combinatorial optimization and path optimization problems. First, it can adaptively optimize in complex environments without explicit modeling, and is suitable for nonlinear and multi-constrained problems. Second, DRL learns decision strategies end-to-end, avoiding the tedious feature engineering of traditional heuristic algorithms. Compared with traditional methods that are prone to computational bottlenecks in large-scale search spaces, DRL can efficiently explore the optimal solution, which is particularly suitable for optimization tasks such as the traveling salesman problem (TSP) and the vehicle routing problem (VRP). In addition, DRL combined with graph neural networks (GNN) or Transformer can also enhance the feature representation capability and capture the complex structural relationships in optimization problems. However, there are still many difficulties in applying DRL to the optimization of surface mounting paths of placement machines. First, the surface mounting path optimization problem of placement machines is a special VRP problem, and each mounting point has attribute constraints of the component type; second, as a carrier (mounting head) for transporting mounting components, it cannot be abstracted into a particle during the path optimization process. Instead, it is necessary to reasonably use the spacing between the mounting heads to further optimize the mounting motion path. Finally, in addition to horizontal movement, the placement of each mounting point also needs to consider the mounting angle and other issues.

[0005] The main defects of current research are: most of the existing placement path optimization algorithms rely on manually designed heuristic rules, which have weak data feature extraction capabilities, are heavily dependent on domain knowledge and expert experience, and have high algorithm development costs. In addition, the placement path optimization algorithms designed for specific placement point distribution have weak migration capabilities and poor flexibility. Summary of the invention

[0006] In view of the problem that the existing mounting path optimization method adopts manual design rules with weak ability to extract data features, which affects the production efficiency of the mounting machine, the present invention provides a surface mounting path optimization method based on deep reinforcement learning.

[0007] A surface mounting path optimization method based on deep reinforcement learning of the present invention comprises:

[0008] S1, obtain the parameters of the placement machine and the circuit board production data, and build a candidate node set of placement nodes, including placement nodes and virtual warehouse nodes depot;

[0009] S2, taking the position information and mounting requirements of each node in the candidate node set as input, using an encoder based on self-attention mechanism and combination mask to extract high-dimensional data features, and output node embedding; the mounting node requirement is whether to mount;

[0010] S3, taking the node embedding of S2 as input, uses a decoder based on a recurrent neural network and an attention mechanism to output the patch header-mounting node number assignment result;

[0011] S4, training a policy network, where the policy network includes the encoder and the decoder;

[0012] S5. According to the placement head-placement node number allocation result of each pick-up and placement cycle output by the trained strategy network, the order of placing components in each pick-up and placement cycle is determined using a dynamic programming method to obtain the final placement path optimization result.

[0013] Preferably, S2 comprises:

[0014] S21, linearly map the two-dimensional coordinate position of each node in the candidate node set and the mounting node requirements to a high-dimensional vector space represents each node of the high-dimensional vector space, M represents the number of mounting nodes, and when i=0, the corresponding node is the virtual warehouse node depot;

[0015] S22, combining the components to be mounted in the current pick-and-place cycle with the serial numbers of the completed mounting nodes, generating an encoder mask Encoder_mask; Encoder_mask is used to filter irrelevant mounting nodes in the current pick-and-place cycle;

[0016] S23, initializing the network layer index l of the encoder = 0, where L is the total number of network layers of the encoder;

[0017] S24, determine whether l≤L, if so, execute S25; otherwise, execute S28;

[0018] S25. Execution MHA is the multi-head attention sublayer, and BN is the batch normalization operation;

[0019] S26, according to the obtained implement Where FF represents the feed-forward sublayer;

[0020] S27, update l=l+1, and return to S24;

[0021] S28, output node embedding H L .

[0022] Preferably, S3 includes:

[0023] S31, determine whether all mounting nodes have been assigned; if so, exit the decoder and output the patch header-mounting node sequence number assignment result Π; otherwise, execute S32;

[0024] S32, determine the node embedding H L The currently selected node π tIs it 0 and t≠1? If so, enter the next picking cycle, execute S2 to re-encode, and update the node embedding H L , r t Update to the number of placement nodes for the next pick-up and placement cycle, and then execute S33; otherwise, directly execute S33;

[0025] S33, according to the currently selected node π t The number of components remaining in the current pick-and-place cycle r t , construct a new vector h c :

[0026]

[0027] in, Represents node π t The embedding vector of , T represents the length of the allocation result sequence;

[0028] S34, vector h c After linear mapping, it is used as the input of the recurrent neural network, and the output of the recurrent neural network is q g =GRU(h c W G ), where W G is the learnable network weight parameter, GRU represents a recurrent neural network;

[0029] S35, use attention mechanism to obtain q g Related information about the currently unassigned placement nodes c ;

[0030] S36, q c As the context embedding information in the decoding process, the probability distribution of candidate placement nodes is calculated Where E = H L W K ,W K represents the learnable network weight parameter, C is a constant, tanh is the tangent function, For the decoding process, the mask is used to shield the already assigned mounting nodes, and softmax represents the normalized exponential function;

[0031] S37, according to Select the next node π for the placement head among the unassigned placement nodes t+1 ;

[0032] S38, update the decoding process mask according to the assigned mounting node Return to S31.

[0033] Preferably, S4 includes:

[0034] S41, input maximum number of iterations Z, batch size X, significance level parameter a, policy network with trainable parameters θ, B The baseline network is the mirror network of the policy network.

[0035] S42. Initialize the parameters θ and θ in the policy network and the baseline network B ;

[0036] S43. Obtaining instances from the training dataset

[0037] S44. Calculate the probability p based on the parameter θ in the policy network θ (Π i |s i ), and then by probability p θ (Π i |s i ) is used to generate the distribution result Π by probability sampling i ={π t ,t=1,...,T} i ;

[0038] S45. According to the parameters θ in the baseline network B Calculating Probability Then pass Greedy sampling of the distribution generates the distribution result The nodes selected for each step obtained for the baseline network;

[0039] S46, according to s i ,Π i , Calculate the equivalent total distance of the placement motion of the strategy network and the baseline network allocation results, and use the negative of the corresponding equivalent total distance as the reward value R(Π i |s i ) and the reward value of the baseline network

[0040] S47. Calculate policy gradient Update the parameters θ in the policy network using the Adam optimizer according to the policy gradient;

[0041] S48. Use the test data set to perform a significance test on the policy network and the baseline network. If the test result is less than the significance level parameter a, update the baseline network parameter θ B =θ;

[0042] S49, repeat S43 to S48 until the maximum number of iterations Z is reached, and the optimal strategy network parameters are obtained.

[0043] Preferably, S5 includes:

[0044] S51, using the optimal strategy network parameters, solve the circuit board instance, and obtain the patch head-mounting node number allocation result Π={π 1 ,π 2 ,...,π T};

[0045] S52, for π t =0 is a separation mark, which is divided into several pick-up and paste cycles to form the placement node allocation result PA of each pick-up and paste cycle;

[0046] S53, according to the serial numbers of the components picked up and pasted in each picking and pasting cycle, a dynamic programming method is used to determine the order PS of placing components in each picking and pasting cycle;

[0047] The placement path optimization results include PA and PS.

[0048] The beneficial effects of the present invention are as follows: the present invention trains and learns the optimization strategy based on deep reinforcement learning, effectively reduces the requirement that the solution strategy is heavily dependent on manual design, and effectively extracts features and learns decoding rules for mounting data through a well-designed decoder and encoder, so that the solution to the surface mounting optimization problem can be obtained quickly and stably, thereby significantly shortening the mounting path. The method provided by the present invention can autonomously learn the solution rules for surface mounting path optimization through deep reinforcement learning without relying on manually designed complex heuristic rules, thereby improving the production efficiency of the mounting machine. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a schematic diagram of the structure of a beam type placement machine;

[0050] Figure 2 This is a flow chart of the traditional placement machine picking process;

[0051] Figure 3 It is a flow chart of the encoder-decoder working process;

[0052] Figure 4 Diagram of the deep reinforcement learning training process. DETAILED DESCRIPTION

[0053] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0054] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0055] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but they are not intended to limit the present invention.

[0056] For a straight-row placement machine, the surface placement path optimization method based on deep reinforcement learning in this embodiment includes:

[0057] Step 1: Obtain placement machine parameters and circuit board production data;

[0058] The placement machine parameters include the total number of placement heads H, the index number of each placement head is h in increasing order along the X axis, h∈[1,…,H], and the placement head interval ΔH;

[0059] The circuit board production data includes: the number of cycles K of the pick-and-place cycle, we call the completion of a pick-and-place process a completion of a pick-and-place cycle, and the index of the pick-and-place cycle is recorded as k∈[1,…,K]; the placement component type group CPGs corresponding to all pick-and-place cycles, CPGs is a K×H array, the row represents the current placement cycle, the column represents the component type corresponding to each patch head, and the component group that needs to be placed in any pick-and-place cycle k is CPG=CPGs[k];

[0060] Candidate node set N = {n i , i=0,1,2,…,M}, the candidate node set includes mounting nodes (i=1,2,…,M) and virtual warehouse node depot (i=0), M represents the number of mounting points, and each node n i A tuple {n i =(x i ,y i ,z i ,c i ,d i ), i = 0, 1, 2, ..., M}, where x i and i The two-dimensional coordinate position of the corresponding node i, z i Indicates the mounting angle, c i Indicates the type of mounted components, d i Indicates the placement point demand. The demand for each unmounted placement point is 1, and the demand after placement is 0. When i = 0, z i =0, c i =0,d i =0.

[0061] Step 2: Taking the position information and mounting requirements of each node in the candidate node set as input, using an encoder based on a self-attention mechanism and a combination mask to extract high-dimensional data features, and outputting node embedding, in a preferred embodiment, specifically includes:

[0062] Step 21: Initialize the features of the candidate node set; the two-dimensional coordinate position and mounting point requirements of each node in the candidate node set are converted into the two-dimensional coordinate position and mounting point requirements of each node in the candidate node set through the learnable parameter W init ∈R 3×128 Linear mapping to high-dimensional vector space Represents each node of the high-dimensional vector space,

[0063] Step 22: Combine the CPG of the current cycle with the serial number of the completed placement points to generate an encoder mask Encoder_mask, which is used to filter the irrelevant placement points of the current cycle; the placement points irrelevant to the current cycle mainly include the placement points that have been allocated and the placement points corresponding to the component types that do not appear in the CPG;

[0064] Step 23: Initialize the encoder network layer index l=0, where L is the total number of encoder network layers;

[0065] Step 24: Determine whether l≤L, if so, execute step 25; otherwise, execute step 28;

[0066] Step 25: Execution MHA is a multi-head attention sublayer with H l Combined with Encoder_mask for input, the relevant information between the mounting points is obtained; BN is a batch normalization operation;

[0067] Step 26: Based on the obtained implement Where FF represents the feed-forward sublayer;

[0068] Step 27, update l=l+1, and return to step 24;

[0069] Step 28: Output node embedding H L .

[0070] Step 2 of this embodiment provides an encoder based on the self-attention mechanism and the combination mask, which effectively extracts and encodes data features for different component types and placement point allocation states in each picking and pasting cycle.

[0071] Step 3: Taking the node embedding of step 2 as input, a decoder based on a recurrent neural network and an attention mechanism is used to output the patch header-mounting node number allocation result; the encoder of step 2 and the decoder of step 3 together constitute a strategy network for patch header-mounting point number allocation; in a preferred embodiment, specifically including:

[0072] Step 31: Determine whether all mounting points have been allocated; if so, exit the decoder and output the patch header-mounting point sequence number allocation result Π; otherwise, execute step 32;

[0073] Step 32: Determine the currently selected node π t Is it 0 and t≠1? If so, enter the next picking cycle, execute steps 22 to 28 for encoding, and update the node embedding H L , r t Update to the number of placement nodes for the next pick-up and placement cycle, and then execute step 33; otherwise, directly execute step 33;

[0074] π t 0, indicating the start of a new picking cycle;

[0075] When t = 1, π 1 =0 means starting from the depot node; π t = 0, it means the current picking cycle has been completed, and you need to refresh r t The number of placement points for the next pick-and-place cycle;

[0076] Step 3. According to the currently selected node π t The number of components remaining in the current pick-and-place cycle r t , construct a new vector h c :As the input of the recurrent neural network;

[0077]

[0078] in, Represents node π t The embedding vector of , T represents the length of the allocation result sequence;

[0079] Step 34: h c After linear mapping, it is used as the input of the recurrent neural network, and the output of the recurrent neural network is q g =GRU(h c W G ), where W G ∈R 129×128 is a learnable linear mapping parameter, GRU represents a recurrent neural network;

[0080] Step 35: Use attention mechanism to obtain q gThe related information between the currently unassigned mounting point is denoted as q c ;

[0081] Step 36: q c As contextual embedding information in the decoding process, the probability distribution of candidate mounting points is calculated Where E = H L W K ,W K ,W K ∈R 128×128 , W K represents the learnable network weight parameter, C is a constant, tanh is the tangent function, The mask is used to mask the assigned mounting points during the decoding process, and softmax represents the normalized exponential function;

[0082] Step 37: According to Select the next node π for the placement head among the unassigned placement nodes t+1 Select the next node;

[0083] In training mode, probability sampling is used to embed H in nodes. L In select the next node for the patch head;

[0084] In inference mode, according to Use a greedy strategy to select the node with the highest probability as the next node;

[0085] Step 38: Update the decoding process mask according to the assigned mounting points Return to step 31;

[0086] Step 3 is used to optimize the placement node number allocation result Π for all patch heads; Step 3 uses the decoder combined with the attention mechanism to output the placement point location allocation result;

[0087] Step 4: Use a reinforcement learning training method to train the policy network. In a preferred embodiment, the method specifically includes:

[0088] Step 41. Input the maximum number of iterations Z, batch size X, significance level parameter a, policy network with trainable parameters θ, and policy network with trainable parameters θ. B The baseline network is the mirror network of the policy network.

[0089] Step 4.2: Initialize the parameters θ and θ in the policy network and the baseline network B ;

[0090] Step 43: Get instances from the training dataset

[0091] Step 4. Calculate the probability p based on the parameter θ in the policy network θ (Π i |s i ), and then randomly sample the probability distribution to generate the distribution result Π i ={π t ,t=1,...,T} i , where π t The nodes selected for each step;

[0092] Step 4 and 5: According to the parameters θ in the baseline network B Calculating Probability Then greedy sampling is performed through the probability distribution to generate the distribution result

[0093] Step 46: According to s i ,Π i , Calculate the reward value R(Π i |s i )and

[0094] Taking the reward value of the strategy network allocation result as an example, the reward value R(Π i |s i ) is defined as the equivalent total distance L(Π i |s i )

[0095] R(Π i |s i )=-L(Π i |s i )

[0096]

[0097] When π t =0, it means returning to the depot node; Represented as π t Node and π t+1 The equivalent distance of the node, where π t Corresponding to the hth patch head, π t+1 Corresponding to the fth patch head, ΔH is the distance between two adjacent patch heads in the X direction;

[0098] Reward value for baseline network allocation results Calculate similarly;

[0099] Step 47: Calculate the policy gradient Update the parameters θ in the policy network using the Adam optimizer according to the policy gradient;

[0100] Step 48: Use the test data set to perform a significance test on the policy network and the baseline network. If the test result is less than the significance level parameter a, update the baseline network parameter θ B =θ; the significance test of this embodiment adopts the T test method;

[0101] Step 49: repeat steps 43 to 48 until the maximum number of iterations is reached to obtain the optimal policy network parameters;

[0102] Step 5: Taking the result of the patch head-placing point number outputted by the strategy network in each pick-up cycle as input, the order of placement in each pick-up cycle is determined by the dynamic programming method to obtain the final placement path optimization result; in the preferred embodiment, it specifically includes:

[0103] Step 51: Use the optimal strategy network parameters obtained through training to solve the PCB instance, and generate the allocation result Π={π 1 ,π 2 ,...,π T};

[0104] Step 52: Convert π to π t =0 is a separation mark, which is divided into several pick-up and paste cycles to form the placement node allocation result PA of each pick-up and paste cycle;

[0105] Example: Π={0,2,3,5,7,9,1,0,4,6,11,8,10,12,0}, divided into two picking cycles: {2,3,5,7,9,1} and {4,6,11,8,10,12}

[0106] Step 53: According to the serial number of the components picked up in each pick-up cycle, the order PS of mounting components in each pick-up cycle is determined by dynamic programming method, and reference may be made to step 3 in a clustering-based LED placement machine pick-up path optimization method with publication number CN111465210A;

[0107] Step 54: Obtain the final placement path optimization results PA and PS;

[0108] This implementation method relies on deep reinforcement learning technology and can autonomously train parameterized surface mounting path optimization rules. This implementation method uses carefully designed decoders and encoders to effectively extract features from mounting data and learn decoding rules, and can quickly and stably obtain solutions to surface mounting optimization problems, thereby significantly shortening the mounting path and improving the production efficiency of the placement machine.

[0109] Although the present invention is described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. It should therefore be understood that many modifications may be made to the exemplary embodiments and that other arrangements may be devised without departing from the spirit and scope of the present invention as defined by the appended claims. It should be understood that the various dependent claims and features described herein may be combined in a manner different from that described in the original claims. It should also be understood that features described in conjunction with individual embodiments may be used in other described embodiments.

Claims

1. A surface mounting path optimization method based on deep reinforcement learning, characterized in that: include: S1, obtain the parameters of the placement machine and the circuit board production data, and build a candidate node set of placement nodes, including placement nodes and virtual warehouse nodes depot; S2, taking the position information and mounting requirements of each node in the candidate node set as input, using an encoder based on self-attention mechanism and combination mask to extract high-dimensional data features, and output node embedding; The mounting node requirement is whether to mount; S3, taking the node embedding of S2 as input, uses a decoder based on a recurrent neural network and an attention mechanism to output the patch header-mounting node number assignment result; S4, training a policy network, where the policy network includes the encoder and the decoder; S5. According to the placement head-placement node number allocation result of each pick-up and placement cycle output by the trained strategy network, the order of placing components in each pick-up and placement cycle is determined using a dynamic programming method to obtain the final placement path optimization result.

2. The surface mounting path optimization method based on deep reinforcement learning according to claim 1 is characterized in that: The S2 includes: S21, linearly map the two-dimensional coordinate position of each node in the candidate node set and the mounting node requirements to a high-dimensional vector space represents each node of the high-dimensional vector space, M represents the number of mounting nodes, and when i=0, the corresponding node is the virtual warehouse node depot; S22, combining the components to be mounted in the current pick-and-place cycle with the serial numbers of the completed mounting nodes, generating an encoder mask Encoder_mask; Encoder_mask is used to filter irrelevant mounting nodes in the current pick-and-place cycle; S23, initializing the network layer index l of the encoder = 0, where L is the total number of network layers of the encoder; S24, determine whether l≤L, if so, execute S25; otherwise, execute S28; S25. Execution MHA is the multi-head attention sublayer, and BN is the batch normalization operation; S26, according to the obtained implement Where FF represents the feed-forward sublayer; S27, update l=l+1, and return to S24; S28, output node embedding H L .

3. The surface mounting path optimization method based on deep reinforcement learning according to claim 1, characterized in that S3 include: S31, determine whether all mounting nodes have been assigned; if so, exit the decoder and output the patch header-mounting node sequence number assignment result Π; otherwise, execute S32; S32, determine the node embedding H L The currently selected node π t Is it 0 and t≠1? If so, enter the next picking cycle, execute S2 to re-encode, and update the node embedding H L , r t Update to the number of placement nodes for the next pick-up and placement cycle, and then execute S33; otherwise, directly execute S33; S33, according to the currently selected node π t The number of components remaining in the current pick-and-place cycle r t , construct a new vector h c : in, Represents node π t The embedding vector of , T represents the length of the allocation result sequence; S34, vector h c After linear mapping, it is used as the input of the recurrent neural network, and the output of the recurrent neural network is q g =GRU(h c W G ), where W G is the learnable network weight parameter, GRU represents a recurrent neural network; S35, use attention mechanism to obtain q g Related information about the currently unassigned placement nodes c ; S36, q c As the context embedding information in the decoding process, the probability distribution of candidate placement nodes is calculated Where E=H L W K ,W K represents the learnable network weight parameter, C is a constant, tanh is the tangent function, For the decoding process, the mask is used to shield the already assigned mounting nodes, and softmax represents the normalized exponential function; S37, according to Select the next node π for the placement head among the unassigned placement nodes t+1 ; S38, update the decoding process mask according to the assigned mounting node Return to S31.

4. The surface mounting path optimization method based on deep reinforcement learning according to claim 3 is characterized in that: In S37, in the training mode, according to The probability sampling method is used to embed H in the node L Select the next node for the patch header.

5. The surface mounting path optimization method based on deep reinforcement learning according to claim 3 is characterized in that: In S37, in the inference mode, according to Use a greedy strategy to select the node with the highest probability as the next node.

6. The surface mounting path optimization method based on deep reinforcement learning according to claim 3, characterized in that S4 include: S41, input maximum number of iterations Z, batch size X, significance level parameter a, policy network with trainable parameters θ, B The baseline network is the mirror network of the policy network. S42. Initialize the parameters θ and θ in the policy network and the baseline network B ; S43. Obtaining instances from the training dataset S44. Calculate the probability p based on the parameter θ in the policy network θ (Π i |s i ), and then by probability p θ (Π i |s i ) is used to generate the distribution result Π by probability sampling i ={π t ,t=1,...,T} i ; S45. According to the parameters θ in the baseline network B Calculating Probability Then pass Greedy sampling of the distribution generates the distribution result The nodes selected for each step obtained for the baseline network; S46, according to s i , Π i , Calculate the equivalent total distance of the placement motion of the strategy network and the baseline network allocation results, and use the negative of the corresponding equivalent total distance as the reward value R(Π i |s i ) and the reward value of the baseline network S47. Calculate policy gradient Update the parameters θ in the policy network using the Adam optimizer according to the policy gradient; S48. Use the test data set to perform a significance test on the policy network and the baseline network. If the test result is less than the significance level parameter a, update the baseline network parameter θ B =θ; S49, repeat S43 to S48 until the maximum number of iterations Z is reached, and the optimal strategy network parameters are obtained.

7. The surface mounting path optimization method based on deep reinforcement learning according to claim 3, characterized in that S5 include: S51, using the optimal strategy network parameters, solve the circuit board instance and obtain the patch header-mounting node number allocation result Π={π1,π2,...,π T }; S52, for π t =0 is a separation mark, which is divided into several pick-up and paste cycles to form the placement node allocation result PA of each pick-up and paste cycle; S53, according to the serial numbers of the components picked up and pasted in each picking and pasting cycle, a dynamic programming method is used to determine the order PS of placing components in each picking and pasting cycle; The placement path optimization results include PA and PS.

8. A computer-readable storage device storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the surface mounting path optimization method based on deep reinforcement learning are implemented as described in any one of claims 1 to 7.

9. A surface mount path optimization device based on deep reinforcement learning, comprising a storage device, a processor, and a computer program stored in the storage device and executable on the processor, characterized in that: The processor executes the computer program to implement the steps of the surface mounting path optimization method based on deep reinforcement learning as described in any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the surface mounting path optimization method based on deep reinforcement learning as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Clustering-based LED chip mounter pick-and-mount path optimization method

    CN111465210A

  • A Cluster-Based Optimization Method for Pick-up and Placement Path of LED Chip and Mount Machine

    CN111465210B

  • Straight-line chip mounter surface mounting process optimization method based on layered heuristic

    CN116056442A

Cited By

  • Electronic component classification prediction method and device

    CN120449780A