Dispensing path planning method, network training method, dispensing equipment and medium
By optimizing the dispensing path planning using deep Q-networks, and taking into account the acceleration, deceleration, and uniform speed movement time of the dispensing equipment as well as obstacle avoidance energy consumption, the problem of insufficient dispensing time optimization in existing technologies is solved, and more efficient dispensing path planning is achieved.
Patent Information
- Application Number
- CN202510972083.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies only consider path length in dispensing path planning, failing to effectively optimize dispensing time. This is especially problematic when there are frequent short-distance movements, leading to frequent start-stop of the dispensing head and impacting efficiency.
A deep Q-network is used for path planning. The dispensing motion time is optimized through a reward function, taking into account the acceleration, deceleration and uniform speed motion of the dispensing equipment, and combining obstacle avoidance and energy consumption factors to generate the optimal path.
It improves dispensing efficiency, avoids dispensing losses caused by frequent short-distance movements, generates adaptive optimal paths, and enhances the production efficiency of dispensing equipment.
Smart Images

Figure CN120993823A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of automation control technology, and in particular to a dispensing path planning method, a network training method, a dispensing device and a medium. BACKGROUND
[0002] In the field of automation control technology such as electronic manufacturing and semiconductor packaging, dispensing machines are usually used to accurately apply glue, solder paste or other fluids to multiple preset points on a workpiece. The movement path of the dispensing head during dispensing directly affects the duration and efficiency of dispensing.
[0003] Prior art usually uses simple scanning, nearest neighbor method or heuristic algorithm for dispensing path planning. However, the traditional method of prior art usually only considers the straight-line distance between points to minimize path length as the optimization objective for path planning.
[0004] The inventors have found that dispensing efficiency in a dispensing task depends not only on path length but also on the speed curve of each movement segment; simply minimizing distance does not necessarily guarantee minimizing dispensing time, especially when the dispensing head moves frequently at short distances, the frequent start and stop of the dispensing head may actually lead to an increase in dispensing time. Therefore, the inventors provide a dispensing path planning method that takes into account the movement process of the dispensing device based on this finding. SUMMARY
[0005] The present application provides a dispensing path planning method, a network training method, a dispensing device and a medium to improve dispensing efficiency.
[0006] According to an aspect of the present application, a dispensing path planning method is provided, the method comprising:
[0007] obtaining the current position of the dispensing head and the access state of each dispensing point in a dispensing task;
[0008] inputting the current position and the access state of each dispensing point into the main network of a pre-trained deep Q network, and determining the next dispensing point according to the main network Q value output by the main network;
[0009] wherein, when training the deep Q network, the reward function in the deep Q network is determined in the following manner to make the deep Q network take the reward function as the optimization objective: when the dispensing head moves from the current position to the next position corresponding to the next dispensing point, the dispensing movement time from the current position to the next position is determined according to the acceleration, deceleration and uniform motion process of the dispensing device; and the reward function in the deep Q network is determined according to the dispensing movement time;
[0010] According to the next point position updating the current position of the dispensing head and the access state of each to-be-dispensed point position, a next point position is determined by a main network Q value output by a main network of the deep Q network until all the point positions are accessed, and a dispensing path planning result is obtained.
[0011] According to an aspect of the present application, a dispensing path planning method is provided, which comprises:
[0012] obtaining the current position of the dispensing head and the access state of each to-be-dispensed point position in a dispensing task;
[0013] constructing a main network of the deep Q network, and taking the current position and the access state of each to-be-dispensed point position as input data of the main network;
[0014] in the iterative decision of the main network, a next point position is determined according to a main network Q value output by the main network;
[0015] when the dispensing head moves from the current position to a next position corresponding to the next point position, dispensing motion time from the current position to the next position is determined according to the acceleration, deceleration and uniform motion process of the dispensing equipment, and a reward function in the deep Q network is determined according to the dispensing motion time;
[0016] generating an experience pool according to the current position, the next point position, the next position, the reward function and the access state of each to-be-dispensed point position;
[0017] constructing a target network of the deep Q network, and calculating a target Q value according to the reward function, a discount factor and a target network output Q value of a sample in the experience pool;
[0018] wherein, parameters in the main network are copied every preset step length to determine parameters in the target network;
[0019] iterative training is performed by determining a loss function according to the main network Q value and the target Q value, and a trained deep Q network is obtained to plan a dispensing path according to the trained deep Q network.
[0020] According to another aspect of the present application, a dispensing path planning device is provided, which comprises:
[0021] a dispensing task data acquisition module for obtaining the current position of the dispensing head and the access state of each to-be-dispensed point position in a dispensing task;
[0022] a next point position determination module for inputting the current position and the access state of each to-be-dispensed point position into a main network of a pre-trained deep Q network, and determining a next point position according to a main network Q value output by the main network;
[0023] The reward function in the deep Q network is determined by the following manner when the deep Q network is trained, so that the deep Q network takes the reward function as an optimization target: when the dispensing head moves from the current position to a next position corresponding to a next dispensing point, the dispensing motion time from the current position to the next position is determined according to the acceleration, deceleration and constant speed motion process of the dispensing equipment; and the reward function in the deep Q network is determined according to the dispensing motion time.
[0024] The dispensing path planning result determination module is configured to update the current position of the dispensing head and the access state of each dispensing point to be dispensed according to the next dispensing point, and return a step of determining the next dispensing point by the main network Q value output by the main network of the deep Q network, until all dispensing points are accessed, to obtain the dispensing path planning result.
[0025] According to another aspect of the present application, a deep Q network training device for dispensing path planning is provided, which comprises:
[0026] The dispensing task data acquisition module is configured to acquire the current position of the dispensing head and the access state of each dispensing point to be dispensed in the dispensing task;
[0027] The main network construction module is configured to construct the main network of the deep Q network, and take the current position and the access state of each dispensing point to be dispensed as input data of the main network;
[0028] The next dispensing point determination module is configured to determine the next dispensing point according to the main network Q value output by the main network in the iterative decision of the main network;
[0029] The reward function determination module is configured to determine the dispensing motion time from the current position to the next position according to the acceleration, deceleration and constant speed motion process of the dispensing equipment when the dispensing head moves from the current position to the next position corresponding to the next dispensing point, and determine the reward function in the deep Q network according to the dispensing motion time;
[0030] The experience pool generation module is configured to generate an experience pool according to the current position, the next dispensing point, the next position, the reward function and the access state of each dispensing point to be dispensed;
[0031] The target Q value calculation module is configured to construct a target network of the deep Q network, and calculate a target Q value according to the reward function, the discount factor and the Q value output by the target network of the sample in the experience pool;
[0032] The target network is configured to copy the parameters from the main network every preset step length to determine the parameters in the target network;
[0033] The deep Q network training module is configured to obtain the trained deep Q network by iterative training according to a loss function determined according to the main network Q value and the target Q value, so as to plan the dispensing path according to the trained deep Q network.
[0034] According to another aspect of the present application, a dispensing device is provided, comprising a motion platform, a dispensing head, a position sensor, a motion controller, at least one processor, and a memory in communication with the at least one processor; wherein:
[0035] The motion platform is configured to place a workpiece to be dispensed; the dispensing device performs dispensing operation on the workpiece to be dispensed while the motion platform moves;
[0036] The dispensing head is configured to perform dispensing operation on the workpiece to be dispensed on the motion platform at a dispensing point;
[0037] The position sensor is configured to obtain the position of the dispensing head;
[0038] The motion controller is configured to control the movement of the dispensing head;
[0039] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the dispensing path planning method according to any one of the embodiments of the present application, or the deep Q network training method for dispensing path planning.
[0040] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to implement the dispensing path planning method according to any one of the embodiments of the present application, or the deep Q network training method for dispensing path planning.
[0041] According to another aspect of the present application, a computer program product is provided, comprising a computer program for enabling a processor to implement the dispensing path planning method according to any one of the embodiments of the present application, or the deep Q network training method for dispensing path planning.
[0042] The technical scheme of the embodiment of the application comprises the following steps: obtaining a current position of a dispensing head in a dispensing task and access states of each to-be-dispensed point; inputting the current position and the access states of each to-be-dispensed point into a main network of a pre-trained deep Q network, and determining a next dispensing point according to a main network Q value output by the main network; wherein, when training the deep Q network, a reward function in the deep Q network is determined in the following manner, so that the deep Q network takes the reward function as an optimization target: when the dispensing head moves from the current position to a next position corresponding to the next dispensing point, the dispensing motion time from the current position to the next position is determined according to the acceleration, deceleration and uniform motion process of the dispensing equipment; and the reward function in the deep Q network is determined according to the dispensing motion time; the current position of the dispensing head and the access states of each to-be-dispensed point are updated according to the next dispensing point, and the step of determining the next dispensing point by returning the main network Q value output by the main network of the deep Q network is repeated until all dispensing points are accessed, so that a dispensing path planning result is obtained, the path planning problem during dispensing of a workpiece is solved, the dispensing motion time is determined by considering the speed change of each motion section of the dispensing equipment during path planning, and then the reward function is generated, the reward function can adaptively perform optimal path planning, and the dispensing efficiency is improved, and in particular, dispensing loss caused by frequent movement of a short distance can be avoided.
[0043] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the application, nor is it used to limit the scope of the application. Other features of the application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0045] Figure 1a is a flow chart of a dispensing path planning method according to an embodiment of the application;
[0046] Figure 1b shows an application flow chart of a dispensing path planning method;
[0047] Figure 2a is a flow chart of a deep Q network training method for dispensing path planning according to an embodiment of the application;
[0048] Figure 2b shows a dispensing path planning application structure schematic diagram;
[0049] Figure 3is a structural schematic diagram of a point dispensing path planning device according to an embodiment three of the present application;
[0050] Figure 4 is a structural schematic diagram of a deep Q network training device for point dispensing path planning according to an embodiment four of the present application;
[0051] Figure 5 A structural schematic diagram of a point dispensing device 10 that can be used to implement embodiments of the present application is shown. DETAILED DESCRIPTION
[0052] In order to make the person skilled in the art better understand the present application scheme, the technical scheme in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by the person skilled in the art without creative labor should belong to the scope of protection of the present application.
[0053] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0054] Embodiment one
[0055] Figure 1a is a flowchart of a point dispensing path planning method according to an embodiment one of the present application. The present embodiment can be applicable to path planning for minimizing dispensing time in a dispensing task. The method can be executed by a point dispensing path planning device, which can be realized in the form of hardware and / or software, and can be configured in a point dispensing device. As shown in the figure, the method comprises: Figure 1a
[0056] Step 110, obtaining the current position of the dispensing head and the access state of each dispensing point in the dispensing task.
[0057] In the present application, in order to facilitate data processing, the dispensing task can be described as a Markov decision process (S, A, P, R, γ). Wherein, S is a state space, which includes the current position of the dispensing head and the access state of each dispensing point to be dispensed; s ∈ S, which can be expressed as s = (pos_current, V). pos_current is the current position of the dispensing head, and V is the access state of each dispensing point to be dispensed, which can be a vector of length N or a bit mask. For example, 0 in V indicates that the corresponding dispensing point to be dispensed is in an unvisited state, and 1 indicates that the corresponding dispensing point to be dispensed is in a visited state. A is an action space, under state s, action a ∈ A(s) indicates that the a-th dispensing point to be visited is selected as the next dispensing point to be visited. P is a state transition, after action a is performed, state s is transferred to s', pos_current is updated to the coordinates of the next dispensing point Pa obtained by performing the dispensing action, and the state corresponding to Pa in V is updated to visited. R is a reward function, which is a function of the cost of performing action a, and is expressed as R(s, a). In the present application, the reward function can be determined according to the time cost when the action a is performed, and the path planning goal is to minimize the dispensing time. γ is a discount factor, which indicates the importance of the time cost of future actions in the reward function. For example, when γ is 1, it indicates that the time cost of all future actions is equally important.
[0058] By describing the dispensing task as a Markov decision process (S, A, P, R, γ), the optimal dispensing path planning can be realized by learning the optimal action value function, i.e. the reward function R(s, a). In the present application, the reward function R(s, a) can be approximated by a deep Q-network (DQN) in reinforcement learning to guide the dispensing equipment to select a moving path that minimizes the dispensing time.
[0059] Step 120, inputting the current position and the access state of each dispensing point to be dispensed into the main network of the pre-trained deep Q-network, and determining the next dispensing point according to the main network Q value output by the main network.
[0060] Wherein, the current position of the dispensing head and the access state of each dispensing point to be dispensed in the dispensing task s = (pos_current, V) can be used as input data during training of the deep Q-network, and the output is the Q value Q(s, a; θ) of each action. θ is the main network parameter. By pre-training the deep Q-network, the current position of the dispensing head and the access state of each dispensing point to be dispensed in the current dispensing task are input into the main network of the pre-trained deep Q-network, and the main network Q value of each action output by the main network can be obtained. For example, the unvisited dispensing point with the maximum main network Q value can be selected as the next dispensing point.
[0061] In the process of planning the dispensing path by using the deep Q network, in order to make the deep Q network take minimizing the dispensing time as the goal, optionally, in the process of training the deep Q network, the reward function in the deep Q network is determined by the following method so as to make the deep Q network take the reward function as the optimization goal: in the process of moving the dispensing head from the current position to the next position corresponding to the next dispensing point, according to the acceleration, deceleration and uniform motion process of the dispensing equipment, the dispensing motion time from the current position to the next position is determined; and the reward function in the deep Q network is determined according to the dispensing motion time.
[0062] The inventor considers that in the process of moving the dispensing head from the current position to the next position, the dispensing head will experience the acceleration, deceleration and uniform motion process, and the acceleration and deceleration will cause the frequent start and stop of the dispensing equipment, thereby increasing the dispensing time. That is to say, minimizing the dispensing movement distance does not necessarily minimize the dispensing time. Especially in the process of frequent short distance movement, the acceleration and deceleration time may dominate the dispensing time. Therefore, the inventor determines the dispensing motion time from the current position to the next position by considering the acceleration, deceleration and uniform motion process in the process of moving the dispensing head from the current position to the next position, and then determines the reward function in the deep Q network according to the dispensing motion time.
[0063] For example, it is assumed that the dispensing equipment accelerates and decelerates at the maximum acceleration a max , and uniformly moves at the maximum speed v max . The total distance of acceleration and deceleration when the dispensing head accelerates from the initial speed 0 to the maximum speed and then decelerates from the maximum speed to 0 is The movement path from the current position to the next position is L = ||p a -pos current ||2. When L ≥ d ad , the trapezoidal speed curve is used to determine the dispensing motion time from the current position to the next position. When L < d ad , the triangular speed curve is used to determine the dispensing motion time from the current position to the next position. Optionally, the reward function is the negative value of the dispensing motion time, that is, R(s, a) = -T move (pos current , p a ; v max , a max). To consider the special requirements in the dispensing scenario and improve the process performance of dispensing, one or more indicators of dispensing obstacle avoidance, dispensing device energy consumption, and motion smoothness can be considered in the reward function. Optionally, when training the deep Q network, the reward function in the deep Q network is determined in such a way that the deep Q network takes the reward function as an optimization objective: when the dispensing head moves from the current position to the next position corresponding to the next dispensing point, according to the acceleration, deceleration, and constant speed motion process of the dispensing device, the dispensing motion time from the current position to the next position is determined; according to the proximity between the movement path of the dispensing head from the current position to the next position and the preset obstacle region in the dispensing task, the dispensing obstacle avoidance penalty parameter from the current position to the next position is determined; according to the number of accelerations, decelerations, acceleration times, and deceleration times in the process of moving the dispensing head from the current position to the next position, the dispensing energy consumption parameter from the current position to the next position is determined; and according to the dispensing obstacle avoidance penalty parameter and / or the dispensing energy consumption parameter, and the dispensing motion time, the reward function in the deep Q network is determined.
[0064] In the dispensing scenario, there can be regions on the workpiece to be dispensed that are prohibited from being crossed or approached, such as installed high-value components or special marking areas, etc. In order to avoid these regions when planning the dispensing path, the dispensing obstacle avoidance penalty parameter can be set according to the proximity between the movement path and the preset obstacle region. For example, the proximity between the movement path and the preset obstacle region is inversely related to the dispensing obstacle avoidance penalty parameter, that is, the closer the movement path is to the preset obstacle region, the greater the value of the dispensing obstacle avoidance penalty parameter. For example, the dispensing obstacle avoidance penalty parameter can be determined according to the reciprocal of the minimum value of the boundary between the movement path and the preset obstacle region.
[0065] Optionally, the dispensing obstacle avoidance penalty parameter from the current position to the next position is determined according to the proximity between the movement path of the dispensing head from the current position to the next position and the preset obstacle region in the dispensing task, including: calculating the shortest distance from the movement path to the preset obstacle region in the dispensing task according to the movement path of the dispensing head from the current position to the next position; and determining the dispensing obstacle avoidance penalty parameter from the current position to the next position according to the reciprocal of the shortest distance and the reciprocal of the preset dispensing safety margin.
[0066] For example, M circular obstacle regions {O1, O2,..., OM} are predefined on the workpiece to be dispensed, and the center of each obstacle region O M is c k , and the radius is rad k . The line segment L k connecting the current position p i and the next position p j is L ij . The shortest distance from the line segment L ij to each obstacle region O kthe shortest distance d(L ij , c k ) of the mobile path to the boundary of the preset obstacle region k . When the shortest distance d(L safe , c consump ) is less than the safety margin m The point dispensing obstacle avoidance penalty parameter represents that only when the distance of the mobile path to the boundary of the preset obstacle region is less than the safety margin, a non-zero and sharply increasing penalty with the decrease of the distance will be generated.
[0067] By generating the point dispensing obstacle avoidance penalty parameter according to the distance of the mobile path to the preset obstacle region in the point dispensing scene, and then generating the reward function according to the point dispensing obstacle avoidance penalty parameter to plan the point dispensing path, the point dispensing device can learn to detour instead of only taking the shortest straight line, so that the point dispensing path planning is adapted to the actual obstacle avoidance situation in the point dispensing scene.
[0068] In the actual point dispensing scene, the energy consumption of the point dispensing device is also a factor that needs to be considered in the point dispensing path planning. The inventors found in actual research that the energy consumption of the point dispensing device is closely related to the acceleration and deceleration process of the motor. Therefore, according to the number of accelerations, the number of decelerations, the acceleration time and the deceleration time in the process of moving the point dispensing head from the current position to the next position, the point dispensing energy consumption parameter from the current position to the next position is determined to avoid the increase of the energy consumption of the point dispensing device caused by frequent start and stop.
[0069] For example, the energy consumption of the point dispensing device is proportional to the integral of the square of the acceleration (representing power) with respect to time in the acceleration and deceleration process. Considering that the movement from the current position to the next position must include an acceleration and a deceleration, the point dispensing energy consumption parameter can be determined according to the acceleration time and the deceleration time. Exemplarily, the point dispensing energy consumption parameter is represented as E consump (i, j) = C motor · (T accel_phase + T decel_phase ). Wherein, T accel_phase is the acceleration time, T decel_phase is the deceleration time, and C motor is a constant related to the characteristics of the motor, which is related to the motor of the point dispensing device.
[0070] By determining the reward function according to the point dispensing energy consumption parameter to plan the point dispensing path, the point dispensing device can be guided to select a smoother and more uniform motion path, i.e. when facing multiple point positions with similar distances, the point dispensing device tends to select a path combination that can form a longer uniform motion, thereby avoiding frequent start and stop.
[0071] In practical applications, the obstacle avoidance and energy consumption problems can be considered according to the actual dispensing scene, so as to determine the reward function in the deep Q network according to the dispensing obstacle avoidance penalty parameter and / or the dispensing energy consumption parameter, and the dispensing motion time. For example, the reward function can be expressed as r t =-(w t ·T move (i,j)+w o ·P obstacle (i,j)+w e ·E consump (i,j))· t wherein w o , w e and w t are weight coefficients, which can be adjusted according to the actual dispensing application scene. For example, the time priority, safety priority or energy saving priority of the dispensing path can be considered in the dispensing scene through weight coefficient adjustment. By considering the minimization of dispensing motion time, avoiding the path close to the obstacle area, and minimizing the energy consumption of the dispensing device in the dispensing path planning, the dispensing performance can be improved.
[0072] Step 130, updating the current position of the dispensing head and the access state of each to-be-dispensed point according to the next dispensing point, returning to the step of determining the next dispensing point through the main network Q value output by the main network of the deep Q network, until all dispensing points are accessed, and obtaining the dispensing path planning result.
[0073] Figure 1b An application flowchart of a dispensing path planning method is shown. As shown in Figure 1b , for a new dispensing task, the pre-trained main network Q(s, a; θ) of the deep Q network can be loaded. The exploration rate ∈=0 is set, that is, the pure exploitation strategy is used. The current position of the dispensing head and the access state of each to-be-dispensed point are obtained. The dispensing path sequence is initialized as empty. The following steps are repeated until all dispensing points are accessed: input the current state into the main network of the deep Q network to obtain the Q value corresponding to each unvisited dispensing point; select the unvisited dispensing point corresponding to the maximum Q value as the next dispensing point a t =argmax a∈unvisited Q(s t , a; θ); add the selected next dispensing point a to the path sequence; update the current state. After the repetition is completed, the generated dispensing path planning result is output, and the dispensing device is controlled to move and dispense in sequence according to the dispensing path planning result.
[0074] Since the deep Q network determines the reward function in the deep Q network by the point adhesion obstacle avoidance penalty parameter and / or the point adhesion energy consumption parameter, and the point adhesion movement time during training, the point adhesion movement time, the obstacle avoidance requirement on the specific workpiece, and the point adhesion equipment energy consumption can be comprehensively considered, and the quality of the point adhesion path is improved.
[0075] The technical scheme of the embodiment obtains the current position of the point adhesion head and the access state of each to-be-point-adhesion point in a point adhesion task, inputs the current position and the access state of each to-be-point-adhesion point into a main network of a pre-trained deep Q network, and determines a next point adhesion point according to a main network Q value output by the main network. During training of the deep Q network, the reward function in the deep Q network is determined in the following manner, so that the deep Q network takes the reward function as an optimization target: when the point adhesion head moves from the current position to a next position corresponding to the next point adhesion point, the point adhesion movement time from the current position to the next position is determined according to the acceleration, deceleration, and uniform speed movement process of the point adhesion equipment; and the reward function in the deep Q network is determined according to the point adhesion movement time; the current position of the point adhesion head and the access state of each to-be-point-adhesion point are updated according to the next point adhesion point, and the step of determining the next point adhesion point by outputting the main network Q value by the main network of the deep Q network is returned until all point adhesion points are accessed, and a point adhesion path planning result is obtained, thereby solving the path planning problem during point adhesion of a workpiece. By considering the speed change of each movement section of the point adhesion equipment to determine the point adhesion movement time during path planning, and then generating the reward function, optimal path planning can be adaptively performed, the point adhesion efficiency is improved, and point adhesion loss caused by frequent movement of a short distance can be avoided.
[0076] Embodiment Two
[0077] Figure 2a A flowchart of a deep Q network training method for point adhesion path planning according to Embodiment Two of the present application is shown. The embodiment can be applicable to path planning for minimizing point adhesion time in a point adhesion task through network training. The method can be performed by a deep Q network training device for point adhesion path planning. The deep Q network training device for point adhesion path planning can be realized in the form of hardware and / or software, and can be configured in a point adhesion equipment. The embodiment is a further addition and refinement of the above technical scheme. The technical scheme in the embodiment can be combined with each optional scheme in one or more of the above embodiments. As shown in the figure, the method comprises: Figure 2a
[0078] Step 210: obtaining the current position of the point adhesion head and the access state of each to-be-point-adhesion point in a point adhesion task.
[0079] As described above, the dispensing task can be described as a Markov decision process (S, A, P, R, γ), which will not be repeated here.
[0080] Step 220, constructing a main network of the deep Q network, and taking the current position and the access state of each dispensing point as input data of the main network.
[0081] The main network can include an input layer, a hidden layer and an output layer. The current position of the dispensing head and the access state of each dispensing point in the dispensing task can be taken as input data of the input layer of the main network. The hidden layer can be one or more fully connected layers, and a nonlinear activation function such as ReLU can be used for data processing in the fully connected layer. The output layer can be N neurons, which can use linear activation, and each neuron outputs a Q value Q(s, a; θ) of an action. When constructing the main network, the main network parameters θ can be initialized.
[0082] Step 230, in the iterative decision of the main network, determining the next dispensing point according to the main network Q value output by the main network.
[0083] In the deep Q network, an exploration rate ∈ can be specified. In the iterative training of the main network, the exploration rate can be initialized to a value greater than a preset parameter and decayed with training. In the training process, the main network can randomly select an unvisited action a t , i.e., exploration, with a probability of ∈, and select an unvisited action a t that maximizes the output of the main network Q(s t , a; θ) with a probability of 1-∈, i.e., using a∈unvisited t .
[0084] When determining the next dispensing point, action masking can be performed, i.e., only selecting the next dispensing point from unvisited dispensing points.
[0085] Step 240, when the dispensing head moves from the current position to the next position corresponding to the next dispensing point, determining the dispensing motion time from the current position to the next position according to the acceleration, deceleration and uniform motion process of the dispensing equipment; and determining the reward function in the deep Q network according to the dispensing motion time.
[0086] The determination of the dispensing motion time is the same as the foregoing, which will not be described herein. Optionally, when training the deep Q network, the reward function in the deep Q network is determined in the following manner so as to make the deep Q network take the reward function as an optimization target: when the dispensing head moves from the current position to a next position corresponding to a next dispensing point, the dispensing motion time from the current position to the next position is determined according to the acceleration, deceleration and uniform motion process of the dispensing equipment; the dispensing obstacle avoidance penalty parameter from the current position to the next position is determined according to the proximity between the movement path of the dispensing head from the current position to the next position and the preset obstacle region in the dispensing task; the dispensing energy consumption parameter from the current position to the next position is determined according to the number of acceleration times, the number of deceleration times, the acceleration time and the deceleration time in the process of moving the dispensing head from the current position to the next position; and the reward function in the deep Q network is determined according to the dispensing obstacle avoidance penalty parameter and / or the dispensing energy consumption parameter and the dispensing motion time.
[0087] Optionally, the dispensing obstacle avoidance penalty parameter from the current position to the next position is determined according to the proximity between the movement path of the dispensing head from the current position to the next position and the preset obstacle region in the dispensing task, including: the shortest distance from the movement path to the preset obstacle region in the dispensing task is calculated according to the movement path of the dispensing head from the current position to the next position; and the dispensing obstacle avoidance penalty parameter from the current position to the next position is determined according to the reciprocal of the shortest distance and the reciprocal of the preset dispensing safety margin.
[0088] Step 250, generating an experience pool according to the current position, the next dispensing point, the next position, the reward function and the access state of each dispensing point.
[0089] wherein the experience pool can be expressed as (s t , a t , r t , s t+1 , done t ), wherein s t is the current position, a t is the action of executing the next dispensing point, r t is the reward function when moving from the current position to the next position, s t+1 is the next position, and done t marks whether s t+1 is a terminal state.
[0090] Step 260, constructing a target network of the deep Q network, and calculating a target Q value according to the reward function of the sample in the experience pool, the discount factor and the output Q value of the target network.
[0091] The target network Q target (s, a; θ -The structure is the same as the main network, but the parameters are different. The parameters θ of the target network are... - The parameters θ can be periodically copied from the main network for calculating the target Q-value during stable training. During initialization, θ... - =θ, and initialize the experience pool, learning rate α, discount factor γ, exploration rate ∈, training batch size B of the target network, and parameter update frequency C of the target network.
[0092] During training, the target network can replay experiences from the experience pool. When there are enough samples in the experience pool, a training batch of size B is randomly sampled, with a size {(s)}. j ,a j ,r j ,s j+1 ,done j For each sample, the target Q-value is calculated based on the reward function, discount factor, and target network output Q-value. Where, r j For the reward function, This is the maximum Q value output by the target network for unvisited points when the next glue point is not in a stopped state.
[0093] When determining the target Q value, a static discount factor γ can be used, such as setting it to 1 to indicate that the reward function of all future actions is given equal importance. However, in order to make the dispensing path planning have different decision emphases at different stages, the discount factor can be adjusted according to the current state characteristics of dispensing. Through the dynamic discount factor, the dispensing path planning decision can be dynamically adjusted.
[0094] The inventors discovered in their research that decisions made in densely populated areas have a significant impact on the overall path, requiring greater attention to immediate rewards and thus allowing for a reduction in the discount factor. Conversely, in sparsely populated areas, where the dispensing head needs to make long-distance transfers, future path selection becomes more crucial, necessitating a higher discount factor. In other words, the discount factor can be inversely correlated with the density of unvisited points; higher density results in a smaller discount factor.
[0095] Optionally, when training a deep Q-network, the discount factor is determined as follows: the dot density of the area where the dispensing head is located is determined based on the current position of the dispensing head and the unvisited dispensing dots within the preset area of the current position; the discount factor is then determined based on the dot density.
[0096] For example, define a current position p i Centered on a radius of R local The area where the dispensing head is located. Count the number N of unvisited points within the area where the dispensing head is located. local =|{k|mask t [k] = 0, and, dist(pi ,p k )≤R local}|. Among them, dist(p i p k ) represents the distance between the current location and the unvisited point k. mask t [k] = 0 indicates an unvisited point. The point density in the area where the dispensing head is located is... The dynamic discount factor can decrease as the point density increases. For example, the dynamic discount factor can be expressed as γ. t =γ max -(γ max -γ min )·tanh(α·ρ). Where, γ max γ min γ represents the upper and lower bounds of the discount factor. For example, the upper bound of the discount factor can be a value in the range of 0.99 to 1; the lower bound of the discount factor can be a value in the range of 0.9 to 0.91. α is the coefficient for adjusting sensitivity. tanh(·) is the hyperbolic tangent function. At points with dense concentrations, i.e., the larger ρ is, the greater γ becomes. t Tend to γ min In areas where the points are sparse, i.e., the smaller ρ is, the greater γ is. t Tend to γ max By using dynamic discount factors, the selection of dispensing points can be dynamically adjusted in the dispensing path planning process, making the determination of the points more adaptable to the specific dispensing scenario.
[0097] Step 270: Iteratively train the deep Q-network by determining the loss function based on the Q-value of the main network and the target Q-value, and then perform dispensing path planning based on the trained deep Q-network.
[0098] For example, mean squared error, mean absolute error, or Hubble loss can be used. The loss function can be expressed as follows: Gradient descent can be used to update the parameters of the main network. α is the learning rate. Parameters are copied from the main network at preset step intervals to determine the parameters θ in the target network. - ←θ.
[0099] To refine the process of determining the dispensing path, in the early stages of training a deep Q-network, when the loss decreases rapidly and the total reward increases significantly, the preset step size C can be appropriately increased to allow the main network to explore more thoroughly. When training enters the later stages and the loss and reward tend to stabilize, the preset step size C can be decreased to allow the target network to synchronize more frequently for fine-tuning, helping the model converge to the optimal policy faster.
[0100] Optionally, in training the deep Q network, the preset step length in copying the main network parameters is determined according to the loss function change value and / or the reward function change value in the deep Q network training process.
[0101] In the above method, the preset step length can be increased when the loss function change value is negative and / or the reward function change value is positive, and the preset step length can be decreased when the loss function change value is positive and / or the reward function change value is negative.
[0102] For example, a counter update_counter = 0 and a preset step length C (for example, C = 100) are initialized. After each update of the main network parameters, the update_counter is increased by one. When the update_counter reaches C, the main network parameters are copied to the target network: θ - ← θ, and the update_counter is reset to 0. After the end of each round, the update step length C of the next round is adjusted according to the change of the average loss L ep and the total reward R ep of the round. For example, a reference loss L base and a reference reward R base may be set. ΔL = L ep - L last_ep , and ΔR = R ep - R last_ep . The update of the preset step length is C ← C - β L · sgn(ΔL) + β R · sgn(ΔR). Wherein, β L , β R is the learning rate. If the loss is decreasing, i.e. ΔL < 0, and / or the reward is increasing, i.e. ΔR > 0, C can be appropriately increased; on the contrary, if the model performance is fluctuating or decreasing, i.e. ΔL > 0 and / or ΔR < 0, C is decreased to stabilize. At the same time, C can be limited within a reasonable range [C min , C max ], C min is the lower limit of the preset step length, and C max is the upper limit of the preset step length.
[0103] The technical scheme of the embodiment of the application trains a deep Q network to plan a dispensing path, considers the dynamic motion time of the dispensing equipment in the dispensing path planning, and optimizes the time on the path planning, so as to balance the influence of long-distance uniform speed and short-distance frequent start-stop; and combines obstacle avoidance to ensure the safety of the dispensing equipment, and combines the energy consumption of the dispensing equipment to consider the economy of dispensing, so that the comprehensive benefit of dispensing is far superior to the algorithm based on distance or general reward, and the production efficiency of dispensing is improved; by considering obstacle avoidance and energy consumption, the optimization of the specific needs of dispensing can be tailored, and the specificity and creativity of dispensing are improved. In the deep Q network training, a dynamic discount factor and an adaptive target network update step are used to make the training process of the deep Q network more intelligent and efficient, and the self hyperparameter adjustment is performed according to the learning progress. The dispensing path planning result is quickly converged and has high quality. In application, only one network forward propagation and argmax operation are needed, which is fast and suitable for real-time or near real-time application. In the subsequent process, the deep Q network can be retrained or fine-tuned offline or online by collecting new data, so as to adapt to equipment aging or process changes.
[0104] Figure 2b A dispensing path planning application structure schematic diagram is shown. As shown in Figure 2b , the dispensing path planning method can be applied to a dispensing equipment. A motion controller can be provided in the dispensing equipment, and a motion control module can be provided in the motion controller to control the dispensing head. A point storage module can also be provided in the motion controller to store the current position of the dispensing task and the access state of each to-be-dispensed point. The deep Q network training for dispensing path planning can be performed by an external or internal high-performance computing server in the dispensing path planning. A time calculation module can be provided in the computing server to calculate the dispensing motion time of the kinematics of the reward function. A deep Q network training module can also be provided in the computing server to train the deep Q network. A storage unit can also be provided in the dispensing equipment to store dispensing data, model parameters, device parameters, and experience pool. A human-computer interaction interface can be provided in or outside the dispensing equipment to load tasks, monitor deep Q network training, and visualize dispensing paths.
[0105] Embodiment three
[0106] Figure 3 A dispensing path planning device structure schematic diagram is provided according to the embodiment three of the application. As shown in Figure 3 , the device includes a dispensing task data acquisition module 310, a next dispensing point determination module 320, a deep Q network training module 330, and a dispensing path planning result determination module 340. Wherein:
[0107] The point gluing task data acquisition module 310 is configured to acquire the current position of the point gluing head and the access state of each to-be-point-glued point in the point gluing task.
[0108] The next point gluing point determination module 320 is configured to input the current position and the access state of each to-be-point-glued point into a main network of a pre-trained deep Q network, and determine the next point gluing point according to a main network Q value output by the main network.
[0109] The deep Q network training module 330 is configured to determine a reward function in the deep Q network when training the deep Q network, so as to make the deep Q network take the reward function as an optimization target, by the following manner: determining a point gluing motion time from the current position to a next position corresponding to the next point gluing point according to the acceleration, deceleration and constant speed motion process of the point gluing device when the point gluing head moves from the current position to the next position; and determining the reward function in the deep Q network according to the point gluing motion time.
[0110] The point gluing path planning result determination module 340 is configured to update the current position of the point gluing head and the access state of each to-be-point-glued point according to the next point gluing point, return the step of determining the next point gluing point by the main network Q value output by the main network of the deep Q network, and repeat until all the point gluing points are accessed, so as to obtain a point gluing path planning result.
[0111] Optionally, the deep Q network training module 330 comprises:
[0112] The point gluing motion time determination unit is configured to determine a point gluing motion time from the current position to a next position corresponding to the next point gluing point according to the acceleration, deceleration and constant speed motion process of the point gluing device when the point gluing head moves from the current position to the next position.
[0113] The point gluing obstacle avoidance penalty parameter determination unit is configured to determine a point gluing obstacle avoidance penalty parameter from the current position to the next position according to the proximity between the movement path of the point gluing head from the current position to the next position and a preset obstacle region in the point gluing task.
[0114] The point gluing energy consumption parameter determination unit is configured to determine a point gluing energy consumption parameter from the current position to the next position according to the number of acceleration times, the number of deceleration times, the acceleration time and the deceleration time in the process of moving the point gluing head from the current position to the next position.
[0115] The reward function determination unit is configured to determine the reward function in the deep Q network according to the point gluing obstacle avoidance penalty parameter and / or the point gluing energy consumption parameter, and the point gluing motion time.
[0116] Optionally, the point gluing obstacle avoidance penalty parameter determination unit is specifically configured to:
[0117] According to a movement path of the dispensing head from the current position to the next position, a shortest distance from the movement path to a preset obstacle region in the dispensing task is calculated;
[0118] According to a reciprocal of the shortest distance and a reciprocal of a preset dispensing safety margin, a dispensing obstacle avoidance penalty parameter from the current position to the next position is determined.
[0119] Optionally, the deep Q network training module 330 comprises:
[0120] A main network construction unit is configured to construct a main network of the deep Q network, and take the current position and the access states of the dispensing points as input data of the main network;
[0121] A next dispensing point determination unit is configured to determine, in iterative decision of the main network, a next dispensing point according to a main network Q value output by the main network;
[0122] A reward function determination unit is configured to determine a reward function of the dispensing action according to the current position and a next position of the dispensing head when moving to the next dispensing point;
[0123] An experience pool generation unit is configured to generate an experience pool according to the current position, the next dispensing point, the next position, the reward function, and the access states of the dispensing points;
[0124] A target Q value calculation unit is configured to construct a target network of the deep Q network, and calculate a target Q value according to the reward function, a discount factor, and a target network output Q value of a sample in the experience pool;
[0125] Wherein, parameters in the target network are determined by copying parameters from the main network every preset step length;
[0126] A deep Q network training unit is configured to obtain a trained deep Q network by iterative training according to a loss function determined by the main network Q value and the target Q value.
[0127] Optionally, the target Q value calculation unit is specifically configured to determine the discount factor in the following manner:
[0128] According to the current position of the dispensing head and dispensing points not accessed in a preset region of the current position, a point density of a region where the dispensing head is located is determined;
[0129] The discount factor is determined according to the point density.
[0130] Optionally, the target Q value calculation unit is specifically configured to determine the preset step length when copying the main network parameters in the following manner:
[0131] The preset step length when copying the main network parameters is determined according to a loss function change value and / or a reward function change value in the deep Q network training process.
[0132] The dispensing path planning device provided by the embodiments of the present application can execute the dispensing path planning method provided by any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.
[0133] Embodiment Four
[0134] Figure 4 is a structural schematic diagram of a deep Q network training device for dispensing path planning provided by Embodiment Four of the present application. As shown in the figure, the device comprises a dispensing task data acquisition module 410, a main network construction module 420, a next dispensing point position determination module 430, a reward function determination module 440, an experience pool generation module 450, a target Q value calculation module 460, and a deep Q network training module 470. Among them: Figure 4
[0135] The dispensing task data acquisition module 410 is configured to acquire the current position of the dispensing head and the access state of each to-be-dispensed point in the dispensing task.
[0136] The main network construction module 420 is configured to construct the main network of the deep Q network, and take the current position and the access state of each to-be-dispensed point as the input data of the main network.
[0137] The next dispensing point position determination module 430 is configured to determine the next dispensing point position according to the main network Q value output by the main network in the iterative decision of the main network.
[0138] The reward function determination module 440 is configured to determine the dispensing motion time from the current position to the next position according to the acceleration, deceleration, and uniform motion process of the dispensing equipment when the dispensing head moves from the current position to the next position corresponding to the next dispensing point position, and determine the reward function in the deep Q network according to the dispensing motion time.
[0139] The experience pool generation module 450 is configured to generate an experience pool according to the current position, the next dispensing point position, the next position, the reward function, and the access state of each to-be-dispensed point.
[0140] The target Q value calculation module 460 is configured to construct the target network of the deep Q network, and calculate the target Q value according to the reward function, the discount factor, and the target network output Q value of the sample in the experience pool.
[0141] Among them, the parameters in the target network are determined by copying the parameters from the main network every preset step length.
[0142] The deep Q network training module 470 is configured to perform iterative training by determining the loss function according to the main network Q value and the target Q value, obtain the trained deep Q network, and perform dispensing path planning according to the trained deep Q network.
[0143] The deep Q network training device for dispensing path planning provided by the embodiments of the present application can execute the deep Q network training method for dispensing path planning provided by any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.
[0144] Embodiment five
[0145] Figure 5 A structural schematic diagram of a dispensing device 10 that can be used to implement embodiments of the present application is shown. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0146] As shown in Figure 5 The dispensing device 10 includes a motion platform 21, a dispensing head 22, a position sensor 23, a motion controller 24, at least one processor 11, and a memory in communication with the at least one processor 11. The motion platform is used to place a workpiece to be dispensed; the dispensing device performs dispensing operations on the workpiece to be dispensed while the motion platform moves; the dispensing head is used to perform dispensing operations on the workpiece to be dispensed on the motion platform at dispensing points;
[0147] The position sensor is used to obtain the position of the dispensing head; the motion controller is used to control the movement of the dispensing head; the memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the dispensing path planning method provided by any of the embodiments of the present application, or the deep Q network training method for dispensing path planning.
[0148] The memory can be a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the processor 11 can perform various appropriate actions and processes according to a computer program stored in the read-only memory (ROM) 12 or a computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the dispensing device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0149] A plurality of components in the dispensing apparatus 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the dispensing apparatus 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0150] The processor 11 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the dispensing path planning method, or the deep Q-network training method for dispensing path planning.
[0151] In some embodiments, the dispensing path planning method, or the deep Q-network training method for dispensing path planning can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the dispensing apparatus 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded onto the RAM 13 and executed by the processor 11, one or more steps of the dispensing path planning method, or the deep Q-network training method for dispensing path planning described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the dispensing path planning method, or the deep Q-network training method for dispensing path planning by any other appropriate means, such as by means of firmware.
[0152] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a load programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0153] Computer programs used to implement the processes of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program
[0154] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store computer programs for use by or in connection with an instruction execution system, apparatus, or device. Computer-readable storage media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0155] To provide for interaction with a user, the systems and techniques described here can be implemented on a dispensing device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the dispensing device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0156] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0157] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.
[0158] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in sequence, or executed in a different order, as long as the desired results of the present disclosure are achieved, and the present disclosure is not limited herein.
[0159] The specific embodiments described above are not intended to be limiting, and persons skilled in the art will appreciate that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the spirit and principles of the disclosure. Accordingly, the disclosure is not limited to the specific embodiments described above, but only by the scope of the appended claims.
Claims
1. A dispensing path planning method, characterized in that, include: Get the current position of the dispensing head and the access status of each dispensing point in the dispensing task; The current position and the access status of each glue dot to be applied are input into the main network of the pre-trained deep Q network, and the next glue dot is determined based on the Q value of the main network output by the main network. Specifically, when training a deep Q-network, the reward function in the deep Q-network is determined in the following way so that the deep Q-network uses the reward function as the optimization objective: when the dispensing head moves from the current position to the next position corresponding to the next dispensing point, the dispensing movement time from the current position to the next position is determined according to the acceleration, deceleration and constant speed movement process of the dispensing device; and the reward function in the deep Q-network is determined according to the dispensing movement time. Update the current position of the dispensing head and the access status of each dispensing point according to the next dispensing point. Return the main network Q value output by the main network of the deep Q network to determine the next dispensing point step, until all dispensing points are accessed, and obtain the dispensing path planning result.
2. The method according to claim 1, characterized in that, When training a deep Q-network, the reward function in the deep Q-network is determined in the following way so that the deep Q-network optimizes the reward function: When the dispensing head moves from the current position to the next position corresponding to the next dispensing point, the dispensing movement time from the current position to the next position is determined based on the acceleration, deceleration and uniform speed movement process of the dispensing equipment. Based on the proximity between the dispensing head's movement path from the current position to the next position and the preset obstacle area in the dispensing task, determine the dispensing obstacle avoidance penalty parameters from the current position to the next position; Based on the number of accelerations, decelerations, acceleration time, and deceleration time during the process of the dispensing head moving from the current position to the next position, determine the dispensing energy consumption parameters from the current position to the next position; The reward function in the deep Q-network is determined based on the dispensing obstacle avoidance penalty parameters and / or dispensing energy consumption parameters, as well as the dispensing motion time.
3. The method according to claim 2, characterized in that, Based on the proximity of the dispensing head's movement path from the current position to the next position to the preset obstacle area in the dispensing task, determine the dispensing obstacle avoidance penalty parameters from the current position to the next position, including: Based on the movement path of the dispensing head from the current position to the next position, calculate the shortest distance from the movement path to the preset obstacle area in the dispensing task; The dispensing obstacle avoidance penalty parameters from the current position to the next position are determined based on the reciprocal of the shortest distance and the reciprocal of the preset dispensing safety margin.
4. The method according to claim 1, characterized in that, Train a deep Q-network using the following method: Construct a deep Q-network as the main network, and use the current position and the access status of each glue dispensing point as the input data of the main network; In the iterative decision-making of the main network, the next glue point is determined based on the Q value of the main network output. The reward function for the dispensing action is determined based on the current position and the next position when the dispensing head moves to the next dispensing point. An experience pool is generated based on the current position, the next glue point, the next position, the reward function, and the access status of each glue point to be applied. Construct the target network of the deep Q-network, and calculate the target Q-value based on the reward function, discount factor, and Q-value of the samples in the experience pool; In this process, parameters are copied from the main network at preset intervals to determine the parameters in the target network; By iteratively training the loss function determined based on the Q-value of the main network and the target Q-value, a trained deep Q-network is obtained.
5. The method according to claim 4, characterized in that, When training a deep Q-network, the discount factor is determined as follows: Determine the dot density of the area where the dispensing head is located based on the current position of the dispensing head and the unvisited dispensing points within the preset area of the current position. The discount factor is determined based on the point density.
6. The method according to claim 4, characterized in that, When training a deep Q-network, the preset step size for copying the main network parameters is determined as follows: The preset step size for replicating the main network parameters is determined based on the changes in the loss function and / or reward function during the training process of the deep Q-network.
7. A deep Q-network training method for dispensing path planning, characterized in that, include: Get the current position of the dispensing head and the access status of each dispensing point in the dispensing task; Construct a deep Q-network as the main network, and use the current position and the access status of each glue dispensing point as the input data of the main network; In the iterative decision-making of the main network, the next glue point is determined based on the Q value of the main network output. When the dispensing head moves from the current position to the next position corresponding to the next dispensing point, the dispensing movement time from the current position to the next position is determined based on the acceleration, deceleration and constant speed movement process of the dispensing equipment; and the reward function in the deep Q network is determined based on the dispensing movement time. An experience pool is generated based on the current position, the next glue point, the next position, the reward function, and the access status of each glue point to be applied. Construct the target network of the deep Q-network, and calculate the target Q-value based on the reward function, discount factor, and Q-value of the samples in the experience pool; In this process, parameters are copied from the main network at preset intervals to determine the parameters in the target network; By iteratively training the loss function determined based on the Q-value of the main network and the target Q-value, a trained deep Q-network is obtained, which is then used for dispensing path planning.
8. A dispensing device, characterized in that, The dispensing device includes: a motion platform, a dispensing head, a position sensor, a motion controller, at least one processor, and a memory communicatively connected to the at least one processor; wherein: The motion platform is used to place the workpiece to be glued; the glue dispensing equipment moves on the motion platform to perform the glue dispensing operation on the workpiece to be glued. A dispensing head is used to dispense adhesive onto a workpiece on a moving platform at a dispensing point. A position sensor is used to determine the position of the dispensing head; Motion controller, used to control the movement of the dispensing head; The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the dispensing path planning method according to any one of claims 1-6, or the deep Q-network training method for dispensing path planning according to claim 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the dispensing path planning method according to any one of claims 1-6, or the deep Q-network training method for dispensing path planning according to claim 7.
10. A computer program product comprising a computer program that, when executed by a processor, implements the dispensing path planning method according to any one of claims 1-6, or the deep Q-network training method for dispensing path planning according to claim 7.
Citation Information
Cited By
Dispensing path planning method and system for random bulk materials
CN122239581A