Water-based adhesive spraying path planning method and system based on reinforcement learning
Through the reinforcement learning water-based glue spraying path planning method, the strategy network and value network are used to optimize the spraying path and generate a global guidance trajectory, which solves the problem of dynamic adjustment of the spray gun posture and speed in the spraying of complex workpieces, realizes the uniformity control of coating thickness, and improves the spraying quality.
Patent Information
- Application Number
- CN202511120276.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-08-12
AI Technical Summary
Existing spray path planning makes it difficult to dynamically adjust the spray gun posture and speed on complex workpieces, and cannot accurately control the uniformity of coating thickness. In addition, reinforcement learning lacks guidance in spraying quality and cannot effectively improve spraying quality.
By obtaining the three-dimensional model of the workpiece and the spraying process parameters, the strategy network is used to interact with the environment to generate trajectory data, the variance of the advantage estimate is calculated, the training data set is constructed to update the value network, the strategy network is optimized, and the global guidance trajectory is generated by combining the landmark set and the node heat map. The landmark contribution value is adjusted to generate an executable spray path instruction sequence.
The learning efficiency and quality control capabilities of spray path planning are improved, ensuring that the spray quality meets the process requirements and improving the spraying effect of complex workpieces.
Smart Images

Figure CN120605849A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to a method and system for water-based glue spraying path planning based on reinforcement learning. Background Art
[0002] Water-based glue is an environmentally friendly adhesive. In order to ensure the quality of water-based glue spraying, high requirements are placed on coating thickness, coverage uniformity, etc. during the spraying process. However, in the spraying of complex workpieces such as deep cavities and sharp edges, manual spraying has problems such as unstable glue coating and spray leakage, and the quality of spraying cannot be guaranteed. Spraying path planning can generate a spraying path based on the characteristics of the workpiece, and compared to manual spraying, it can accurately control the spraying quality. It is particularly suitable for water-based glue spraying of complex workpieces. However, one problem faced is that the existing spray path planning focuses on path coverage, and lacks support for how to dynamically adjust the spray gun posture, speed, and distance from the workpiece to accurately control the uniformity of the coating thickness. Reinforcement learning's self-learning and optimization capabilities in complex decision-making problems can be applied to robot path planning. However, there are still many problems in directly applying reinforcement learning to actual spraying. The value function or strategy obtained by reinforcement learning is usually an evaluation of local state or action. How to convert it into a spraying path that guides the overall situation and meets process constraints, and when the spraying quality does not meet the standards, how to extract valuable information from it to improve the strategy and value judgment in a targeted manner and avoid repeated and inefficient exploration in invalid areas, is the key to improving the quality of water-based adhesive spraying. Summary of the Invention
[0003] To address the problems of low learning efficiency, insufficient utilization of valuable information, and insufficient guidance on spraying quality in reinforcement learning in spraying path planning, this application proposes a water-based adhesive spraying path planning method based on reinforcement learning, including: Obtaining a three-dimensional model of the workpiece and preset spraying process parameters, the policy network interacts with the spraying environment simulation and generates trajectory data containing states, actions, and rewards. The value network then calculates an advantage estimate and the variance of the advantage estimate. When the monitored spraying quality index does not meet the preset performance benchmark, the trajectory that does not meet the quality standard is obtained from the trajectory, a training data set is constructed based on the variance of the advantage estimate, and the error of the training data set is used to update the value network; the advantage estimate is calculated based on the updated value network and the parameters of the strategy network are optimized; The workpiece surface is converted into a discrete graph structure, a landmark set is established, and the cost between each point is pre-calculated; a node heat map of the workpiece surface is obtained using the value network; a landmark subset corresponding to the current spraying area is determined from the landmark set based on the geometric properties of the current spraying area and the spraying accuracy requirements, and the contribution value of the landmarks in the landmark subset in the path search of the current spraying area is obtained using the node heat map; When the actual surface morphology of the workpiece deviates from the three-dimensional model beyond a preset value, the contribution value of the landmark related to the target spraying area is adjusted; the global guidance trajectory is calculated on the graph structure using the landmark subset and the contribution value, and an executable spray path instruction sequence is generated based on the strategy network and the global guidance trajectory.
[0004] Optionally, constructing a training data set according to the variance of the advantage estimate includes: Compute the variance of the advantage estimate for each state-action pair within a time segment in the trajectory; All trajectory data with a variance of the advantage estimate greater than the variance threshold, or trajectory data ranked in the top N percent of the variance, are added to the training data set.
[0005] Optionally, obtaining a node heat map of the workpiece surface using the value network includes: Extract the three-dimensional spatial position coordinates, surface normal vector, and preset spraying process parameters of each node to obtain the state vector; Each state vector is input into the trained value network to obtain the evaluation value; the evaluation values of all nodes are normalized to obtain the normalized node heat map.
[0006] Optionally, obtaining contribution values of landmarks in the landmark subset in the current spraying area path search using the node heat map includes: From a pre-set global landmark set, all landmarks whose geometric positions are located within the current spraying area and landmarks whose distance from the current spraying area is less than a predetermined threshold are obtained, thereby obtaining a current landmark subset; For each landmark in the landmark subset , determine the set of neighboring nodes N with the nearest neighbor in Euclidean distance; Calculate the average value of each node evaluation value in the neighborhood node set and use the average value as the landmark Contribution value in the path search process .
[0007] Optionally, adjusting the contribution value of the landmark associated with the target spraying area includes: Get the upper threshold of positive deviation and the negative deviation lower threshold ; For each landmark in the target spraying area , calculated with The average value Avg of the normal deviation of all actual surface points in the influence area with a radius of R as the center; If Avg is greater than , then according to the formula To reduce the landmark The contribution value of is the contribution value attenuation coefficient; If Avg is less than , then according to the formula To reduce the landmark The contribution value of is the contribution value attenuation coefficient.
[0008] Optionally, calculating a global guidance trajectory on the graph structure using the landmark subset and the contribution value includes: Based on the discretized graph structure G=(V,E) of the workpiece surface, V is a set of nodes representing each tiny area on the surface, and E is a set of edges representing nodes that can be directly connected for spraying operations; each edge e(u,v) connects node u and node v and has a basic path cost value ,The basic path cost value is determined based on the three-dimensional Euclidean distance between nodes u and v or the estimated spraying time; Determine the starting node and ending node of the current spraying task in the graph structure G; For each edge e(u,v) in the graph structure G, get all landmarks whose distance to the midpoint of the edge e(u,v) is less than the influence radius R , the edge adjustment factor is calculated according to the following formula ,in Is a nearby landmark The contribution value of It is a landmark The weight of the influence on edge e(u,v); Calculate the cost of each edge in the graph structure G , where is the basic path cost of edge e(u,v); A path search method is used to search for a path with the minimum cumulative cost from the starting node to the ending node as the global guiding trajectory.
[0009] Optionally, generating an executable spray path instruction sequence according to the policy network and the global guidance trajectory includes: The global guidance trajectory is refined into a series of path points through an interpolation algorithm. Each path point contains the three-dimensional position coordinates, the desired spray gun end posture, and the spray start and stop status flag; the current state of the spray gun is initialized; Starting from the first path point, in each discrete servo control cycle, the current state of the spray gun and the current path point information are combined into a state vector. After receiving the input state vector, the policy network outputs an action instruction vector. The target speed and posture instructions in the action instruction vector are converted into specific joint motion instructions. The generated specific joint motion instructions are added to a first-in-first-out instruction sequence buffer. After the robot motion controller executes the instructions of the current cycle, the current state of the spray gun is updated. Calculate the distance between the current gun position and the current path point. If the distance is less than the tolerance and the end point of the trajectory has not been reached, the next path point will be used as the current path point, until the last path point.
[0010] This application also proposes a water-based glue spraying path planning system based on reinforcement learning, including: A data generation unit is used to obtain a three-dimensional model of the workpiece and preset spraying process parameters. The strategy network interacts with the spraying environment simulation and generates trajectory data including states, actions, and rewards. The value network calculates the advantage estimate and the variance of the advantage estimate. an iterative optimization unit for obtaining, when the monitored spray quality indicator fails to meet a preset performance benchmark, a trajectory that fails to meet the quality benchmark from the trajectory, constructing a training data set based on the variance of the advantage estimate, and updating the value network using the error of the training data set; and optimizing the parameters of the strategy network based on the advantage estimate calculated by the updated value network; A graph construction unit is configured to convert the workpiece surface into a discrete graph structure, establish a landmark set, and pre-calculate the cost between each point; utilize the value network to obtain a node heat map of the workpiece surface; determine a landmark subset corresponding to the current spraying area from the landmark set based on the geometric properties of the current spraying area and the spraying accuracy requirements; and utilize the node heat map to obtain the contribution value of the landmarks in the landmark subset in the path search of the current spraying area; A path generation unit is used to adjust the contribution values of landmarks related to the target spraying area when the actual surface morphology of the workpiece and the three-dimensional model exceed a preset deviation; use the landmark subset and the contribution values to calculate the global guidance trajectory on the graph structure, and generate an executable spray path instruction sequence based on the strategy network and the global guidance trajectory.
[0011] Optionally, constructing a training data set according to the variance of the advantage estimate includes: Compute the variance of the advantage estimate for each state-action pair within a time segment in the trajectory; All trajectory data with a variance of the advantage estimate greater than the variance threshold, or trajectory data ranked in the top N percent of the variance, are added to the training data set.
[0012] Optionally, obtaining a node heat map of the workpiece surface using the value network includes: Extract the three-dimensional spatial position coordinates, surface normal vector, and preset spraying process parameters of each node to obtain the state vector; Each state vector is input into the trained value network to obtain the evaluation value; the evaluation values of all nodes are normalized to obtain the normalized node heat map.
[0013] Optionally, obtaining contribution values of landmarks in the landmark subset in the current spraying area path search using the node heat map includes: From a pre-set global landmark set, all landmarks whose geometric positions are located within the current spraying area and landmarks whose distance from the current spraying area is less than a predetermined threshold are obtained, thereby obtaining a current landmark subset; For each landmark in the landmark subset , determine the set of neighboring nodes N with the nearest neighbor in Euclidean distance; Calculate the average value of each node evaluation value in the neighborhood node set and use the average value as the landmark Contribution value in the path search process .
[0014] Optionally, adjusting the contribution value of the landmark associated with the target spraying area includes: Get the upper threshold of positive deviation and the negative deviation lower threshold ; For each landmark in the target spraying area , calculated with The average value Avg of the normal deviation of all actual surface points in the influence area with a radius of R as the center; If Avg is greater than , then according to the formula To reduce the landmark The contribution value of is the contribution value attenuation coefficient; If Avg is less than , then according to the formula To reduce the landmark The contribution value of is the contribution value attenuation coefficient.
[0015] Optionally, calculating a global guidance trajectory on the graph structure using the landmark subset and the contribution value includes: Based on the discretized graph structure G=(V,E) of the workpiece surface, V is a set of nodes representing each tiny area on the surface, and E is a set of edges representing nodes that can be directly connected for spraying operations; each edge e(u,v) connects node u and node v and has a basic path cost value ,The basic path cost value is determined based on the three-dimensional Euclidean distance between nodes u and v or the estimated spraying time; Determine the starting node and ending node of the current spraying task in the graph structure G; For each edge e(u,v) in the graph structure G, get all landmarks whose distance to the midpoint of the edge e(u,v) is less than the influence radius R , the edge adjustment factor is calculated according to the following formula ,in Is a nearby landmark The contribution value of It is a landmark The weight of the influence on edge e(u,v); Calculate the cost of each edge in the graph structure G , where is the basic path cost of edge e(u,v); A path search method is used to search for a path with the minimum cumulative cost from the starting node to the ending node as the global guiding trajectory.
[0016] Optionally, generating an executable spray path instruction sequence according to the policy network and the global guidance trajectory includes: The global guidance trajectory is refined into a series of path points through an interpolation algorithm. Each path point contains the three-dimensional position coordinates, the desired spray gun end posture, and the spray start and stop status flag; the current state of the spray gun is initialized; Starting from the first path point, in each discrete servo control cycle, the current state of the spray gun and the current path point information are combined into a state vector. After receiving the input state vector, the policy network outputs an action instruction vector. The target speed and posture instructions in the action instruction vector are converted into specific joint motion instructions. The generated specific joint motion instructions are added to a first-in-first-out instruction sequence buffer. After the robot motion controller executes the instructions of the current cycle, the current state of the spray gun is updated. Calculate the distance between the current gun position and the current path point. If the distance is less than the tolerance and the end point of the trajectory has not been reached, the next path point will be used as the current path point, until the last path point.
[0017] The present invention constructs a training data set based on the variance of the advantage estimate after the strategy network interacts with the environment, so that the value network can preferentially learn from trajectory data with higher uncertainty or potential for improvement, thereby improving the pertinence and efficiency of the learning process; the trained value network is used to obtain a node heat map of the workpiece surface, and the node heat map is used to obtain the contribution value of the landmarks in the landmark subset in the path search of the current spraying area, and the learning-based value judgment is integrated into the generation of the global guidance trajectory, so that path planning no longer relies solely on geometric information, thereby planning a global path that is more conducive to improving the spraying quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a flow chart of Example 1; Figure 2 This is a visualization diagram of the node heat map on the workpiece surface; Figure 3 Schematic diagram of the global guidance trajectory; Figure 4 Schematic diagram of the global guidance trajectory and path segments. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0020] In a specific embodiment, this application proposes a water-based glue spraying path planning method based on reinforcement learning, such as Figure 1 Shown, including: Step 1: Obtain a 3D model of the workpiece and preset spraying process parameters. The policy network interacts with the spraying environment simulation and generates trajectory data containing states, actions, and rewards. The value network then calculates the advantage estimate and the variance of the advantage estimate. Obtain a three-dimensional model of the workpiece to be sprayed, such as a standard format file generated by CAD software, and load preset spray process parameters, such as the flow rate of the paint, the rated movement speed of the spray gun, the target coating thickness, etc. The policy network included in the reinforcement learning interacts with the spray simulation environment. The simulation environment can simulate the actual physical process of spraying, including fluid dynamics effects and the kinematic characteristics of the robot manipulator. During the interaction process, the policy network generates a series of empirical trajectories through continuous trial and error. Each trajectory contains state information such as the current coverage of the workpiece surface, the real-time position and posture of the spray gun, such as the control instructions for the spray gun movement, and the actions taken by the policy network, as well as feedback reward signals, such as evaluation values quantified based on the spray quality and efficiency. Among them, the policy network is a deep neural network, preferably a policy network with an Actor-Critic architecture.
[0021] The value network is used to assess the expected cumulative reward from a specific state or action. Based on the value network's evaluation results, a dominance estimate is calculated for each state-action pair in the trajectory. Preferably, a generalized dominance estimate is used. This dominance value describes the superiority of the selected action relative to the average action in the current state. The statistical variance of the dominance estimate is calculated over the entire trajectory data or within a predefined segment. The value network preferably employs an actor-critic architecture.
[0022] Step 2: When the monitored spray quality indicator does not meet the preset performance benchmark, the trajectory that does not meet the quality standard is obtained from the trajectory, a training data set is constructed based on the variance of the advantage estimate, and the error of the training data set is used to update the value network; the advantage estimate is calculated based on the updated value network and the parameters of the strategy network are optimized; When online monitoring reveals that spray quality indicators fail to meet preset performance benchmarks, such as the uniformity of the coating, the effective utilization of the paint, or the coverage integrity of a specific area, these can be obtained through virtual sensors or analytical models, triggering a learning and optimization mechanism. First, the empirical trajectory that led to the substandard spray results is identified. To guide subsequent training, a training data set is constructed based on the variance of the advantage estimate calculated in the previous step. Preferably, the variance of the advantage estimate is used to prioritize the empirical samples, implementing a prioritized experience replay mechanism, that is, samples with larger variances are sampled for training with a higher probability.
[0023] The constructed training dataset value network is used to update parameters by minimizing the loss function. After the value network is improved, the parameters of the policy network will also be further optimized. For example, gradient algorithms such as proximal policy optimization are used. The more accurate advantage estimates provided by the updated value network are used to guide the iterative improvement of the strategy, thereby guiding the intelligent agent to learn better spraying behavior.
[0024] Step 3: Convert the workpiece surface into a discrete graph structure, establish a landmark set, and pre-calculate the cost between each point; use the value network to obtain a node heat map of the workpiece surface; determine the landmark subset corresponding to the current spraying area from the landmark set based on the geometric properties and spraying accuracy requirements of the current spraying area, and use the node heat map to obtain the contribution value of the landmarks in the landmark subset in the path search of the current spraying area; The complex three-dimensional surface of the workpiece is converted into a discrete graph structure representation. In one embodiment, the workpiece model is voxelized, or a surface mesh graph is generated directly from its CAD mesh data, where the graph nodes represent discrete units or patches on the surface, and the edges represent the adjacency relationship or sprayable connectivity between them. At the same time, a set of landmarks is defined. The landmarks are key points selected based on the geometric features of the workpiece, or points in areas that are difficult to spray well. The initial travel costs between these landmarks, that is, between the graph nodes, are pre-calculated, such as the surface geodesic distance or the estimated spraying time. The information in the reinforcement learning process is used to generate a node heat map of the workpiece surface. In one embodiment, the entropy value of the action probability distribution output by the policy network at each discrete surface point is calculated, and the high entropy area is displayed as the focus area on the heat map. In one embodiment, the node heat map is a graph for the node, rather than a continuous graph, such as Figure 2 As shown in Figure 2, all nodes and their evaluation values constitute a node heat map.
[0025] For the spraying area currently requiring path planning, a subset of relevant landmarks is selected from the global landmark set based on geometric range, shape characteristics, and spraying accuracy requirements. A node heat map is used to determine the contribution of each landmark in the path search for the current spraying area. In one optional embodiment, the contribution value is calculated using a kernel-based function. In another embodiment, the local heat map pattern around the landmark and the geometric properties of the landmark itself are used as input to a neural network, and the output is the contribution value.
[0026] Step 4: When the actual surface morphology of the workpiece deviates from the three-dimensional model beyond a preset deviation, the contribution value of the landmark related to the target spraying area is adjusted; the global guidance trajectory is calculated on the graph structure using the landmark subset and the contribution value, and an executable spray path instruction sequence is generated based on the strategy network and the global guidance trajectory.
[0027] If an online inspection method such as 3D scanning finds that the actual surface morphology of the workpiece deviates from its theoretical 3D model beyond a preset tolerance, the contribution value of the landmark associated with the target spraying area where the deviation occurs is adjusted. In one embodiment, Gaussian process regression is used for this adjustment. Specifically, the impact on the difficulty of spraying is predicted based on the specific information of the observed surface deviation, and the contribution value of the landmark is updated accordingly. Using the updated landmark contribution value and the previously selected landmark subset, a global guidance trajectory is calculated on the discrete graph structure of the workpiece surface, such as Figure 3 As shown, the preferred search method is RRT. When constructing the search tree, RRT prioritizes exploring and connecting areas influenced by high-contribution landmarks through a biased sampling process or node connection strategy. During each control cycle, the policy network not only outputs instantaneous actions for individual waypoints on the global trajectory, but also predicts the optimal action sequence within a short future time window. This allows the spray path to follow global guidance while making continuous local adjustments and quickly responding to small disturbances, ensuring effective spraying.
[0028] In an optional embodiment, constructing a training data set according to the variance of the advantage estimate includes: Compute the variance of the advantage estimate for each state-action pair within a time segment in the trajectory; All trajectory data with a variance of the advantage estimate greater than the variance threshold, or trajectory data ranked in the top N percent of the variance, are added to the training data set.
[0029] When constructing the training dataset, each state-action pair is extracted from the trajectory generated by the interaction between the policy network and the environment. For example, in a spray painting task, the policy network selects the speed and direction of the spray gun as the action in a certain state and returns the corresponding reward. For each state-action pair, its advantage estimate is calculated—that is, the degree of superiority of the action relative to the average action in the current state. The variance of the advantage estimate is then calculated within a preset time period or time slice. Samples are filtered based on a preset variance threshold or ranking percentage. For example, if the variance threshold is set to 0.1, all state-action pairs with a variance greater than 0.1 are added to the training dataset. Alternatively, samples ranked in the top 10% of the variance can be selected. These filtered samples are then used to update the value network and optimize the policy network, thereby improving the model's performance in complex spray painting tasks. This allows the model to better learn samples with high uncertainty, thereby improving the overall spray path planning performance.
[0030] In an optional embodiment, obtaining a node heat map of the workpiece surface using the value network includes: Extract the three-dimensional spatial position coordinates, surface normal vector, and preset spraying process parameters of each node to obtain the state vector; Each state vector is input into the trained value network to obtain the evaluation value; the evaluation values of all nodes are normalized to obtain the normalized node heat map.
[0031] Specifically, in the process of generating a node heat map, information of each node is extracted from the three-dimensional model of the workpiece. For example, assume that the surface of the workpiece is discretized into multiple nodes, each node has its three-dimensional coordinates and surface normal vector. In addition, preset spraying process parameters such as the flow rate of the paint and the movement speed of the spray gun are also included in the state vector, and this information together constitutes a state vector. Each state vector is input into the trained value network. The value network will output an evaluation value based on the input state vector, which reflects the expected effect of spraying under the current state. For example, if the evaluation value of a node is high, it means that the area may obtain better coverage during the spraying process. The evaluation values of all nodes are normalized to generate a normalized node heat map. This heat map shows the spraying priority of each area on the workpiece surface.
[0032] In an optional embodiment, the step of using the node heat map to obtain the contribution value of a landmark in the landmark subset in the current spraying area path search includes: From a pre-set global landmark set, all landmarks whose geometric positions are located within the current spraying area and landmarks whose distance from the current spraying area is less than a predetermined threshold are obtained, thereby obtaining a current landmark subset; For each landmark L_s in the landmark subset, determine the set of neighboring nodes N that are closest to it in Euclidean distance; Calculate the average value of each node evaluation value in the neighborhood node set and use the average value as the landmark Contribution value in the path search process .
[0033] Specifically, in the process of calculating the landmark contribution value, the landmarks related to the current spraying area are screened from the pre-set global landmark set. For example, assuming that the current spraying area is a specific workpiece surface area, the geometric position of each landmark is checked to determine whether it is located in the area or whether the distance to the area is less than a preset threshold such as 5 mm, and a current landmark subset is constructed. For each landmark in the landmark subset, , determine the set of neighboring nodes N with the nearest Euclidean distance. For example, if the landmark The coordinates are (x, y, z), the system will calculate all nodes to The Euclidean distance is calculated and the k nodes with the smallest distance are selected as the neighborhood node set N. Then, the average evaluation value of each node in the neighborhood node set is calculated and the average value is used as the landmark The contribution value C_s during the path search process. For example, if the node evaluation values in the neighborhood node set N are 0.8, 0.7, and 0.9, respectively, then the contribution value C_s of the landmark L_s will be (0.8+0.7+0.9) / 3=0.8. This contribution value can effectively guide the path search, ensure that the spray path covers key areas, and improve the overall spray effect.
[0034] In an optional embodiment, adjusting the contribution value of the landmark associated with the target spraying area includes: Get the upper threshold of positive deviation and the negative deviation lower threshold ; For each landmark in the target spraying area , calculated with The average value Avg of the normal deviation of all actual surface points in the influence area with a radius of R as the center; If Avg is greater than , then according to the formula To reduce the landmark The contribution value of is the contribution value attenuation coefficient; If Avg is less than , then according to the formula To reduce the landmark The contribution value of is the contribution value attenuation coefficient.
[0035] Specifically, in the process of adjusting the landmark contribution value, it is first necessary to obtain the upper limit threshold of the positive deviation and the negative deviation lower threshold Assumptions is 0.1, -0.1 indicates the maximum positive and negative deviation allowed. , calculated with The average value Avg of the normal deviation of all actual surface points within the influence area with a radius R of 10 mm as the center. For example, if there are 5 points in the influence area and their normal deviations are 0.05, 0.08, 0.12, 0.15 and 0.18 respectively, then Avg is 0.116. If Avg is greater than ,For example, , using the formula To reduce the contribution value of the landmark L_s, where is the contribution value attenuation coefficient, preferably 0.5. For example, if The initial value is 0.8, so the new contribution value will be 0.736. If Avg is less than ,For example, , using the formula To reduce the landmark The contribution value of is the contribution value attenuation coefficient, which is preferably 0.5. For example, if The initial value is 0.8, so the new contribution value In this way, the contribution value of the landmark is dynamically adjusted to ensure the adaptability and accuracy of the spraying path.
[0036] In an optional embodiment, calculating a global guidance trajectory on the graph structure using the landmark subset and the contribution value includes: Based on the discretized graph structure G=(V,E) of the workpiece surface, V is a set of nodes representing each tiny area on the surface, and E is a set of edges representing nodes that can be directly connected for spraying operations; each edge e(u,v) connects node u and node v and has a basic path cost value ,The basic path cost value is determined based on the three-dimensional Euclidean distance between nodes u and v or the estimated spraying time; Determine the starting node and ending node of the current spraying task in the graph structure G; For each edge e(u,v) in the graph structure G, get all landmarks whose distance to the midpoint of the edge e(u,v) is less than the influence radius R , the edge adjustment factor is calculated according to the following formula ,in Is a nearby landmark The contribution value of It is a landmark The weight of the influence on edge e(u,v); Calculate the cost of each edge in the graph structure G , where is the basic path cost of edge e(u,v); A path search method is used to search for a path with the minimum cumulative cost from the starting node to the ending node as the global guiding trajectory.
[0037] Specifically, in the process of calculating the global guidance trajectory, according to the discretized graph structure G=(V,E) of the workpiece surface, V is a set of nodes representing each tiny area on the surface, and E is a set of edges representing the nodes that can be directly connected for spraying operations. Assume that the workpiece surface is discretized into multiple nodes, each node represents a tiny spraying area, and the edge represents the connection relationship between these areas. Each edge e(u,v) connects node u and node v and has a basic path cost value , the cost value can be determined based on the three-dimensional Euclidean distance between nodes u and v or the estimated spraying time. For example, if the Euclidean distance between nodes u and v is 5 mm, then the basic path cost value is 5. Determine the starting node and the end node of the current spraying task in the graph structure G. For example, the starting node is a specific position on the surface of the workpiece, and the end node is another position, which are the starting point and end point of the spraying task respectively. For each edge e(u,v) in the graph structure G, the system will obtain all landmarks whose distance to the midpoint of the edge e(u,v) is less than the influence radius R For example, if the influence radius R is 10 mm, find the landmarks whose distance from the edge midpoint is less than 10 mm. Based on the contribution of these landmarks and weights , calculate the adjustment factor of the edge For example, if the nearby landmark Contribution value is 0.8, weight If 0.5, then the adjustment factor AF is 0.714. Calculate the cost of each edge in the graph structure G , and uses the path search method to search for a path with the minimum cumulative cost from the starting node to the end node as the global guiding trajectory.
[0038] In an optional embodiment, generating an executable spray path instruction sequence based on the policy network and the global guidance trajectory includes: The global guidance trajectory is refined into a series of path points through interpolation algorithm, such as Figure 4 As shown, each path point contains three-dimensional position coordinates, the desired spray gun end posture and the spray start and stop status flag; initialize the current state of the spray gun; Starting from the first path point, in each discrete servo control cycle, the current state of the spray gun and the current path point information are combined into a state vector. After receiving the input state vector, the policy network outputs an action instruction vector. The target speed and posture instructions in the action instruction vector are converted into specific joint motion instructions. The generated specific joint motion instructions are added to a first-in-first-out instruction sequence buffer. After the robot motion controller executes the instructions of the current cycle, the current state of the spray gun is updated. Calculate the distance between the current gun position and the current path point. If the distance is less than the tolerance and the end point of the trajectory has not been reached, the next path point will be used as the current path point, until the last path point.
[0039] Specifically, during the process of generating an executable spray path instruction sequence, the global guidance trajectory is refined into a series of path points using an interpolation algorithm. For example, assuming the global guidance trajectory is a path from the starting point to the end point, the system will use an interpolation algorithm to refine it into multiple path points, each of which contains three-dimensional position coordinates, the desired spray gun end posture, and the spray start and stop status flag. If the starting coordinates of the global guidance trajectory are (0,0,0) and the end coordinates are (10,10,10), it can be refined into 10 path points, with the coordinates of each point being (1,1,1), (2,2,2),..., (10,10,10), and the corresponding spray gun posture and start and stop status are set.
[0040] The spray gun's current state is initialized, and starting from the first pathpoint, motion commands are generated within each discrete servo control cycle. For example, assuming the coordinates of the current pathpoint are (3,3,3) and the spray gun's current state is (2,2,2), the current state of the spray gun and the current pathpoint information are combined into a state vector and input into the policy network. After receiving the input state vector, the policy network outputs an action command vector, such as a target velocity of (1,1,1) and a posture command of (0,0,0). This action command vector is converted into specific joint motion commands and added to a first-in-first-out instruction sequence buffer. After the robot motion controller completes the execution of the current cycle's commands, it updates the spray gun's current state and calculates the distance between the current spray gun position and the current pathpoint. If the distance is less than the preset tolerance and the trajectory endpoint has not been reached, the next pathpoint is set as the current pathpoint, and motion command generation continues until the last pathpoint.
[0041] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the above embodiments, it should be understood by those skilled in the art that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. In addition, the various different implementations of the embodiments of the present invention can also be arbitrarily combined, as long as they do not violate the ideas of the embodiments of the present invention, and they should also be regarded as the contents disclosed in the embodiments of the present invention.
Claims
1. A water-based adhesive spraying path planning method based on reinforcement learning, characterized in that: include: Obtaining a three-dimensional model of the workpiece and preset spraying process parameters, the policy network interacts with the spraying environment simulation and generates trajectory data containing states, actions, and rewards. The value network then calculates an advantage estimate and the variance of the advantage estimate. When the monitored spraying quality index does not meet the preset performance benchmark, the trajectory that does not meet the quality standard is obtained from the trajectory, a training data set is constructed based on the variance of the advantage estimate, and the error of the training data set is used to update the value network; the advantage estimate is calculated based on the updated value network and the parameters of the strategy network are optimized; The workpiece surface is converted into a discrete graph structure, a landmark set is established, and the cost between each point is pre-calculated; a node heat map of the workpiece surface is obtained using the value network; a landmark subset corresponding to the current spraying area is determined from the landmark set based on the geometric properties of the current spraying area and the spraying accuracy requirements, and the contribution value of the landmarks in the landmark subset in the path search of the current spraying area is obtained using the node heat map; When the actual surface morphology of the workpiece deviates from the three-dimensional model beyond a preset value, the contribution value of the landmark related to the target spraying area is adjusted; the global guidance trajectory is calculated on the graph structure using the landmark subset and the contribution value, and an executable spray path instruction sequence is generated based on the strategy network and the global guidance trajectory.
2. The water-based glue spraying path planning method based on reinforcement learning according to claim 1 is characterized in that: The step of constructing a training data set according to the variance of the advantage estimate comprises: Compute the variance of the advantage estimate for each state-action pair within a time segment in the trajectory; All trajectory data with a variance of the advantage estimate greater than the variance threshold, or trajectory data ranked in the top N percent of the variance, are added to the training data set.
3. The water-based glue spraying path planning method based on reinforcement learning according to claim 1 is characterized in that: The step of obtaining a node heat map of a workpiece surface by using the value network includes: Extract the three-dimensional spatial position coordinates, surface normal vector, and preset spraying process parameters of each node to obtain the state vector; Each state vector is input into the trained value network to obtain the evaluation value; the evaluation values of all nodes are normalized to obtain the normalized node heat map.
4. The water-based glue spraying path planning method based on reinforcement learning according to claim 1 is characterized in that: The step of using the node heat map to obtain the contribution value of the landmarks in the landmark subset in the current spraying area path search includes: From a pre-set global landmark set, all landmarks whose geometric positions are located within the current spraying area and landmarks whose distance from the current spraying area is less than a predetermined threshold are obtained, thereby obtaining a current landmark subset; For each landmark in the landmark subset , determine the set of neighboring nodes N with the nearest neighbor in Euclidean distance; Calculate the average value of each node evaluation value in the neighborhood node set and use the average value as the landmark Contribution value in the path search process .
5. The water-based adhesive spraying path planning method based on reinforcement learning according to claim 1 is characterized in that: The adjusting the contribution value of the landmark associated with the target spraying area includes: Get the upper threshold of positive deviation and the negative deviation lower threshold ; For each landmark in the target spraying area , calculated with The average value Avg of the normal deviation of all actual surface points in the influence area with a radius of R as the center; If Avg is greater than , then according to the formula To reduce the landmark The contribution value of is the contribution value attenuation coefficient; If Avg is less than , then according to the formula To reduce the landmark The contribution value of is the contribution value attenuation coefficient.
6. The water-based glue spraying path planning method based on reinforcement learning according to claim 1 is characterized in that: The calculating a global guidance trajectory on the graph structure using the landmark subset and the contribution value includes: Based on the discretized graph structure G=(V,E) of the workpiece surface, V is a set of nodes representing each tiny area on the surface, and E is a set of edges representing nodes that can be directly connected for spraying operations; each edge e(u,v) connects node u and node v and has a basic path cost value ,The basic path cost value is determined based on the three-dimensional Euclidean distance between nodes u and v or the estimated spraying time; Determine the starting node and ending node of the current spraying task in the graph structure G; For each edge e(u,v) in the graph structure G, get all landmarks whose distance to the midpoint of the edge e(u,v) is less than the influence radius R , the edge adjustment factor is calculated according to the following formula ,in Is a nearby landmark The contribution value of It is a landmark The weight of the influence on edge e(u,v); Calculate the cost of each edge in the graph structure G , where is the basic path cost of edge e(u,v); A path search method is used to search for a path with the minimum cumulative cost from the starting node to the ending node as the global guiding trajectory.
7. The water-based adhesive spraying path planning method based on reinforcement learning according to claim 1 is characterized in that: The method of generating an executable spray path instruction sequence based on the strategy network and the global guidance trajectory includes: The global guidance trajectory is refined into a series of path points through an interpolation algorithm. Each path point contains the three-dimensional position coordinates, the desired spray gun end posture, and the spray start and stop status flag; the current state of the spray gun is initialized; Starting from the first path point, in each discrete servo control cycle, the current state of the spray gun and the current path point information are combined into a state vector. After receiving the input state vector, the policy network outputs an action instruction vector. The target speed and posture instructions in the action instruction vector are converted into specific joint motion instructions. The generated specific joint motion instructions are added to a first-in-first-out instruction sequence buffer. After the robot motion controller executes the instructions of the current cycle, the current state of the spray gun is updated. Calculate the distance between the current gun position and the current path point. If the distance is less than the tolerance and the end point of the trajectory has not been reached, the next path point will be used as the current path point, until the last path point.
8. A water-based glue spraying path planning system based on reinforcement learning, characterized in that: include: A data generation unit is used to obtain a three-dimensional model of the workpiece and preset spraying process parameters. The strategy network interacts with the spraying environment simulation and generates trajectory data including states, actions, and rewards. The value network calculates the advantage estimate and the variance of the advantage estimate. an iterative optimization unit for obtaining, when the monitored spray quality indicator fails to meet a preset performance benchmark, a trajectory that fails to meet the quality benchmark from the trajectory, constructing a training data set based on the variance of the advantage estimate, and updating the value network using the error of the training data set; and optimizing the parameters of the strategy network based on the advantage estimate calculated by the updated value network; A graph construction unit is configured to convert the workpiece surface into a discrete graph structure, establish a landmark set, and pre-calculate the cost between each point; utilize the value network to obtain a node heat map of the workpiece surface; determine a landmark subset corresponding to the current spraying area from the landmark set based on the geometric properties of the current spraying area and the spraying accuracy requirements; and utilize the node heat map to obtain the contribution value of the landmarks in the landmark subset in the path search of the current spraying area; A path generation unit is used to adjust the contribution values of landmarks related to the target spraying area when the actual surface morphology of the workpiece and the three-dimensional model exceed a preset deviation; use the landmark subset and the contribution values to calculate the global guidance trajectory on the graph structure, and generate an executable spray path instruction sequence based on the strategy network and the global guidance trajectory.
9. The water-based adhesive spraying path planning system based on reinforcement learning according to claim 8 is characterized in that: The step of constructing a training data set according to the variance of the advantage estimate comprises: Compute the variance of the advantage estimate for each state-action pair within a time segment in the trajectory; All trajectory data with a variance of the advantage estimate greater than the variance threshold, or trajectory data ranked in the top N percent of the variance, are added to the training data set.
10. The water-based adhesive spraying path planning system based on reinforcement learning according to claim 8, characterized in that: The step of obtaining a node heat map of a workpiece surface by using the value network includes: Extract the three-dimensional spatial position coordinates, surface normal vector, and preset spraying process parameters of each node to obtain the state vector; Each state vector is input into the trained value network to obtain the evaluation value; the evaluation values of all nodes are normalized to obtain the normalized node heat map.
Citation Information
Patent Citations
Intelligent coating track planning method based on deep reinforcement learning
CN115408813A
Unmanned aerial vehicle adaptive information path planning method based on deep reinforcement learning
CN116088579A
Vector propeller control system and method based on flexible shaft
CN117666355A
RRT algorithm spraying mechanical arm path planning method and system based on dynamic step length
CN119704181A
Sleeper labeling method based on deep learning algorithm
CN119888642A
Cited By
Reinforcement learning training method and system for space-time interaction dynamic three-dimensional reconstruction
CN121562717A