A cast cleaning trajectory planning method based on point cloud processing and reinforcement learning

CN121468563BActive Publication Date: 2026-08-21CRRC DALIAN INST CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511954941.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-08-21
Estimated Expiration
2045-12-23

AI Technical Summary

Technical Problem

[0003]本发明提供一种基于点云处理与强化学习的铸件清理轨迹规划方法,以克服智能体难以直接学习铸件点云数据的有效特征表示问题,以及因缺乏将先验的几何知识融入学习框架的有效机制,导致学习效率低下且训练不稳定的问题

Benefits of technology

[0010]有益效果:本发明一种基于点云处理与强化学习的铸件清理轨迹规划方法,通过三维相机获取点云数据和基于深度学习的点云分割方法,铸件清理系统能够自动识别并定位每一个铸件的待清理区域,无需预先编程;即使是不同的铸件,铸件清理系统也能够自动重新规划轨迹,极大提升了生产线的灵活性与适应性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121468563B_ABST
    Figure CN121468563B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on point cloud processing and reinforcement learning's castings cleaning trajectory planning method, comprising: the original point cloud data of castings surface is collected by three-dimensional camera;Original point cloud data is identified using point cloud segmentation method based on deep learning, obtain castings main point set and to be cleaned area point set;Geometric feature extraction is carried out to the point set, and point cloud geometric feature is obtained;Reinforcement learning agent is introduced, in combination with the geometric feature and robot state, reinforcement learning method of policy gradient is used to train the reinforcement learning agent, and the policy network of the reinforcement learning agent after training is deployed to actual robot system, in combination with PD controller, output expected joint position instruction, realize the automation trajectory planning of castings cleaning.The method effectively improves the precision and adaptability of trajectory planning, significantly improves cleaning efficiency and automation level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of casting cleaning and robotic arm trajectory planning, and in particular to a casting cleaning trajectory planning method based on point cloud processing and reinforcement learning. Background Technology

[0002] In the casting production process, the surface of castings often has excess structures such as risers, flash, and burrs. These areas need to be cleaned by cutting, grinding, or polishing in post-processing. Traditional casting cleaning methods mainly fall into two categories: The first is manual cleaning, which relies on workers using grinding wheels or cutting tools. This method is not only labor-intensive and inefficient, but also produces a noisy and dusty working environment, posing serious occupational health and safety risks. The second is automated cleaning based on robot teaching. This method requires engineers to pre-generate cleaning trajectories through teaching programming or offline programming. Although it can replace manual labor to some extent, its disadvantages are: the teaching trajectory is usually specific to a particular casting model, lacking flexibility and generalization ability; when the batch, size, or orientation of the casting changes, it is necessary to re-teach or modify the program, resulting in low efficiency and high maintenance costs. With the development of 3D vision and intelligent control, point cloud processing technology has become an important means of perceiving the surface of castings. Point cloud data acquired through 3D cameras can be used for region segmentation and geometric fitting, thereby automatically extracting areas to be cleaned. On the other hand, reinforcement learning has shown good adaptive capabilities in the fields of robot control and trajectory optimization. By constructing diverse scenarios in a simulation environment to train agents, they can learn to generate efficient trajectories under different constraints. However, directly applying reinforcement learning to such tasks has the following bottlenecks: the high dimensionality and sparsity of casting point cloud data make it difficult for agents to directly learn effective feature representations; the lack of an effective mechanism to integrate prior geometric knowledge into the learning framework leads to low learning efficiency and unstable training. Summary of the Invention

[0003] This invention provides a casting cleaning trajectory planning method based on point cloud processing and reinforcement learning, which overcomes the problem that the agent has difficulty in directly learning the effective feature representation of casting point cloud data, as well as the problem that the lack of an effective mechanism to integrate prior geometric knowledge into the learning framework leads to low learning efficiency and unstable training.

[0004] To achieve the above objectives, the technical solution of the present invention is as follows: A casting cleaning trajectory planning method based on point cloud processing and reinforcement learning includes: S1. Acquire the original point cloud data of the casting surface using a 3D camera and perform preprocessing to obtain preprocessed point cloud data; S2. Based on the deep learning-based point cloud segmentation method and the preprocessed point cloud data, obtain the category probability of the preprocessed point cloud data to determine the category of the points in the preprocessed point cloud data, and obtain the main body point set of the casting and the point set of the area to be cleaned. S3. Extract the geometric features of the main point set of the casting and the point set of the area to be cleaned to obtain the geometric features of the point cloud; the geometric features of the point cloud include the multi-plane polar coordinate vector of the main point set of the casting and the multi-plane polar coordinate vector of the point set of the area to be cleaned. S4. Introduce a reinforcement learning agent. Based on the geometric features of the point cloud and the robot's state information, train the reinforcement learning agent using a policy gradient reinforcement learning method to obtain a trained reinforcement learning agent. The robot's state information includes the joint positions and velocities of the robotic arm and the three-dimensional pose of the end effector. S5. Deploy the policy network of the trained reinforcement learning agent into the actual robot casting cleaning system. Based on the collected original point cloud data of the surface of the casting to be cleaned, obtain the geometric features of the point cloud of the surface of the casting to be cleaned and output the desired joint position. Based on the desired joint position, obtain the desired joint driving torque signal through the PD controller, which is used to drive the robotic arm to move in order to realize automated trajectory planning.

[0005] Furthermore, the deep learning-based point cloud segmentation method is either a point-based method or a Voxel-based method, wherein the objective function of the deep learning-based point cloud segmentation method or the Voxel-based method is a learning function, expressed as: (1) In the formula, Indicates the main area of ​​the casting; Indicates the area to be cleaned; Represents the set of learnable parameters of a network model; Indicates by parameters The defined point cloud segmentation model; This represents the set of point clouds that were input.

[0006] Furthermore, the specific steps for obtaining the geometric features of the point cloud include: S31. Based on the segmented point cloud data, obtain the point set of the main casting body and the point set of the area to be cleaned, expressed as: (2) (3) In the formula, Represents the main point set of the casting; This represents the set of points in the area to be cleared. Indicates the first One point; Represents three-dimensional Euclidean space; Point Type; S32. Perform principal component analysis on the preprocessed point cloud data to obtain three mutually orthogonal principal directions, and then construct three projection planes; S33. Project the main point set of the casting and the point set of the area to be cleaned onto three projection planes to obtain the following projected point set: (4) In the formula, Indicates the index of the selected projection plane; and These represent the values ​​of the point cloud on the two coordinate axes within the projection plane; The point cloud of the main body of the casting is shown in the first position. A set of projection points on a projection plane; This indicates the point cloud area to be cleaned in the [number]th [year]. A set of projection points on a projection plane; S34. Based on the set of projection points, calculate the polar coordinate radius function on the projection plane, expressed as: (5) (6) In the formula, The centroid coordinates of the projection point set of the main body of the casting; Represents the centroid coordinates of the projection point set of the area to be cleaned; Indicates the polar angle within the projection plane; Indicates the first Along the polar angle direction on each projection plane The maximum radial distance function from the boundary of the casting body to the centroid of the casting body point set is the polar coordinate radius function of the casting body point set. Indicates the first Along the polar angle direction on each projection plane The maximum radial distance function from the boundary of the region to be cleared to the centroid of the point set of the region to be cleared is the polar coordinate radius function of the point set of the region to be cleared. The polar angle within the projection plane is sampled by using intervals... Evenly divided into The angle is expressed as: (7) In the formula, Represents the set of polar angles used for sampling; Indicates the first The polar angle of each sampling point; Indicates the index of the angle sampling point; Indicates the total number of samples; S35. Based on the polar coordinate radius function, the discrete radius sequence of the projection plane is obtained, and its expression is: (8) (9) In the formula, Indicates the main body point set of the casting at the th Polar angle of each angle sampling point on the projection plane The calculated sequence of discrete radius values; Indicates the point set of the area to be cleaned at the th Polar angle of each angle sampling point on the projection plane The calculated sequence of discrete radius values; S36. Concatenate the discrete radius sequences of each projection plane to obtain the geometric features of the point cloud, expressed as: (10) (11) In the formula, This represents the multi-plane polar coordinate feature vector extracted and stitched together from the three orthogonal projection planes of the casting body; This represents the multi-plane polar coordinate feature vector obtained by extracting and stitching the region to be cleaned on three projection planes.

[0007] Furthermore, the state space of the reinforcement learning agent is: (12) In the formula, Represents the state vector; Indicates the current joint position; Indicates the current joint velocity; Represents the three-dimensional pose of the end effector; The action space of the reinforcement learning agent is: (13) In the formula, For the joint degrees of freedom of the robotic arm; The target joint position for the nth degree of freedom; The reward function for the reinforcement learning agent is designed as follows: (14) In the formula, Indicates coverage bonus; Indicates a collision penalty; Indicates a penalty for trajectory smoothness; Indicates energy consumption constraints; The expression for the coverage reward is as follows: (15) In the formula, This represents the coverage bonus coefficient; The set of coverage points that overlaps with the end-effector's area of ​​action and the area to be cleaned; the end-effector's area of ​​action is a neighborhood sphere centered on the trajectory points of the casting cleaning trajectory; The expression for the collision penalty is: (16) In the formula, Indicates the penalty coefficient; This represents the set of coverage points that overlap with the end effect area and the main casting area; The expression for the smoothness penalty is: (17) In the formula, Indicates the smoothness coefficient; This indicates the agent's current time step. The amount of change in joint position; This represents the change in joint position of the agent in the previous time step; The expression for the energy consumption constraint is: (18) In the formula, Indicates energy consumption weight; This indicates the current joint driving torque.

[0008] Furthermore, the objective function of the policy gradient-based reinforcement learning method is: (19) In the formula, This represents the optimization objective function based on the pruning strategy; This represents the expected operation on the upsampled data at each time step; This indicates a function clipping operation; Indicates the clipping parameters; This represents the estimation of the advantage function; Represents the learnable parameters of the policy network; This represents the probability ratio, and ,in, Indicates an action, Indicates state, Indicates the current strategy parameters Under these circumstances, the agent is in a state Choose action The probability of; This indicates the parameters before the policy update. Under these conditions, the agents are in the same state. Choose action The probability of; The advantage function estimate is calculated using generalized advantage estimation, and its expression is as follows: (20) In the formula, Represents relative to the current time step Future step size index; Indicates the smoothing parameter; This represents the value function estimate given by the value network; The reward function represents the reward for a future time step.

[0009] Furthermore, based on the desired joint position, the desired joint driving torque signal is obtained through the PD controller, expressed as: (twenty one) In the formula, This represents the joint drive torque signal output by the controller; and Indicates the proportional and differential gain coefficients; Indicates the position of the target joint; Indicates the target joint velocity.

[0010] Beneficial effects: This invention provides a casting cleaning trajectory planning method based on point cloud processing and reinforcement learning. By acquiring point cloud data through a 3D camera and using a point cloud segmentation method based on deep learning, the casting cleaning system can automatically identify and locate the area to be cleaned for each casting without pre-programming. Even for different castings, the casting cleaning system can automatically replan the trajectory, greatly improving the flexibility and adaptability of the production line. By extracting the compact geometric feature of the multi-plane polar coordinate vectors of the casting body and the area to be cleaned, the complex 3D shape information is encoded into low-dimensional, regular vectors, which greatly reduces the dimensionality of the state space and provides a more discriminative input signal for the reinforcement learning agent. This enables the agent to learn cleaning strategies more efficiently. By training the agent using a policy gradient reinforcement learning method, the agent is able to self-optimize through trial and error in a simulation environment, thus receiving sufficient training. Finally, the agent's policy network is deployed into a real casting cleaning system, effectively overcoming the "simulation-reality" gap and making the application of this technical solution in actual industrial scenarios highly feasible. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart of the casting cleaning trajectory planning method of the present invention; Figure 2 This is a schematic diagram of reinforcement learning agent training in an embodiment of the present invention. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0014] This embodiment provides a casting cleaning trajectory planning method based on point cloud processing and reinforcement learning, such as... Figure 1 As shown, it includes: S1. Acquire the original point cloud data of the casting surface using a 3D camera and perform preprocessing to obtain preprocessed point cloud data; S2. Based on the deep learning-based point cloud segmentation method and the preprocessed point cloud data, obtain the category probability of the preprocessed point cloud data to determine the category of the points in the preprocessed point cloud data, and obtain the main body point set of the casting and the point set of the area to be cleaned. S3. Extract the geometric features of the main point set of the casting and the point set of the area to be cleaned to obtain the geometric features of the point cloud; the geometric features of the point cloud include the multi-plane polar coordinate vector of the main point set of the casting and the multi-plane polar coordinate vector of the point set of the area to be cleaned. S4. Introduce a reinforcement learning agent. Based on the geometric features of the point cloud and the robot's state information, train the reinforcement learning agent using a policy gradient reinforcement learning method to obtain a trained reinforcement learning agent. The robot's state information includes the joint positions and velocities of the robotic arm and the three-dimensional pose of the end effector. S5. Deploy the policy network of the trained reinforcement learning agent into the actual robot casting cleaning system. Based on the collected original point cloud data of the surface of the casting to be cleaned, obtain the geometric features of the point cloud of the surface of the casting to be cleaned and output the desired joint position. Based on the desired joint position, obtain the desired joint driving torque signal through the PD controller, which is used to drive the robotic arm to move in order to realize automated trajectory planning.

[0015] Specifically, the original point cloud data is represented as follows: (1) In the formula, Represents the point cloud set on the surface of the casting; Represents the three-dimensional spatial coordinates of a point; This represents the index of a discrete sampling point in a point cloud set; Indicates the number of point clouds; The preprocessing is used to remove isolated noise points with insufficient neighborhood points in the point cloud set by using radius filtering or statistical filtering methods, thereby obtaining an effective point set and improving the stability of subsequent contour extraction.

[0016] Preferably, the deep learning-based point cloud segmentation method is a point-based method (PointNet++) or a Voxel-based method (3D-CNN). In this embodiment, the point-based method (PointNet++) is used for segmentation, and its objective function is a learning function, expressed as: (2) In the formula, Indicates the main area of ​​the casting; Indicates the area to be cleaned; It represents the set of learnable parameters of the network model, including the weights and biases of each layer of the neural network. It is continuously updated through the backpropagation algorithm during the training process to minimize the segmentation error function and improve the point cloud recognition accuracy. Indicates by parameters The defined point cloud segmentation model is used to segment the input point cloud set. The mapping is to a corresponding set of regional category labels, and learning enables automatic identification and classification of different spatial regions. This represents the set of point clouds that were input.

[0017] Specifically, PointNet++ is used to perform layer-by-layer feature fusion on the preprocessed point cloud data to obtain point features, and a Softmax classifier is used to output the class probabilities of the point cloud; the expression for the class probabilities of the point cloud is: (3) In the formula, This represents the probability distribution of a point belonging to a category (the main casting body and the area to be cleaned). This represents the weight matrix of the Softmax classification layer; Representing point features; This represents the bias vector of the Softmax classification layer.

[0018] Preferably, the specific steps for obtaining the geometric features of a point cloud include: S31. Based on the segmented point cloud data, obtain the point set of the main casting body and the point set of the area to be cleaned, expressed as: (4) (5) In the formula, Represents the main point set of the casting; This represents the set of points in the area to be cleared. Indicates the first One point; Represents three-dimensional Euclidean space; Point Type; S32. Perform principal component analysis on the preprocessed point cloud data to obtain three mutually orthogonal principal directions, and then construct three projection planes; S33. Project the main point set of the casting and the point set of the area to be cleaned onto three projection planes to obtain the following projected point set: (6) In the formula, Indicates the index of the selected projection plane; and These represent the values ​​of the point cloud on the two coordinate axes within the projection plane; The point cloud of the main body of the casting is shown in the first position. A set of projection points on a projection plane; This indicates the point cloud area to be cleaned in the [number]th [year]. A set of projection points on a projection plane; S34. Based on the set of projection points, calculate the polar coordinate radius function on the projection plane, expressed as: (7) (8) In the formula, The centroid coordinates of the projection point set of the main body of the casting; Represents the centroid coordinates of the projection point set of the area to be cleaned; Indicates the polar angle within the projection plane; Indicates the first Along the polar angle direction on each projection plane The maximum radial distance function from the boundary of the casting body to the centroid of the casting body point set is the polar coordinate radius function of the casting body point set. Indicates the first Along the polar angle direction on each projection plane The maximum radial distance function from the boundary of the region to be cleared to the centroid of the point set of the region to be cleared is the polar coordinate radius function of the point set of the region to be cleared. The polar angle within the projection plane is sampled by sampling the interval... Evenly divided into The angle is expressed as: (9) In the formula, Represents the set of polar angles used for sampling; Indicates the first The polar angle of each sampling point; Indicates the index of the angle sampling point; Indicates the total number of samples; S35. Based on the polar coordinate radius function, the discrete radius sequence of the projection plane is obtained, and its expression is: (10) (11) In the formula, Indicates the main body point set of the casting at the th Polar angle of each angle sampling point on the projection plane The calculated sequence of discrete radius values; Indicates the point set of the area to be cleaned at the th Polar angle of each angle sampling point on the projection plane The calculated sequence of discrete radius values; S36. Concatenate the discrete radius sequences of each projection plane to obtain the geometric features of the point cloud, expressed as: (12) (13) In the formula, This represents the multi-plane polar coordinate feature vector extracted and stitched together from the three orthogonal projection planes of the casting body; This represents the multi-plane polar coordinate feature vector obtained by extracting and stitching the region to be cleaned on three projection planes.

[0019] Preferably, the state space of the reinforcement learning agent is: (14) In the formula, Represents the state vector; Indicates the current joint position; Indicates the current joint velocity; Represents the three-dimensional pose of the end effector; The action space of the reinforcement learning agent is: (15) In the formula, For the joint degrees of freedom of the robotic arm; The target joint position for the nth degree of freedom; Specifically, the reinforcement learning agent outputs the target joint position at a higher level through the action space, while the stability and execution effect of the underlying dynamics are guaranteed by the PD controller. The parameters of the PD controller can be independently adjusted according to different working conditions, thereby enhancing the robustness and adaptability of the action output. The reward function for the reinforcement learning agent is designed as follows: (16) In the formula, Indicates coverage bonus; Indicates a collision penalty; Indicates a penalty for trajectory smoothness; Indicates energy consumption constraints; The expression for the coverage reward is as follows: (17) In the formula, This represents the coverage bonus coefficient; The set of coverage points that overlaps with the end-effector's area of ​​action and the area to be cleaned; the end-effector's area of ​​action is a neighborhood sphere centered on the trajectory points of the casting cleaning trajectory; The expression for the collision penalty is: (18) In the formula, Indicates the penalty coefficient; This represents the set of coverage points where the end effect area coincides with the main casting area; The expression for the smoothness penalty is: (19) In the formula, Indicates the smoothness coefficient; This indicates the agent's current time step. The amount of change in joint position; This represents the change in joint position of the agent in the previous time step; The expression for the energy consumption constraint is: (20) In the formula, Indicates energy consumption weight; This indicates the current joint driving torque.

[0020] Specifically, in this embodiment, the reinforcement learning method for policy gradients adopts the PPO (Proximal Policy Optimization) algorithm. Training the reinforcement learning agent through the policy gradient reinforcement learning method is an existing technology, and will only be briefly described here. like Figure 2 As shown, the training steps of the reinforcement learning agent include: S41. Construct a simulation environment and initialize the policy network and value network of the reinforcement learning agent; S42. The strategy network performs actions and outputs them to the simulation environment based on the current state, namely the geometric features of the point cloud and the robot's state information. S43. The policy network obtains the state and reward at the next moment through feedback from the simulation environment; and calculates the cumulative discount reward based on the reward. S44. The value network designs its objective function based on the current state, action, next moment state, and reward using a reinforcement learning method with policy gradient. S45. Update the policy network and value network based on the objective function and cumulative discount reward of the value network; S46. Repeat steps S42 to S45 until performance converges and training is complete.

[0021] Specifically, the cumulative discount return is calculated based on the reward function, and the current strategy is evaluated based on the cumulative discount return. The expression for calculating the cumulative discount return is as follows: (twenty one) In the formula, Indicates the time step from the current time. The cumulative discounted reward obtained by the agent at each future time step; T represents the total number of time steps for a complete training trajectory; Indicates relative to the current time step Future step size index; This represents the discount factor, and ; This indicates that the agent is at time step The instant reward received.

[0022] Preferably, the objective function of the policy gradient-based reinforcement learning method is: (twenty two) In the formula, This represents the optimization objective function based on the pruning strategy; This represents the expected operation on the upsampled data at each time step; This indicates a function clipping operation; Indicates the clipping parameters; This represents the estimation of the advantage function; Represents the learnable parameters of the policy network; Represents the probability ratio, and ,in, Indicates an action, Indicates state, Indicates the current strategy parameters Under these circumstances, the agent is in a state Choose action The probability of; This indicates the parameters before the policy update. Under these conditions, the agents are in the same state. Choose action The probability of; The advantage function estimate is calculated using generalized advantage estimation, and its expression is as follows: (twenty three) In the formula, Indicates the smoothing parameter; This represents the value function estimate given by the value network; The reward function represents the reward for a future time step.

[0023] Specifically, the learnable parameters are updated through gradient ascent, expressed as: (twenty four) In the formula, Indicates the learning rate; express The gradient.

[0024] Specifically, the trained reinforcement learning agent is deployed into the actual robot's casting cleaning system. During the deployment phase, the reinforcement learning agent takes the geometric features of the point cloud (converted from the point cloud data collected by the sensor) and the robot's state information as inputs and outputs the desired joint position command. Through the underlying PD controller, the desired joint position command is converted into joint driving torque, thereby driving the robotic arm to perform the cleaning action.

[0025] Preferably, based on the desired joint position, the desired joint driving torque signal is obtained through the PD controller, and the expression is: (25) In the formula, This represents the joint drive torque signal output by the controller; and Indicates the proportional and differential gain coefficients; Indicates the position of the target joint; Indicates the target joint velocity.

[0026] The present invention has the following beneficial effects: The present invention provides a casting cleaning trajectory planning method based on point cloud processing and reinforcement learning. By acquiring point cloud data through a 3D camera and using a point cloud segmentation method based on deep learning, the casting cleaning system can automatically identify and locate the area to be cleaned for each casting without pre-programming; even for different castings, the casting cleaning system can automatically replan the trajectory, greatly improving the flexibility and adaptability of the production line. By extracting the compact geometric feature of the multi-plane polar coordinate vectors of the casting body and the area to be cleaned, the complex 3D shape information is encoded into low-dimensional, regular vectors, which greatly reduces the dimensionality of the state space and provides a more discriminative input signal for the reinforcement learning agent. This enables the agent to learn cleaning strategies more efficiently. By training the agent using a policy gradient reinforcement learning method, the agent is able to self-optimize through trial and error in a simulation environment, thus receiving sufficient training. Finally, the agent's policy network is deployed into a real casting cleaning system, effectively overcoming the "simulation-reality" gap and making the application of this technical solution in actual industrial scenarios highly feasible.

[0027] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for planning casting cleaning trajectories based on point cloud processing and reinforcement learning, characterized in that, include: S1. Acquire the original point cloud data of the casting surface using a 3D camera and perform preprocessing to obtain preprocessed point cloud data; S2. Based on the deep learning-based point cloud segmentation method and the preprocessed point cloud data, obtain the category probability of the preprocessed point cloud data to determine the category of the points in the preprocessed point cloud data, and obtain the main body point set of the casting and the point set of the area to be cleaned. S3. Extract the geometric features of the main point set of the casting and the point set of the area to be cleaned to obtain the geometric features of the point cloud; the geometric features of the point cloud include the multi-plane polar coordinate vector of the main point set of the casting and the multi-plane polar coordinate vector of the point set of the area to be cleaned. S4. Introduce a reinforcement learning agent. Based on the geometric features of the point cloud and the robot's state information, train the reinforcement learning agent using a policy gradient reinforcement learning method to obtain a trained reinforcement learning agent. The robot's state information includes the joint positions and velocities of the robotic arm and the three-dimensional pose of the end effector. The state space of the reinforcement learning agent is as follows: (1) In the formula, Represents the state vector; Indicates the current joint position; Indicates the current joint velocity; Represents the three-dimensional pose of the end effector; This represents the multi-plane polar coordinate feature vector extracted and stitched together from the three orthogonal projection planes of the casting body; This represents the multi-plane polar coordinate feature vector obtained by extracting and stitching the region to be cleaned on three projection planes; The action space of the reinforcement learning agent is: (2) In the formula, For the joint degrees of freedom of the robotic arm; The target joint position for the nth degree of freedom; The reward function for the reinforcement learning agent is designed as follows: (3) In the formula, Indicates coverage bonus; Indicates a collision penalty; Indicates a penalty for trajectory smoothness; Indicates energy consumption constraints; The expression for the coverage reward is as follows: (4) In the formula, This represents the coverage bonus coefficient; The set of coverage points that overlaps with the end-effector's area of ​​action and the area to be cleaned; the end-effector's area of ​​action is a neighborhood sphere centered on the trajectory points of the casting cleaning trajectory; This represents the set of points in the area to be cleared. The expression for the collision penalty is: (5) In the formula, Indicates the penalty coefficient; This represents the set of coverage points that overlap with the end effect area and the main casting area; Represents the main point set of the casting; The expression for the smoothness penalty is: (6) In the formula, Indicates the smoothness coefficient; This indicates the agent's current time step. The amount of change in joint position; This represents the change in joint position of the agent in the previous time step; The expression for the energy consumption constraint is: (7) In the formula, Indicates energy consumption weight; This indicates the current joint driving torque; The objective function of the policy gradient-based reinforcement learning method is: (8) In the formula, This represents the optimization objective function based on the pruning strategy; This represents the expected operation on the upsampled data at each time step; This indicates a function clipping operation; Indicates the clipping parameters; This represents the estimation of the advantage function; Represents the learnable parameters of the policy network; This represents the probability ratio, and ,in, Indicates an action, Indicates state, Indicates the current strategy parameters Under these circumstances, the agent is in a state Choose action The probability of; This indicates the parameters before the policy update. Under these conditions, the agents are in the same state. Choose action The probability of; The advantage function estimate is calculated using generalized advantage estimation, and its expression is as follows: (9) In the formula, Represents relative to the current time step Future step size index; Indicates the smoothing parameter; This represents the value function estimate given by the value network; The reward function representing a future time step; S5. Deploy the policy network of the trained reinforcement learning agent into the actual robot casting cleaning system. Based on the collected original point cloud data of the surface of the casting to be cleaned, obtain the geometric features of the point cloud of the surface of the casting to be cleaned and output the desired joint position. Based on the desired joint position, obtain the desired joint driving torque signal through the PD controller, which is used to drive the robotic arm to move in order to realize automated trajectory planning.

2. The casting cleaning trajectory planning method based on point cloud processing and reinforcement learning according to claim 1, characterized in that, The deep learning-based point cloud segmentation method is either a point-based method or a voxel-based method, wherein the objective function of the deep learning-based point cloud segmentation method is a learning function, expressed as: (10) In the formula, Indicates the main area of ​​the casting; Indicates the area to be cleaned; Represents the set of learnable parameters of a network model; Indicates by parameters The defined point cloud segmentation model; This represents the set of point clouds that were input.

3. The casting cleaning trajectory planning method based on point cloud processing and reinforcement learning according to claim 2, characterized in that, The specific steps to obtain the geometric features of a point cloud include: S31. Based on the segmented point cloud data, obtain the point set of the main casting body and the point set of the area to be cleaned, expressed as: (11) (12) In the formula, Indicates the first One point; Represents three-dimensional Euclidean space; Point Type; S32. Perform principal component analysis on the preprocessed point cloud data to obtain three mutually orthogonal principal directions, and then construct three projection planes; S33. Project the main point set of the casting and the point set of the area to be cleaned onto three projection planes to obtain the following projected point set: (13) In the formula, Indicates the index of the selected projection plane; and These represent the values ​​of the point cloud on the two coordinate axes within the projection plane; The point cloud of the main body of the casting is shown in the first position. A set of projection points on a projection plane; This indicates the point cloud area to be cleaned in the [number]th [year]. A set of projection points on a projection plane; S34. Based on the set of projection points, calculate the polar coordinate radius function on the projection plane, expressed as: (14) (15) In the formula, The centroid coordinates of the projection point set of the main body of the casting; Represents the centroid coordinates of the projection point set of the area to be cleaned; Indicates the polar angle within the projection plane; Indicates the first Along the polar angle direction on each projection plane The maximum radial distance function from the boundary of the casting body to the centroid of the casting body point set is the polar coordinate radius function of the casting body point set. Indicates the first Along the polar angle direction on each projection plane The maximum radial distance function from the boundary of the region to be cleared to the centroid of the point set of the region to be cleared is the polar coordinate radius function of the point set of the region to be cleared. The polar angle within the projection plane is sampled by sampling the interval... Evenly divided into The angle is expressed as: (16) In the formula, Represents the set of polar angles used for sampling; Indicates the first The polar angle of each sampling point; Indicates the index of the angle sampling point; Indicates the total number of samples; S35. Based on the polar coordinate radius function, the discrete radius sequence of the projection plane is obtained, and its expression is: (17) (18) In the formula, Indicates the main body point set of the casting at the th Polar angle of each angle sampling point on the projection plane The calculated sequence of discrete radius values; Indicates the point set of the area to be cleaned at the th Polar angle of each angle sampling point on the projection plane The calculated sequence of discrete radius values; S36. Concatenate the discrete radius sequences of each projection plane to obtain the geometric features of the point cloud, expressed as: (19) (20)。 4. The casting cleaning trajectory planning method based on point cloud processing and reinforcement learning according to claim 1, characterized in that, Based on the desired joint position, the desired joint driving torque signal is obtained through the PD controller, and its expression is: (21) In the formula, This represents the joint drive torque signal output by the controller; and Indicates the proportional and differential gain coefficients; Indicates the position of the target joint; Indicates the target joint velocity.

Citation Information

Patent Citations

  • Intelligent tool scheduling control system for casting cleaning robot

    CN117260728A

  • Low-temperature cooling water cooling air-cooled island cooling robot path planning method and system

    CN120991884A