A contact machining robot motion planning method based on deep reinforcement learning

Through the method of deep reinforcement learning, the problems of path continuity and smoothness in contact processing of multi-joint robots are solved, intelligent and safe planning of contact processing is realized, and the processing quality and efficiency are improved. It is suitable for processes such as grinding, polishing, and welding.

CN119501934BActive Publication Date: 2025-10-10BEIHANG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411662463.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-20
Publication Date
2025-10-10
Estimated Expiration
2044-11-20

AI Technical Summary

Technical Problem

Multi-joint industrial robots suffer from unexpected rigid contact effects during contact processing, resulting in poor processing quality and low efficiency. Existing methods fail to effectively solve the problems of robot path continuity and smoothness, and real-time planning is insufficient.

Method used

A method based on deep reinforcement learning is adopted. By establishing the forward and inverse kinematic models of the robotic arm, designing a multi-constraint reward function, and constructing a double-delayed deep deterministic policy gradient network, real-time optimization and safety planning of contact machining parameters are achieved.

Benefits of technology

It realizes intelligent safety planning of contact processing, avoids self-excited vibration, improves processing quality and efficiency, and is suitable for grinding, polishing, welding and other contact processing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119501934B_ABST
    Figure CN119501934B_ABST
Patent Text Reader

Abstract

The application relates to a contact type machining robot motion planning method based on deep reinforcement learning, which comprises the following steps: firstly, based on DH parameters and homogeneous transformation matrix, a position level forward and inverse kinematics model between a six-degree-of-freedom mechanical arm end state vector and a base and joint motion vector is established; secondly, considering constraints such as a machining task, a feeding speed, a feeding angle and the like, a multi-stage reward function is designed for operation parameter planning; then, a network structure is designed based on a double-delay deep deterministic policy gradient, and three-level optimization is carried out on the algorithm; finally, according to the network structure, a planning process is designed, and the application is applied to a planning task of an intelligent contact type machining mechanical arm system taking grinding and polishing machining as an example. The application can realize real-time optimization and control of various machining parameters of the contact type machining process through autonomous learning, and can be used for intelligent and safe planning under the condition of contact type machining process dynamics coupling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of motion planning for intelligent contact-type processing robots, and specifically relates to a motion planning method for contact-type processing robots based on deep reinforcement learning. Background Art

[0002] Multi-joint industrial robots operate under multiple cross-coupling constraints. Furthermore, contact-based machining tasks such as grinding, polishing, and hole-making can be subject to unexpected rigid contact effects. Improper machining parameter settings can easily induce spatiotemporal self-excited tangential chatter, resulting in poor workpiece quality and low machining efficiency. Therefore, it is imperative to develop a real-time safety planning method and parameter optimization technology to track and plan the entire machining process.

[0003] In the motion planning of contact machining robots, traditional methods typically plan tool path points. Three main approaches exist: the equal chord length error method, the equal chord height error method, and the equal parameter method. However, these traditional methods solely consider surface curvature changes and normal direction as criteria. While they aim to optimize paths and improve machining efficiency from an algorithmic perspective, they fail to consider the continuity and smoothness of the robot path during actual machining. With the advancement of machine learning, a Chinese patent application (Application No. CN202410322517.6) proposes a learning and planning method for the labeled 2D data after converting workpiece data into a 2D model. However, this method ignores the issue of trajectory continuity during machining and cannot perform real-time planning. For welding robots, a Chinese patent application (Application No. CN202410667799.3) proposes a data processing and planning method based on deep learning. However, this method requires multiple model calls, and the time it takes to call the model can lead to significant variability in planning results. Regarding the tool planning problem, a Chinese patent application (application number CN202410291785.6) proposed a method of using a deep reinforcement learning algorithm to compensate for machining errors. Although this method effectively reduces machining errors, its pre-planning of the tool relies on the machining model and is not accurate enough. Summary of the Invention

[0004] To address the unstructured dynamic characteristics of the robot's operating environment in different contact machining tasks, and the impact of operating conditions on factors such as process parameters and robot parameters, this paper proposes a motion planning method for contact machining robots based on deep reinforcement learning. This method eliminates the need for precise modeling of the machining process and enables real-time optimization and control of multiple machining parameters in the contact machining process through autonomous learning. This method can be used for intelligent and safe planning in the context of dynamic coupling between contact machining robots and robots.

[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] A motion planning method for a contact processing robot based on deep reinforcement learning, the method comprising the following steps:

[0007] Firstly, based on the DH parameters and homogeneous transformation matrix, the position-level forward and inverse kinematic models between the state variables of the six-degree-of-freedom robot end and the base and joint motion variables were established; secondly, considering the constraints such as processing tasks, feed speed, and feed angle, a multi-level reward function based on multiple constraints was designed to plan the operation parameters; then, based on the double-delayed deep deterministic policy gradient, the network structure was designed, and a triple-improved double-delayed deep deterministic policy gradient network was established; finally, the triple-improved double-delayed deep deterministic policy gradient network was trained and loaded to plan the robot's motion trajectory and optimize the intelligent safety parameters of contact processing. Taking grinding and polishing as an example, it was applied to the planning task of the intelligent contact processing robot arm system.

[0008] The first step is specifically implemented as follows:

[0009] Establish the forward kinematics model of the manipulator, and transform the i-1th coordinate system of the manipulator link to the i-th coordinate system according to the modified DH parameter representation The general formula is:

[0010] ,

[0011] In DH parameter representation, represents the angle between the joint of the i-th coordinate system and the two adjacent equivalent straight rods, represents the length of the equivalent straight rod of the connecting rod in the i-th coordinate system, represents the angle between two adjacent joint axes of the link in the i-th coordinate system, represents the distance between two adjacent equivalent straight rods connected to the joint in the i-th coordinate system;

[0012] Homogeneous change matrix of the tool coordinate system at the end of the robot arm relative to the inertial coordinate system for:

[0013] ,

[0014] Where, is the homogeneous transformation matrix from the inertial coordinate system to the base mass center coordinate system, is the homogeneous transformation matrix from the base mass center coordinate system to the robot arm installation coordinate system, is the homogeneous transformation matrix from the end joint coordinate system to the robot arm installation coordinate system, is the homogeneous transformation matrix from the end joint coordinate system to the tool coordinate system at the end of the robot arm;

[0015] Establish an inverse kinematics model for the manipulator. Classify the joints of the manipulator with an SRS configuration (three ball joints and three revolute joints) into a non-coplanar joint group and a coplanar parallel axis joint group. Solve the non-coplanar joint group by multiplying the inverse matrix on the left and making the corresponding elements equal. Then, substitute the solution of the non-coplanar joint group into the coplanar parallel axis joint group according to the properties of trigonometric functions.

[0016] The following assumptions are made for the grinding and polishing process:

[0017] (1) The contact area is small compared to the disc and workpiece geometry;

[0018] (2) The workpiece surface is smooth and the surface equation is continuous to the second order;

[0019] (3) The polishing disc is much softer than the workpiece, so there is no deformation of the workpiece;

[0020] (4) The deformation of the disk is caused only by the normal contact force, without considering the tangential contact force;

[0021] (5) Compared with the workpiece curvature radius, the material removal depth during polishing is negligible.

[0022] Establish a machining contact surface model, and equate the geometric structure of the flexible grinding disc of the machining tool to a disc. Assume that the disc with a radius of R contacts the workpiece at the initial point O. Establish a standard orthogonal coordinate system O-XYZ with O as the origin. The Z axis is the direction of the surface normal, the X axis is the direction of the tool feed, Od is the center of the disc bottom surface, and the Y axis is given by the right-hand rule. The contact depth of the disc is Defined as the displacement of the disk along the surface normal, the disk tilt angle Defined as the angle between the bottom surface of the disk and the tangent plane of the workpiece, the flexible disk is tilted at an angle Press into the workpiece to the contact depth After that, the bottom of the disk produces a local deformation at the initial point O, and the plane expressed by the bottom of the disk Written as follows:

[0023] ,

[0024] x and y represent the two-dimensional coordinates in the plane.

[0025] According to the classical differential geometry theory, the smooth workpiece contour near the contact point O is approximated as a quadratic function:

[0026] ,

[0027] In the formula represents the principal curvature of the contact surface, represents the minor curvature of the contact surface, represents the angle between the feed direction of the flexible grinding disc and the principal curvature of the contact surface, and T represents the transpose of the matrix.

[0028] The second step is specifically implemented as follows:

[0029] (1) Constructing a reward function based on processing task constraints :

[0030] ,

[0031] Where d represents the distance between the end effector and the corresponding processing point, k1 and k2 are the first and second proportional coefficients;

[0032] (2) Constructing a reward function based on feed rate constraints ;

[0033] The magnitude and direction of the feed speed are important factors affecting the grinding and polishing effect. When the speed of the flexible grinding disc at the end is too high, a collision may occur, causing damage to the workpiece and the driving robot arm. To meet this constraint, the speed constraint is ignored when the end is far from the target position. When the distance between the end and the target position is within a safe range, a penalty function for speed and distance is established:

[0034] ,

[0035] Where v e represents the feed speed of the flexible grinding disc, k3 represents the proportional coefficient of the degree to which the speed constraint is affected by the distance, and k4 represents the proportional coefficient that affects the speed constraint reward value.

[0036] (3) Constructing a reward function based on feed angle constraints ;

[0037] The feed angle constraint is essentially a constraint on the end-arm posture. In the contact surface modeling, the angle between the grinding wheel feed direction and the principal curvature of the contact surface is defined. , Too large a value will generate a large tangential force and cause damage to the tool. The designed reward function is as follows:

[0038] ,

[0039] Where k5 represents the proportional coefficient of the degree to which the angle constraint reward value is affected by the distance, and k6 represents the proportional coefficient that affects the angle constraint reward value.

[0040] Taking the triple constraints into consideration, the comprehensive reward function is proposed as follows:

[0041] ,

[0042] The third step is specifically implemented as follows:

[0043] A dual actor-critic framework is used to fit the value and policy functions. An experience replay buffer is introduced to enable batch checking during training. Three improvements are made to the network to improve the algorithm's efficiency.

[0044] First, simplify the two Q networks. In the Actor-Critic network, the update strategy is:

[0045] ,

[0046] Where r represents the cumulative reward value, γ is the discount factor, is the policy function.

[0047] In actual calculations, when the Actor-Critic network updates slowly, the two networks are similar and lack the independence conditions to make independent estimates. Therefore, a biased estimation function is given.

[0048] ,

[0049] 、 Represent the biased estimation functions of the two Actor-Critic networks respectively; 、 Represent the policy functions of the two Actor-Critic networks respectively;

[0050] Select the smaller value between the two as the final estimator .

[0051] Second, a delayed update strategy is proposed. During the learning process, the target network serves as a function approximator, providing a stable target for comparison. This accelerates network convergence and significantly improves stability. To minimize error propagation, the actor network is updated less frequently than the critic network. Lowering the policy update frequency reduces the variance of the value function updates, leading to higher-quality policies. By sufficiently delaying policy updates, the possibility of repeated updates to an unchanged critic is limited.

[0052] The third step is to smooth the target policy so that the actions have a normalized range. A regularization strategy is introduced to map similar values ​​to similar action spaces, normalizing the values ​​to a region in the action space. This is achieved by adding a small amount of normally distributed noise to the actions and averaging them over a pool of mini-batches of experience.

[0053] The fourth step is specifically implemented as follows:

[0054] The process of designing the algorithm is as follows:

[0055] 1) First initialize the Critic network With Actor Network , set the state vector and assign a random initial value, the state vector includes the joint angle and angular velocity of the robot arm, the speed and posture of the robot arm end effector, and the distance between the end effector and the workpiece;

[0056] 2) Initialize the target critic network With Actor Network , initialize the experience buffer pool R;

[0057] 3) Randomly give a target planning position and obtain the current state s of the robot arm;

[0058] 4) Select the current action after adding random noise according to the current state s , Nt is random noise;

[0059] 5) Execute the current action And obtain the reward value r and the new state s' according to the designed reward function, and Stored in the experience buffer;

[0060] 6) Randomly sample N groups of samples in the experience buffer to form a minimum batch set;

[0061] 7) Normalize the actions and streamline the dual Q network using the optimization method described in step 3;

[0062] 8) Update the critic network according to the double-delay update strategy and loss function described in step 3;

[0063] 9) Update the Actor network based on the gradient update function;

[0064] 10) Update the target network;

[0065] 11) End.

[0066] The beneficial effects of the present invention are:

[0067] Through deep reinforcement learning, we achieve intelligent motion planning for the entire contact machining process, ensuring safe and reliable machining parameters and avoiding self-excited chatter between the robot arm and the workpiece due to improper parameters. This invention is not limited to grinding and polishing planning but is also applicable to other contact machining processes such as welding and spraying. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 This is a flowchart of a motion planning method for a contact processing robot based on deep reinforcement learning according to the present invention;

[0069] Figure 2 This is a structural diagram of the DH modeling of the segmented card robot arm system of the present invention;

[0070] Figure 3 Processing a three-dimensional model of the contact surface for the present invention;

[0071] Figure 4 This is the structure diagram of the dual-delay deep deterministic policy gradient network of the present invention. DETAILED DESCRIPTION

[0072] The present invention will be further described below with reference to the accompanying drawings and examples.

[0073] like Figure 1 As shown, the present invention provides a contact processing robot motion planning method based on deep reinforcement learning, comprising:

[0074] In the first step, a general dynamic model of contact machining is obtained based on the position-level forward and inverse robot coupling kinematic models of the manipulator system and the machining contact surface model;

[0075] The second step is to build a multi-level reward function based on multiple constraints and plan the operation parameters;

[0076] The third step is to design an algorithm structure based on the dual Actor-Critic framework based on the general dynamics model of contact processing, and establish a triple-improved dual-delay deep deterministic policy gradient network;

[0077] The fourth step is to train and load the triple-improved dual-delay deep deterministic policy gradient network to plan the robot's motion trajectory and optimize the contact processing intelligent safety parameters.

[0078] Example

[0079] Taking a grinding and polishing robot as an example, the first step is specifically implemented as follows:

[0080] The grinding and polishing robot system studied in this invention selects the joint card robot arm, and the definition of the DH coordinate system is shown in Figure 2 , each joint coordinate system is defined as , the corresponding DH parameters are shown in Table 1, where Indicates the length of the i-th arm, the installation coordinate system of the robot arm Coordinate system with the center of mass of the base Pointing in the same direction, the installation origin is at the center of mass of the base:

[0081] Table 1

[0082]

[0083] Solve the forward kinematics model of the robotic arm:

[0084] Connecting rod transformation based on modified DH parameter representation The general formula is:

[0085] ,

[0086] In the DH parameter notation, θi represents the angle between the two adjacent equivalent straight rods adjacent to joint i, ai represents the length of the equivalent straight rod of link i, αi represents the angle between the two adjacent joint axes of link i, and di represents the distance between the two adjacent equivalent straight rods of the chain of joint i.

[0087] Homogeneous change matrix of the tool coordinate system at the end of the robot arm relative to the inertial coordinate system for:

[0088] ,

[0089] Where, is the homogeneous transformation matrix from the inertial coordinate system to the base mass center coordinate system, is the homogeneous transformation matrix from the base mass center coordinate system to the robot arm installation coordinate system, is the homogeneous transformation matrix from the end joint coordinate system to the robot arm installation coordinate system, is the homogeneous transformation matrix from the end joint coordinate system to the tool coordinate system at the end of the robot arm.

[0090] The inverse kinematics of the manipulator is calculated, and the joints of the manipulator with SRS configuration (three ball joints and three revolute joints) are classified into non-coplanar joint groups and coplanar parallel axis joint groups for solution.

[0091] like Figure 3 As shown in the figure, the geometric modeling process is given by taking grinding and polishing as an example.

[0092] The following assumptions are made for the grinding and polishing process:

[0093] (1) The contact area is small compared to the disc and workpiece geometry;

[0094] (2) The workpiece surface is smooth and the surface equation is continuous to the second order;

[0095] (3) The polishing disc is much softer than the workpiece, so there is no deformation of the workpiece;

[0096] (4) The deformation of the disk is caused only by the normal contact force, without considering the tangential contact force;

[0097] (5) Compared with the workpiece curvature radius, the material removal depth during polishing is negligible.

[0098] Assume that a disk with a radius of R = 200 mm and a height of H = 50 mm contacts the workpiece at the initial point O. A standard orthogonal coordinate system O-XYZ is established with O as the origin, where the Z axis is the surface normal, the X axis is along the tool feed direction, and Od is the center of the disk bottom surface. The Y axis is given by the right-hand rule. The disk contact depth Defined as the displacement of the disk along the surface normal, the disk tilt angle Defined as the angle between the bottom surface of the disk and the tangent plane of the workpiece, the flexible disk is tilted at an angle Press into the workpiece to the contact depth After that, the bottom of the disk produces a local deformation at the initial point O, and the plane expressed by the bottom of the disk Written as follows:

[0099] ,

[0100] x and y represent the two-dimensional coordinates in the plane.

[0101] According to the classical differential geometry theory, the smooth workpiece contour near the contact point O is approximated as a quadratic function:

[0102] ,

[0103] Where, represents the principal curvature of the contact surface, represents the minor curvature of the contact surface, represents the angle between the feed direction of the flexible grinding disc and the principal curvature of the contact surface, and T represents the transpose of the matrix.

[0104] The second step is specifically implemented as follows:

[0105] (1) Task constraints

[0106] The distance between the end effector and the corresponding capture point (denoted as d) is the key to the design of the reward function. In order to make the reward increase as the relative distance decreases, the reward function is taken as A direct proportional function of the form, when the end effector reaches a small area near the target, it is expected to give the manipulator an increasing reward for accelerating to the target point. To this end, a logarithmic function is introduced:

[0107] ,

[0108] Where, The term avoids singularities in the logarithmic function and limits the maximum reward. With k1 = 2, as d decreases, the reward initially increases at a constant rate, ultimately evenly planning to a closer location. With k2 = 0.5, when d decreases to within 1 meter, the reward increases significantly, further improving planning accuracy and achieving the design goal.

[0109] (2) Feed rate constraint

[0110] The magnitude and direction of feed speed are important factors affecting the grinding and polishing effect. When the flexible grinding wheel speed at the end is too high, a collision may occur, causing damage to the workpiece and the driving robot arm. When the end is far from the target position, the speed constraint is ignored. When the distance between the end and the target position is within a safe range, a speed and distance penalty function is established:

[0111] ,

[0112] The penalty function of the analysis design is: e represents the grinding wheel feed speed, and d represents the distance between the grinding wheel and the workpiece. k3 = 1 determines the penalty strength, that is, the degree to which the speed constraint is affected by the distance, and k4 = 0.1 determines the maximum penalty value.

[0113] (3) Feed angle constraint

[0114] The feed angle constraint is essentially a constraint on the end-arm posture. In the contact surface modeling, the angle between the grinding wheel feed direction and the principal curvature of the contact surface is defined. , Too large a force will generate a large tangential force and cause damage to the tool. Therefore, it is necessary to plan the posture of the end effector of the robot arm during processing to control the feed angle so that Within a safe range, the designed reward function is as follows:

[0115] ,

[0116] Analyze the designed reward function. When the distance is large, no constraint is imposed on the robot arm posture. k5=1 determines the extent to which the reward value is affected by the distance, that is, the size of the safety distance. k6=0.2 affects the size of the reward value.

[0117] Taking the triple constraints into consideration, the comprehensive reward function is proposed as follows:

[0118] ,

[0119] The third step is specifically implemented as follows:

[0120] The designed network structure is as follows Figure 4 As shown, a dual actor-critic framework is used to construct the target network and the online network. An experience replay buffer is introduced to enable batch checking during the training process. Three improvements are made to the network to improve the algorithm's efficiency.

[0121] First, simplify the two Q networks. In the Actor-Critic network, the update strategy is:

[0122] ,

[0123] Where r represents the cumulative reward value, γ is the discount factor, is the policy function.

[0124] In actual calculations, when the Actor-Critic network updates slowly, the two networks are similar and lack the independence conditions to make independent estimates. Therefore, a biased estimation function is given.

[0125] ,

[0126] Select the smaller value between the two as the final estimator .

[0127] Second, a delayed update strategy is proposed. During the learning process, the target network serves as a function approximator. This provides a stable target for the learning process, accelerating network convergence and significantly improving stability. To minimize error propagation, the policy network is designed to update at a lower frequency than the value network. Lower policy update frequency reduces the variance of the value function updates, resulting in higher-quality policies. By sufficiently delaying policy updates, the possibility of repeated updates to unchanged critics is limited.

[0128] The third step is to smooth the target policy so that the actions have a normalized range. A regularization strategy is introduced to map similar values ​​to similar action spaces, normalizing the values ​​to a region in the action space. This is achieved by adding a small amount of normally distributed noise to the actions and averaging them over a pool of mini-batches of experience.

[0129] The fourth step is specifically implemented as follows:

[0130] 1) If Figure 4 As shown, first initialize the Critic network of the online network With Actor Network , set the state vector as: the robot arm joint angle and angular velocity, the speed and posture of the robot arm end effector, the distance between the end effector and the workpiece, etc., and assign random initial values ​​to the state;

[0131] 2) Initialize the Critic network of the target network With Actor Network , initialize the experience buffer R and set the upper limit of R to 10000;

[0132] 3) Randomly give a target planning position and obtain the current state s of the robot arm;

[0133] 4) Select the current action after adding detection noise according to the current state s , N t For random noise, set it to Gaussian noise;

[0134] 5) Execute the action and obtain the reward value r and the new state s', and Stored in the experience pool;

[0135] 6) Randomly sample N = 64 groups of samples from the experience pool to form a minimum batch set;

[0136] 7) Normalize the actions and streamline the dual Q network using the optimization method described in step 3;

[0137] 8) Update the critic network based on the loss function according to the delayed update strategy described in step 3;

[0138] 9) Update the Actor network based on the policy gradient;

[0139] 10) Update the target network;

[0140] 11) End.

[0141] Based on the triple improvement, the learning parameters are selected as follows: the Actor network has two hidden layers, each layer has 200 neurons, and the learning rate is 1.0×10-4; the Critic network has two hidden layers, each layer has 200 neurons, and the learning rate is 1.0×10-4; the total number of learning rounds is 5000 rounds, and the maximum number of steps in a single round is 128 steps; after the network converges, the network is loaded into the joint card robot arm and planning instructions are issued. During the planning process, the robot arm will update the information in the state vector in real time after each step of action, and call the Actor-critic network based on the current state value to obtain the next step instruction. In this way, a safe and stable approach to the workpiece can be achieved, and the feed speed and feed angle are the target values. The technology of the present invention is not only limited to grinding and polishing planning, but is also applicable to other types of contact processing processes such as welding and spraying.

[0142] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A motion planning method for a contact processing robot based on deep reinforcement learning, characterized in that: The following steps are involved: In the first step, a general dynamic model of contact machining is obtained based on the position-level forward and inverse robot coupling kinematic models of the manipulator system and the machining contact surface model; The second step is to build a multi-level reward function based on multiple constraints and plan the operation parameters; this includes: Constructing a reward function based on processing task constraints : , Where d represents the distance between the end effector and the corresponding processing point, k1 and k2 are the first and second proportional coefficients; Constructing a reward function based on feedrate constraints : , Where, v e represents the feed speed of the flexible grinding disc, k3 represents the proportional coefficient of the speed constraint affected by the distance, and k4 represents the proportional coefficient affecting the speed constraint reward value; Constructing a reward function based on feed angle constraints : , Where, It represents the angle between the grinding wheel feed direction and the principal curvature of the contact surface, k5 represents the proportional coefficient of the degree to which the angle constraint reward value is affected by the distance, and k6 represents the proportional coefficient that affects the angle constraint reward value; Give a comprehensive reward function for: ; The third step is to design an algorithm structure based on the dual Actor-Critic framework based on the general dynamics model of contact processing, and establish a triple-improved dual-delay deep deterministic policy gradient network; The first improvement includes: both Q networks are Actor-Critic networks: , 、 Represent the biased estimation functions of the two Actor-Critic networks respectively; 、 Represent the policy functions of the two Actor-Critic networks respectively; Select 、 The smaller value between the two is used as the estimated function estimator ; represents the discount factor; The second improvement includes: during the learning process, the target network is used as a function approximator to update the network, and the update frequency of the actor network is lower than that of the critic network; The third improvement includes: introducing a regularization strategy to regularize the value to a region in the action space; The fourth step is to train and load the triple-improved dual-delay deep deterministic policy gradient network to plan the robot's motion trajectory and optimize the contact processing intelligent safety parameters.

2. A contact processing robot motion planning method based on deep reinforcement learning according to claim 1, characterized in that: The first step is specifically implemented as follows: Establish the forward kinematics model of the manipulator, and transform the i-1th coordinate system of the manipulator link to the i-th coordinate system according to the modified DH parameter representation The general formula is: , In DH parameter representation, represents the angle between the joint of the i-th coordinate system and the two adjacent equivalent straight rods, represents the length of the equivalent straight rod of the connecting rod in the i-th coordinate system, represents the angle between two adjacent joint axes of the link in the i-th coordinate system, represents the distance between two adjacent equivalent straight rods connected to the joint in the i-th coordinate system; Homogeneous change matrix of the tool coordinate system at the end of the robot arm relative to the inertial coordinate system for: , Where, is the homogeneous transformation matrix from the inertial coordinate system to the base mass center coordinate system, is the homogeneous transformation matrix from the base mass center coordinate system to the robot arm installation coordinate system, is the homogeneous transformation matrix from the end joint coordinate system to the robot arm installation coordinate system, is the homogeneous transformation matrix from the end joint coordinate system to the tool coordinate system at the end of the robot arm; Establish an inverse kinematics model for the robotic arm, classify the robotic arm joints of the SRS configuration into a non-coplanar joint group and a coplanar parallel axis joint group for solution. The SRS configuration includes three ball joints and three revolute joints. The non-coplanar joint group is solved by multiplying the inverse matrix on the left and making the corresponding elements equal. Then, the solution of the non-coplanar joint group is brought in and the coplanar parallel axis joint group is solved according to the properties of trigonometric functions. Establish a machining contact surface model, and equate the geometric structure of the flexible grinding disc of the machining tool to a disc. Assume that the disc with a radius of R contacts the workpiece at the initial point O. Establish a standard orthogonal coordinate system O-XYZ with O as the origin. The Z axis is the direction of the surface normal, the X axis is the direction of the tool feed, Od is the center of the disc bottom surface, and the Y axis is given by the right-hand rule. The contact depth of the disc is Defined as the displacement of the disk along the surface normal, the disk tilt angle Defined as the angle between the bottom surface of the disk and the tangent plane of the workpiece, the flexible disk is tilted at an angle Press into the workpiece to the contact depth After that, the bottom of the disk produces a local deformation at the initial point O, and the plane expressed by the bottom of the disk Written as follows: , x, y represent the two-dimensional coordinates in the plane; According to the classical differential geometry theory, the smooth workpiece contour within the preset range of the initial point O is approximated as a quadratic function : , Where, represents the principal curvature of the contact surface, represents the minor curvature of the contact surface, represents the angle between the feeding direction of the flexible grinding disc and the principal curvature of the contact surface, and T represents the transpose of the matrix; Based on the established manipulator position-level forward and inverse robot kinematic models and machining contact surface model, a general contact machining model of the contact machining robot is obtained.

3. The motion planning method for a contact processing robot based on deep reinforcement learning according to claim 1, characterized in that: The fourth step is specifically implemented as follows: 1) Initialize the online network's critic network With Actor Network , set the state vector and assign a random initial value, the state vector includes the joint angle and angular velocity of the robot arm, the speed and posture of the end effector of the robot arm, and the distance between the end effector and the workpiece; 2) Initialize the Critic network of the target network With Actor Network , initialize the experience buffer; 3) Randomly give the target planning position and obtain the current state s of the robot arm; 4) Select the current action after adding random noise according to the current state s , N t is the random noise added; 5) Execute the current action And according to the designed reward function Get the reward value and the new state s', and Stored in the experience buffer; 6) Randomly sample N groups of samples in the experience buffer to form a minimum batch set; 7) Based on the algorithm structure based on the dual Actor-Critic framework described in step 3, normalize the actions and streamline the dual Q network; 8) Update the critic network based on the double-delayed deep deterministic strategy and loss function described in step 3; 9) Update the Actor network based on the policy gradient function; 10) Update the target network; 11) End.

Citation Information

Patent Citations

  • Robot path planning method and device based on reinforcement learning

    CN118163101A

  • A welding method and welding robot based on artificial intelligence

    CN118237825B

  • Cutter motion control method of five-axis machining robot and related device

    CN118238156A

  • Spacecraft attitude fault-tolerant control method based on iterative-learning disturbance observer

    CN107121961A

  • Unmanned vehicle path planning method based on improved A * algorithm and deep reinforcement learning

    CN111780777A