Man-machine cooperation compliance control method based on improved deep reinforcement learning in combination with intention of collaborator

By improving the deep reinforcement learning algorithm combined with the intention of the collaborator, a robotic arm kinematics and dynamics model is established to realize the flexible control of the robotic arm in complex environments, solving the instability and insecurity problems in the human-machine collaboration process, and improving the perception and anti-interference ability of the robotic arm.

CN120395841APending Publication Date: 2025-08-01NANJING TECH UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510601183.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In the prior art, it is difficult to achieve the robotic arms having perception, flexibility and anti-interference capabilities in complex environments in human-machine collaboration, resulting in unstable and unsafe human-machine collaboration process.

Method used

The improved deep reinforcement learning algorithm is used to combine collaborative intent with collaborative intent, and by building a robotic arm kinematics and dynamics model, the mapping relationship between collaborative intent and robotic arm motion and interactive forces is established, and the nonlinear function is approximateed by radial basis neural network, damping parameters are adaptively adjusted, combined with Gaussian process optimization, and strategy switching mechanism is designed to achieve the flexible control of the robotic arm.

Benefits of technology

It improves the stability, safety and efficiency of man-machine collaboration of the robotic arm in complex environments, and enhances the perception and anti-interference ability of the robotic arm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120395841A_ABST
    Figure CN120395841A_ABST
Patent Text Reader

Abstract

The invention provides a man-machine cooperation compliance control method based on improved deep reinforcement learning in combination with intentions of collaborators. Estimating the motion intention of the human in real time based on a radial basis function neural network; a strategy network in a traditional DDPG is replaced with a DDPG algorithm combined with GP, and optimization of impedance parameters in self-adaptive impedance control is achieved. Aiming at hyper-parameter optimization in the GP model, a k-fold cross validation method is adopted; a smooth switching strategy is adopted, and a proper strategy is selected. Finally, the flexibility of the mechanical arm is guaranteed, and meanwhile the man-machine cooperation process is safely and efficiently achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of compliant control methods for robotic arms during human - robot collaboration, and particularly to a compliant control method for robotic arms that senses the intention of the collaborator and combines an improved deep reinforcement learning algorithm. Background Art

[0002] In the context of Industry 4.0, the demand for production methods is getting higher and higher. In traditional situations, for operations such as heavy object handling, when a heavy object needs to be clumsily delivered to a difficult - to - access area, it is very challenging for human operators to keep it in a predetermined position and orientation, which is ergonomically difficult and gradually shows limitations in current industrial production. Traditional control methods in the field of human - robot collaboration applications also cannot meet the current production requirements. Nowadays, robots generally operate in complex environments full of uncertainties, requiring robots to have certain sensing capabilities, compliance, and anti - interference capabilities to ensure the stability, safety, and efficiency of industrial robots during the process of contact and interaction with humans and the environment.

[0003] Human - robot collaboration combined with intelligent control can well solve the above problems. In recent years, reinforcement learning has been regarded as one of the core key technologies of intelligent systems and has been widely applied in different fields such as artificial intelligence, machine learning, and robotics. At the beginning of its development, the research of reinforcement learning mainly focused on small - scale discrete action spaces, and now it has developed into learning control for large - scale continuous state spaces. The deep reinforcement learning formed by the combination of deep learning and reinforcement learning further improves the ability of robots to learn operation skills. Among all current fields of robot control, the robot control strategy based on deep reinforcement learning can make robots perform more intelligently and is the best among all control strategies.

[0004] Therefore, this application proposes an intelligent compliant control method for robotic arms that senses the intention of the collaborator and combines an improved deep reinforcement learning algorithm. Summary of the Invention

[0005] Aiming at the defects existing in the prior art, the purpose of the present invention is to provide a robotic arm intelligent impedance control method that can sense the intention of the collaborator and combine an improved DDPG algorithm. This control method can sense the intention trajectory of the collaborator during the human - robot collaboration process, combine the deep deterministic policy gradient algorithm to optimize the impedance parameters in real - time, and achieve safe, efficient, and compliant control during the human - robot collaboration process.

[0006] To achieve the above purpose,

[0007] A human - robot collaboration compliant control method based on the intention of the collaborator and improved deep reinforcement learning is built using the following steps:

[0008] Step 1: Build the kinematic and dynamic models of the robotic arm in Cartesian space. According to the spring-mass-damper model, establish the upper limb model of the collaborator during the human-robot collaboration process; establish the mapping relationship between the motion intention of the collaborator and the motion and interaction force of the robotic arm during the human-robot collaboration process;

[0009] Step 2: Based on Step 1, establish the non-linear function relationship between the collaborator's intention trajectory, the external interaction force, the end position, and the end velocity. Use a radial basis neural network approximation to obtain the collaborator's intention trajectory;

[0010] Step 3: Establish an adaptive impedance control model as the initial strategy. In the control strategy, adaptively adjust the damping coefficient according to the external force during the end human-robot collaboration process;

[0011] Step 4: Establish an intelligent strategy. Use the DDPG algorithm combined with Gaussian process to optimize the damping parameters. The Gaussian process is used to approximate the policy network in the DDPG algorithm, and the hyperparameters are optimized through k-fold cross-validation;

[0012] Step 5: Input the collaborator's intention trajectory as the ideal trajectory into the intelligent impedance control algorithm. The admittance control module, inverse kinematics module, joint controller module, and forward kinematics module execute this instruction, and smoothly switch through a selector to select appropriate strategies at different stages.

[0013] When executing Step 1, the kinematics of the robotic arm is defined by the following formula:

[0014]

[0015] where \(x(t)\), and represent the position, velocity, and acceleration of the end of the robotic arm in Cartesian space respectively; \(q\), and represent the angular velocity, angular velocity, and angular acceleration of the robotic arm in joint space respectively; \(J(q)\) and represent the Jacobian matrix and the first derivative of the Jacobian matrix respectively; \(\psi(q)\) represents the kinematic model of the robotic arm;

[0016] The dynamic equation of the robotic arm in joint space is as follows:

[0017]

[0018] where \(M(q)\) is the inertia diagonal matrix, is the Coriolis and centrifugal force matrix, \(G(q)\) is the gravity vector matrix, and \(\tau\) is the torque vector of the robotic arm control input;

[0019] Human-robot collaboration occurs in the Cartesian operation space. According to the above kinematic and dynamic equations, the dynamic equation of the robotic arm in the Cartesian space is given as follows:

[0020]

[0021] In the formula,

[0022] where u is the control force of the robotic arm in the Cartesian coordinate system, and f ext is the interaction force with the external environment;

[0023] According to the spring-mass-damper model, an upper limb damping model of the collaborator is established, and the specific formula is as follows:

[0024]

[0025] where x, are the position, velocity, and acceleration of the contact point between the end of the collaborator's limb and the environment; f ext represents the force between the end of the collaborator's upper limb and the environment; x hd represents the motion intention trajectory of the collaborator; is the upper limb damping of the collaborator; K h (x) is the upper limb stiffness of the collaborator; is considered to be the uncertainty caused by the inaccuracy of modeling, external interference, and the time-varying characteristics of impedance parameters;

[0026] The position, velocity, and environmental contact force are obtained through sensors. The motion intention x of the collaborator hd can be obtained through the contact force f at the interaction point ext , the actual position x, and the actual velocity to obtain the following functional mapping form:

[0027]

[0028] where E(·) represents an unknown non-linear function related to f ext , x, .

[0029] 3. A human-robot collaboration compliant control method based on improving deep reinforcement learning by combining the intention of the collaborator as described in claim 1, characterized in that when performing step 2, a radial basis network is used to approximate the non-linear function E(·); specifically including the following steps:

[0030] Step A1: Design the structure of the RBF neural network, including the input layer, hidden layer, and output layer, specifically

[0031] as follows:

[0032] The input signals are the interaction force, the position and velocity of the end effector of the robotic arm, as follows:

[0033]

[0034] The Gaussian function is used as the activation function for the hidden layer. The output of the i-th hidden layer node is:

[0035]

[0036] where, u i and η i are the center and width of the Gaussian function, respectively.

[0037] The output is the estimated value of the motion intention trajectory, as follows:

[0038] x hd =W T S(z)(8)

[0039] where, W is the weight matrix from the hidden layer to the output layer, and S(z) is the output vector of the hidden layer;

[0040] Step A2: Define the cost function and design the cost function as follows:

[0041]

[0042] where, α, β, and γ are weight coefficients, and x k , f k are the position error, velocity error, and interaction force, respectively.

[0043] Step A3: The network weights are adjusted online in the direction of the cost function descent, as follows:

[0044]

[0045] where, the network output u = W T φ(s), φ(s) is the input of the radial basis function, and s is the input vector; and are estimated online; E k , W k are the cost function and weights of the neural network at the k-th time step, respectively; η is the learning rate.

[0046] When performing Step 3, the damping parameter B is adjusted using adaptive impedance control, which specifically includes the following steps:

[0047] Step B1: The robotic impedance model formula is as follows:

[0048]

[0049] Among them, E = X d - X, X d is the desired position of the end of the robotic arm, and X is the actual position of the end of the robotic arm; M d , B d and K d are the inertia matrix, damping matrix, and stiffness matrix;

[0050] Step B2: Based on the feedback of the external force, adjust the damping coefficient. The adaptive impedance control model of the robotic arm is as follows:

[0051]

[0052] Among them, B vi is the variable damping matrix; S is the adjustment matrix; s ij is the element in the S matrix, α is the weight factor; δ ij represents 1 when i = j, otherwise 0; sgn is the sign function; is the i-th component of the Cartesian velocity vector; is the time derivative of the i-th component of the force;

[0053] Step B3: Stability analysis. The impedance models of the collaborator and the human are both linear. The equation of the entire system is expressed as a second-order linear equation, and the equation is as follows:

[0054]

[0055] Rewritten as the state-space equation, as follows:

[0056]

[0057] Among them, A is the system characteristic matrix, T is the time constant of the first-order low-pass filter. When the real part of the system characteristic matrix is less than 0, the system is stable.

[0058] When performing Step 4, the DDPG algorithm combined with the Gaussian process is used to optimize the damping parameters, which specifically includes the following steps:

[0059] Step C1: Use the GP model to approximate the policy function in the DDPG algorithm. The input of the GP model is the state S t , and the output is the action A t ; Define the state S t as the current position, velocity, and contact force of the robot, and the action A t as the damping parameter B;

[0060] The Gaussian process model describes the relationship between the state and the action through the covariance function, and the formula is as follows:

[0061] γ = f(X) + N(0, σn 2 )~N(0, K(X, X)+σ n 2 I)(16)

[0062] A t =u=K(S t , X)(K(X, X)+σ n 2 I) -1 γ(17)

[0063]

[0064] where X = {S1, …, S Ne}, γ = {A1, …, A Ne}, A i =B θi is the action corresponding to s i , f(X) is a normal distribution with mean zero, N(0, σ

[0065] ) is the noise of independent distribution; u is the mean, K(·) is the covariance n 2 matrix, k(·) is the element in the covariance matrix, σ

[0066] 、l f 、l i,j are hyperparameters;

[0067] Step C2: The hyperparameters of the GP model are optimized by k-fold cross-validation: First, the training set is randomly divided into k subsets of equal size; Second, for each combination of hyperparameters, k trainings and validations are performed.

[0068] Each time during training, k - 1 subsets are used as training data, and the remaining 1 subset is used as validation data. Finally, the performance metrics of each validation are calculated, and the average performance of the k validations is taken as the final performance evaluation of this combination of hyperparameters;

[0069] Combined with Markov chain Monte Carlo, determine the possible values of the hyperparameters and their probabilities; Combine active learning with cross-validation, and according to the prediction uncertainty of the current model for different validation sets, preferentially select the validation set that can provide more information; Dynamically adjust the parameter structure in the kernel function during the optimization process;

[0070] Step C3: Design the reward function in DDPG, which includes four parts: force error, speed error, energy consumption, and intention matching reward;

[0071]

[0072] w1(t) = w 1,i *e -βt (20)

[0073] w2(t) = w 2,i *(1 - e -βt )(21)

[0074] where w1(t) is the dynamic weight value of the force error and velocity error terms, w2(t) is the dynamic weight value of the energy consumption term, w 1,i , w 2,i are the initial weights, β is the decay rate; r i is the intention matching reward term, and the formula is as follows:

[0075]

[0076] where x is the actual trajectory of the robotic arm tracking the collaborator in real time, x hd is the trajectory estimated by the radial basis neural network, ||·|| represents the Euclidean distance, γ is the maximum reward value, and σ is the bandwidth parameter.

[0077] When performing step 5, a strategy switching mechanism is designed to gradually transition from the initial strategy to the intelligent strategy through the weight function; the action output π of the hybrid strategy is a weighted combination of the initial strategy action π I and the intelligent strategy action π θ , as follows:

[0078] π(t) = α(t)π I + (1 - α(t))π θ (23)

[0079] where α(t) is the weight function to achieve a smooth transition of the switch, and is designed as an exponential decay function; the dynamic switching condition is set as follows:

[0080]

[0081] where ε and δ are the thresholds of the force error and velocity error respectively.

[0082] The present invention proposes a human - robot collaborative compliant control method based on improved deep reinforcement learning by combining the intentions of collaborators, which solves the problem that in an uncertain complex environment, the robot is required to have certain perception capabilities, compliance, and anti - interference capabilities, and realizes the stability, safety, and efficiency of the robotic arm during the contact and interaction process with humans and the environment. Description of the Drawings

[0083] Figure 1 is the structural diagram of the intelligent admittance control system;

[0084] Figure 2Block diagram of adaptive impedance control based on intention estimation;

[0085] Figure 3 Structural schematic diagram of the DDPG algorithm combined with the GP model;

[0086] Figure 4 Technical roadmap for optimizing the hyperparameters of the GP model by k-fold cross-validation. Specific implementation plan

[0087] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0088] As Figures 1-4 shown, a human-robot collaborative compliant control method based on improved deep reinforcement learning combined with the intention of a collaborator includes the following steps:

[0089] Step 1: Build a kinematic and dynamic model of the robotic arm in the Cartesian space. According to the spring-mass-damper model, establish an upper limb model of the collaborator during the human-robot collaboration process. Finally, establish the mapping relationship between the motion intention of the collaborator and the motion and interaction force of the robotic arm during the human-robot collaboration process.

[0090] Step 2: Based on Step 1, establish the non-linear functional relationship between the collaborator's intention trajectory, the external interaction force, the end position, and the end velocity. Use a radial basis neural network approximation to obtain the collaborator's intention trajectory.

[0091] Step 3: Establish an adaptive impedance control model as the initial strategy. In the control strategy, adaptively adjust the damping coefficient according to the external force during the end human-robot collaboration process.

[0092] Step 4: Establish an intelligent strategy. Use the DDPG algorithm combined with Gaussian process to optimize the damping parameters. The Gaussian process is used to approximate the policy network in the DDPG algorithm, and the hyperparameters are optimized by k-fold cross-validation.

[0093] Step 5: The collaborator's intention trajectory is used as the ideal trajectory and input into the intelligent impedance control algorithm. The admittance control module, inverse kinematics module, joint controller module, and forward kinematics module execute this instruction, and through the selector, smoothly switch and select the appropriate strategy at different stages.

[0094] Preferably, when executing Step 1, the kinematics of the robotic arm is defined by the following formula:

[0095]

[0096] where x(t), and respectively represent the position, velocity, and acceleration of the end of the robotic arm in the Cartesian space; q, and respectively represent the angular velocity, angular velocity and angular acceleration of the robotic arm in the joint space; J(q) and respectively represent the Jacobian matrix and the first derivative of the Jacobian matrix; ψ(q) represents the kinematic model of the robotic arm.

[0097] The dynamic equation of the robotic arm in the joint space is as follows:

[0098]

[0099] where, M(q) is the inertia diagonal matrix, is the Coriolis force and centrifugal force matrix, G(q) is the gravity vector matrix, and τ is the torque vector of the robotic arm control input.

[0100] Human-robot collaboration occurs in the Cartesian operation space. According to the above kinematic and dynamic equations, the dynamic equation of the robotic arm in the Cartesian space is given as follows:

[0101]

[0102] In the formula,

[0103] where, u is the control force of the robotic arm in the Cartesian coordinate system, and f ext is the interaction force of the external environment.

[0104] According to the spring-mass-damper model, an upper limb damping model of the collaborator is established, and the specific formula is as follows:

[0105]

[0106] where, x, are the position, velocity and acceleration of the contact point between the end of the collaborator's limb and the environment; f ext represents the force between the end of the collaborator's upper limb and the environment; x hd represents the motion intention trajectory of the collaborator; is the upper limb damping of the collaborator; K h (x) is the upper limb stiffness of the collaborator; is considered to be the uncertainty caused by the inaccuracy of modeling, external interference and the time-varying characteristics of impedance parameters.

[0107] The position, velocity and environmental contact force are obtained through sensors. The motion intention x hd of the collaborator can be obtained through the contact force f ext at the interaction point, the actual position x and the actual velocity to obtain the following function mapping form:

[0108]

[0109] Among them, E(·) represents an unknown non - linear function related to f ext , x, .

[0110] Preferably, when performing step 2, a radial basis network is used to approximate the non - linear function E(·), which specifically includes the following steps:

[0111] Step A1: Design the structure of the RBF neural network, including the input layer, hidden layer, and output layer, specifically as follows:

[0112] The input signals are the interaction force, the position and velocity of the end - effector of the manipulator, as follows:

[0113]

[0114] The hidden layer uses the Gaussian function as the activation function, and the output of the i - th hidden - layer node is:

[0115]

[0116] Among them, u i and η i are the center and width of the Gaussian function respectively.

[0117] The output is the estimated value of the motion intention trajectory, as follows:

[0118] x hd =W T S(z)(8)

[0119] Among them, W is the weight matrix from the hidden layer to the output layer, and S(z) is the output vector of the hidden layer.

[0120] Step A2: Define the cost function and design the cost function, as follows:

[0121]

[0122] Among them, α, β, γ are weight coefficients, x k , f k are the position error, velocity error, and interaction force respectively.

[0123] Step A3: The network weights are adjusted online in the direction of the cost - function descent, and there is:

[0124]

[0125] Among them, the network output u = W T φ(s), φ(s) is the input of the radial basis function, and s is the input vector; and Obtained by online estimation; E k , W k are the cost function and weights of the neural network at the k-th time step respectively; η is the learning rate.

[0126] Preferably, when performing step 3, an adaptive impedance control is used to adjust the damping parameter B, specifically including

[0127] the following steps:

[0128] Step B1: The robot impedance model formula is as follows:

[0129]

[0130] where, E = X d - X, X d is the desired position of the end of the robotic arm, and X is the actual position of the end of the robotic arm; M d , B d and K d are the inertia matrix, damping matrix and stiffness matrix.

[0131] Step B2: Based on the feedback of the external force, adjust the damping coefficient. The adaptive impedance control model of the robotic arm is as follows:

[0132]

[0133] where, B vi is the variable damping matrix; S is the adjustment matrix; s ij is the element in the S matrix, α is the weight factor; δ ij represents 1 when i = j, otherwise 0; sgn is the sign function; is the i-th component of the Cartesian velocity vector; is the time derivative of the i-th component of the force.

[0134] Step B3: Stability analysis. The impedance models of the collaborator and the human are both linear. The equation of the whole system is expressed as a second-order linear equation, and the equation is as follows:

[0135]

[0136] Rewritten as the state space equation, as follows:

[0137]

[0138] where, A is the system characteristic matrix, T is the time constant of the first-order low-pass filter. When the real part of the system characteristic matrix is less than 0, the system is stable.

[0139] Preferably, when performing step 4, the DDPG algorithm combined with the Gaussian process is used to optimize the damping parameter, specifically including the following steps:

[0140] Step C1: Use the GP model to approximate the policy function in the DDPG algorithm. The input of the GP model is the state S t , and the output is the action A t ; Define the state S t as the current position, speed, and contact force of the robot, and the action A t as the damping parameter B.

[0141] The Gaussian process model describes the relationship between the state and the action through the covariance function, and the formula is as follows:

[0142] γ = f(X) + N(0,σ n 2 ) ∼ N(0, K(X, X) + σ n 2 I)(16)

[0143] A t = u = K(S t , X)(K(X, X) + σ n 2 I) -1 γ(17)

[0144]

[0145] where X = {S1, …, S Ne}, γ = {A1, …, A Ne}, A i = B θi is the action corresponding to s i , f(X) is

[0146] a normal distribution with a mean of zero, N(0, σ n 2 ) is the noise of the independent distribution; u is the mean, K(·) is the covariance

[0147] matrix, k(·) is the element in the covariance matrix, and σ f , l i,j are hyperparameters.

[0148] Step C2: Optimize the hyperparameters of the GP model through k-fold cross-validation: First, randomly divide the training set into k subsets of equal size; Second, for each combination of hyperparameters, perform k times of training and validation.

[0149] Each time during training, use k - 1 subsets as the training data, and the remaining 1 subset as the validation data. Finally, calculate the performance metrics for each validation, and take the average performance of the k validations as the final performance evaluation of this combination of hyperparameters.

[0150] Combined with Markov Chain Monte Carlo, determine the possible values of hyperparameters and their probabilities; combine active learning with cross-validation, and according to the prediction uncertainty of the current model for different validation sets, preferentially select the validation set that can provide more information; dynamically adjust the parameter structure in the kernel function during the optimization process.

[0151] Step C3: Design the reward function in DDPG, which includes four parts: force error, speed error, energy consumption, and intention matching reward.

[0152]

[0153] w1(t) = w 1,i *e -βt (20)

[0154] w2(t) = w 2,i *(1 - e -βt )(21)

[0155] where w1(t) is the dynamic weight value of the force error and speed error terms, w2(t) is the dynamic weight value of the energy consumption term, w 1,i 、w 2,i are the initial weights, β is the decay rate; r i is the intention matching reward term, and the formula is as follows:

[0156]

[0157] where x is the actual trajectory of the robotic arm tracking the collaborator in real time, x hd is the trajectory estimated by the radial basis neural network, ||·|| represents the Euclidean distance, γ is the maximum reward value, and σ is the bandwidth parameter.

[0158] Preferably, when performing step 5, design a strategy switching mechanism to gradually transition from the initial strategy to the intelligent strategy through a weight function. The action output π of the hybrid strategy is composed of the initial strategy action π I and the intelligent strategy action π θ

[0159] weighted combination, as follows:

[0160] π(t) = α(t)π I +(1 - α(t))π θ (23)

[0161] where α(t) is the weight function to achieve a smooth transition of the switch, and is designed as an exponential decay function. Set the dynamic switching conditions as follows:

[0162]

[0163] where ε and δ are the thresholds of force error and velocity error respectively.

Claims

1. A human-robot collaborative compliant control method based on improved deep reinforcement learning by combining the intentions of collaborators, characterized in that It is built by following the steps below: Step 1: Build the kinematic and dynamic models of the robotic arm in the Cartesian space. According to the spring-mass-damper model, establish the upper limb model of the collaborator during the human-robot collaboration process; establish the mapping relationship between the collaborator's motion intention and the motion and interaction force of the robotic arm. Step 2: Based on Step 1, establish the non-linear function relationship between the collaborator's intention trajectory, external interaction force, end position, and end velocity. Use a radial basis neural network approximation to obtain the collaborator's intention trajectory. Step 3: Establish an adaptive impedance control model as the initial strategy. In the control strategy, adaptively adjust the damping coefficient according to the external force during the end human-robot collaboration process. Step 4: Establish an intelligent strategy. Use the DDPG algorithm combined with Gaussian process to optimize the damping parameters. The Gaussian process approximates the policy network in the DDPG algorithm, and the hyperparameters therein are optimized through k-fold cross-validation. Step 5: The collaborator's intention trajectory is input as the ideal trajectory into the intelligent impedance control algorithm. The instruction is executed by the admittance control module, inverse kinematics module, joint controller module, and forward kinematics module. Smooth switching is performed through a selector to select appropriate strategies at different stages.

2. The human - machine collaborative compliant control method based on improved deep reinforcement learning in combination with the intentions of collaborators according to claim 1, wherein, When performing Step 1, the kinematics of the robotic arm is defined by the following formula: where \(x(t)\), and represent the position, velocity, and acceleration of the end - effector of the robotic arm in Cartesian space, respectively; \(q\), and represent the angular velocity, angular velocity, and angular acceleration of the robotic arm in joint space, respectively; \(J(q)\) and represent the Jacobian matrix and the first - order derivative of the Jacobian matrix, respectively; \(\psi(q)\) represents the kinematic model of the robotic arm. The dynamic equation of the robotic arm in joint space is as follows: where \(M(q)\) is the inertia diagonal matrix, is the Coriolis force and centrifugal force matrix, \(G(q)\) is the gravity vector matrix, and \(\tau\) is the torque vector of the manipulator control input; Human-robot collaboration occurs in the Cartesian operation space. According to the above kinematic and dynamic equations, the dynamic equation of the robotic arm in Cartesian space is given as follows: In the formula, where u is the control force of the robotic arm in the Cartesian coordinate system, and f ext is the interaction force with the external environment; According to the spring-mass-damper model, establish the damping model of the collaborator's upper limb. The specific formula is as follows: where x, are the position, velocity, and acceleration of the contact point between the end of the collaborator's limb and the environment; f ext represents the force exerted by the end of the collaborator's upper limb on the environment; x hd represents the motion intention trajectory of the collaborator; is the damping of the collaborator's upper limb; K h (x) is the stiffness of the collaborator's upper limb; is considered to be the uncertainty caused by the inaccuracy of modeling, external disturbances, and the time-varying characteristics of impedance parameters; The position, velocity, and environmental contact force are obtained through sensors, and the motion intention x of the collaborator hd can be determined by the contact force f at the interaction point ext , the actual position x, and the actual velocity to obtain the function mapping form shown below: where E(·) represents an unknown non-linear function related to f ext , x, and 3. The human-robot collaborative compliant control method based on improved deep reinforcement learning in combination with the intentions of collaborators according to claim 1, characterized in that, When performing Step 2, use a radial basis network to approximate the non-linear function E(·). It specifically includes the following steps: Step A1: Design the structure of the RBF neural network, including the input layer, hidden layer, and output layer, as follows: The input signals are the interaction force, the end position and velocity of the robotic arm, as follows: The hidden layer uses the Gaussian function as the activation function. The output of the i-th hidden layer node is: where u i and η i are the center and width of the Gaussian function, respectively. The output is the estimated value of the motion intention trajectory, as follows: x hd = W T S(z) (8) where, W is the weight matrix from the hidden layer to the output layer, and S(z) is the output vector of the hidden layer; Step A2: Define the cost function and design the cost function as follows: where α, β, and γ are weight coefficients, and x k , f k are the position error, velocity error, and interaction force, respectively. Step A3: The network weights are adjusted online in the direction of the cost function decrease, and there is: Among them, the network output u = W T φ(s), where φ(s) is the input of the radial basis function and s is the input vector; and are obtained by online estimation; E k , W k are the cost function and weights of the neural network at the k-th time step, respectively; η is the learning rate.

4. The human - machine collaborative compliant control method based on improved deep reinforcement learning in combination with the collaborator's intention according to claim 1, characterized in that, When performing Step 3, use adaptive impedance control to adjust the damping parameter B. It specifically includes the following steps: Step B1: The formula of the robot impedance model is as follows: where E = X d - X, X d is the desired position of the end of the robotic arm, and X is the actual position of the end of the robotic arm; M d , B d and K d are the inertia matrix, the damping matrix, and the stiffness matrix; Step B2: Based on the feedback of the external force, adjust the damping coefficient. The adaptive impedance control model of the robotic arm is as follows: Among them, B vi is a variable damping matrix; S is an adjustment matrix; s ij is an element in the S matrix, α is a weight factor; δ ij represents 1 when i = j and 0 otherwise; sgn is a sign function; is the i-th component of the Cartesian velocity vector; is the time derivative of the i-th component of the force; Step B3: Stability analysis. The impedance models of both the collaborator and the human are linear. Represent the equation of the entire system as a second-order linear equation. The equation is as follows: Rewrite it as the state space equation as follows: where, A is the system characteristic matrix, T is the time constant of the first-order low-pass filter. When the real part of the system characteristic matrix is less than 0, the system is stable.

5. The human - machine collaborative compliant control method based on improved deep reinforcement learning in combination with the collaborator's intention according to claim 1, wherein When performing Step 4, use the DDPG algorithm combined with Gaussian process to optimize the damping parameters. It specifically includes the following steps: Step C1: Use the GP model to approximate the policy function in the DDPG algorithm. The input of the GP model is the state S t , and the output is the action A t ; Define the state S t as the current position, speed, and contact force of the robot, and the action A t as the damping parameter B; The Gaussian process model describes the relationship between states and actions through the covariance function. The formula is as follows: γ = f(X) + N(0,σ n 2 ) ~ N(0, K(X,X) + σ n 2 I) (16) A t = u = K(S t , X)(K(X, X)+σ n 2 I) -1 γ (17) where X = {S1, …, S Ne}, γ = {A1, …, A Ne}, A i = B θi is the action corresponding to s i , f(X) is a normal distribution with mean zero, N(0, σ n 2 ) is the noise of independent distribution; u is the mean, K(·) is the covariance matrix, k(·) is the element in the covariance matrix, σ f , l i,j are hyperparameters; Step C2: The hyperparameters of the GP model are optimized by k-fold cross-validation: First, the training set is randomly divided into k subsets of equal size; Second, for each combination of hyperparameters, k trainings and validations are performed. Each time during training, k - 1 subsets are used as training data, and the remaining 1 subset is used as validation data. Finally, calculate the performance metrics for each validation, and take the average performance of the k validations as the final performance evaluation of this combination of hyperparameters; Combined with Markov chain Monte Carlo, determine the possible values and their probabilities of the hyperparameters; Combine active learning and cross-validation, and according to the prediction uncertainty of the current model for different validation sets, preferentially select the validation set that can provide more information; Dynamically adjust the parameter structure in the kernel function during the optimization process; Step C3: Design the reward function in DDPG, which includes four parts: force error, speed error, energy consumption, and intention matching reward; w1(t) = w 1,i *e -βt (20) w2(t) = w 2,i *(1 - e -βt ) (21) where, w1(t) is the dynamic weight value of the force error and velocity error terms, w2(t) is the dynamic weight value of the energy consumption term, w 1,i and w 2,i are the initial weights, β is the decay rate; r i is the intention matching reward term, and the formula is as follows: where x is the actual trajectory of the robotic arm tracking the collaborator in real time, and x hd is the trajectory estimated by the radial basis neural network, ||·|| represents the Euclidean distance, γ is the maximum reward value, and σ is the bandwidth parameter.

6. The human-robot collaborative compliant control method based on improved deep reinforcement learning in combination with the intentions of collaborators according to claim 1, characterized in that When performing step 5, a design strategy switching mechanism is designed to gradually transition from the initial strategy to the intelligent strategy through a weight function; the action output π of the hybrid strategy is a weighted combination of the initial strategy action π I and the intelligent strategy action π θ as follows: π(t) = α(t)π I + (1 - α(t))π θ (23) Among them, α(t) is the weight function to achieve a smooth transition of the switch, and is designed as an exponential decay function; Set the dynamic switching conditions as follows: Among them, ε and δ are the thresholds of force error and speed error respectively.

Citation Information

Cited By

  • Collaborative robot control method based on adaptive observer and neural network

    CN120816506A