Robot optimal human-robot interaction impedance control method based on federated learning, storage medium and robot
By combining reinforcement learning and deterministic learning theories, an adaptive neural network impedance controller was designed, which solved the problem of the difficulty in selecting the optimal impedance parameters of the robot, realized online optimal adjustment and reuse of historical experience, and improved the real-time performance and compliance of human-computer interaction.
Patent Information
- Application Number
- CN202310459253.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-26
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2043-04-26
AI Technical Summary
Existing technologies struggle to achieve optimal selection of impedance parameters based on different task requirements and operator characteristics. Furthermore, robot impedance control is difficult to model accurately under nonlinear factors and environmental changes, resulting in poor real-time performance of the control scheme and high computational resource consumption.
Combining reinforcement learning and deterministic learning theories, an adaptive neural network impedance controller is designed. By constructing a task space reference regression trajectory and a second-order impedance model, the controller achieves online optimal impedance parameter adjustment and saves the neural network weights after learning convergence to reuse historical experience.
It enables online optimal selection of impedance parameters under different task scenarios and interaction object conditions, improves the real-time performance and compliance of human-computer interaction control, saves computing resources, and enhances the human-computer interaction experience.
Smart Images

Figure CN116512256B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of human-machine interaction control of robots, specifically to a robot optimal human-machine interaction impedance control method based on joint learning, a storage medium, and a robot. Background Technology
[0002] With the improvement of my country's scientific and technological level and the rapid development of industrial manufacturing capabilities in recent years, robot control technology has also been continuously improving. In addition to traditional industrial production tasks such as polishing, grinding, and assembly, robots are increasingly being used in fields such as rehabilitation medicine and human-robot collaboration, including rehabilitation robots, surgical robots, and assistive exoskeletons. Human-robot interaction systems leverage the advantages of human intelligence and autonomy while fully utilizing the advantages of robots, such as high repeatability, high precision, accurate quantitative evaluation, and suitability for extreme working environments. In the field of human-robot interaction control, the design of compliant interaction control methods is particularly crucial. Impedance control is a commonly used control method for compliant robot control. Unlike traditional force-position control, impedance control directly controls the human-robot interaction process by designing the robot's control torque, achieving better compliant interaction effects. To achieve a higher level of human-robot interaction quality, it is often necessary to rationally design the impedance parameters during the interaction process based on different task requirements and the unique motion characteristics of different operators. However, traditional impedance control techniques struggle to achieve optimal selection of impedance parameters. Furthermore, high-precision robot impedance control often requires accurate dynamic models. However, due to the robot's own nonlinear factors, component wear, and changes in working environment parameters such as temperature and humidity, accurate modeling of the robot is extremely difficult. Existing research typically uses neural networks to accurately approximate the nonlinear unknown dynamics in the robot system, but the neural network weights need to be readjusted for each task execution. This not only reduces the real-time performance of the control scheme but also consumes a significant amount of computational resources. Therefore, it is of great significance to design a high-performance human-robot interaction impedance control method that combines reinforcement learning and deterministic learning theories, enabling online optimal adjustment of human-robot interaction parameters under different task scenarios and interaction objects, while also reusing historical experience knowledge for similar human-robot interaction tasks to save computational resources and shorten adjustment time. Summary of the Invention
[0003] The main objective of this invention is to overcome the shortcomings and deficiencies of existing technologies and provide a robot optimal human-robot interaction impedance control method, storage medium, and robot based on joint learning. For the human-robot interaction compliance control problem, this invention utilizes the concept of impedance control and combines Lyapunov stability theory to propose an adaptive neural network impedance controller. Addressing the problem of online adjustment of human-robot compliance interaction parameters under conditions of unknown kinematic characteristics in different human-robot interaction task scenarios, this invention utilizes reinforcement learning theory to achieve online optimal selection of impedance parameters based on different task scenarios and interaction objects with different motion characteristics. For the unknown nonlinear dynamics in the robot model, this invention utilizes deterministic learning theory to achieve accurate fitting of the unknown nonlinear dynamic model, while saving the neural network weights after learning convergence. For similar human-robot interaction tasks, historical experience knowledge can be reused to save computational resources and shorten adjustment time.
[0004] To achieve the above objectives, the present invention adopts the following technical solution:
[0005] In a first aspect, the present invention provides a method for optimal human-machine interaction impedance control of robots based on joint learning, comprising the following steps:
[0006] S1. Constructing a task space reference regression trajectory, a second-order impedance model for human-computer interaction, and a task space auxiliary trajectory based on robot characteristics:
[0007] The second-order impedance model for human-computer interaction is as follows:
[0008]
[0009] Where t is time, M d (t) is the inertia matrix of the second-order impedance model at time t, B d (t) is the damping matrix of the second-order impedance model at time t, K d (t) is the stiffness matrix of the second-order impedance model at time t, K f (t) represents the human-computer interaction force gain at time t. For robot end-effector acceleration, Let ξ be the robot's end-effector velocity and ξ be the robot's end-effector position. For the robot's task space reference acceleration, ξ is the reference velocity for the robot's task space. d Let f be the reference position in the robot's task space, and let f be the interaction force between the robot and the human operator.
[0010] The task space auxiliary trajectory is as follows:
[0011]
[0012] Where, ξ r1ξ provides auxiliary positioning for the robot's task space. r2 To assist the robot's speed in the task space;
[0013] S2. Establish a human-computer interaction task space augmentation system and its corresponding evaluation index function, and update the parameters of the second-order impedance model online based on the integral reinforcement algorithm until the optimal parameters are obtained, as detailed below:
[0014] Design a human-computer interaction task space augmentation system and its corresponding evaluation index function:
[0015]
[0016] U = KX,
[0017]
[0018] in, To broaden the system state for human-computer interaction tasks. To assist in speeding up the task space, For task space assisted acceleration, k f1 k f2 k f3 Let U be the unknown human-computer interaction force characteristic parameters, K be the augmented system control input, V be the augmented system control gain matrix, t be the performance evaluation index function, and K be the time. q For a symmetric positive definite matrix, by designing K q Matrix elements allow for adjustments to the focus of human-computer interaction tasks, K r It is a symmetric positive definite matrix, and τ is an auxiliary time variable;
[0019] S3. For the second-order impedance model, construct an adaptive neural network impedance controller. Based on deterministic learning theory, the weights of the trained and converged neural network are... Save as constant neural network weights Specifically as follows:
[0020] The impedance error is defined as:
[0021]
[0022] Design an adaptive neural network impedance controller:
[0023]
[0024] Where e is the auxiliary impedance error variable, and the convergence of e indicates the convergence of the impedance error ε, τ f This maps the robot's joint space control torque to the control input in the task space. This is the transpose of the neural network weight estimates. Let θ be a Gaussian radial basis function.k Let ρ be the center point of the distribution, k = i, 2, ..., N. k Where N is the width and N is the number of nodes in the neural network. Where q = [q1, q2, ..., q n ] T Let q be the angular displacement of the robot in joint space. i Let be the angular displacement of the i-th joint, where i = 1, 2, ..., n, and n corresponds to the number of joints in the robot. Let be the angular velocity of the robot in joint space. Let K be the angular velocity of the i-th joint. e It is the gain matrix of the adaptive neural network controller;
[0025] Constructing neural network weight estimates The weight update law is:
[0026]
[0027] Where Γ is the gain term of the weight update law, and σ is the design constant of the weight update law;
[0028] S4. Using constant neural network weights Constructing a constant neural network impedance controller:
[0029]
[0030] Among them, K f It is the optimal human-computer interaction force gain. M d It is the inertia matrix of the second-order impedance model, B d It is the damping matrix of the optimal second-order impedance model, K d It is the stiffness matrix of the optimal second-order impedance model.
[0031] As a preferred technical solution, the robot characteristics are determined by a robot model, which is set as an n-link rigid robotic arm model, specifically including:
[0032] The robot's kinematic model is as follows:
[0033] ξ = g(q),
[0034]
[0035] Where g(·) is the mapping of the robot from the joint space angular displacement to the task space coordinates, and J is the Jacobian matrix of the robot system;
[0036] The robot joint space dynamics model is as follows:
[0037]
[0038] in, M represents the angular acceleration of the robot in joint space. q (q) represents the robot's inertia matrix in joint space. For the centripetal force matrix and G of the robot in joint space q (q) is the robot's gravity matrix in joint space, τ q For joint control torque, Let be the angular acceleration of the i-th joint, i = 1, 2, ..., n.
[0039] As a preferred technical solution, in step S1, the task space reference regression trajectory is:
[0040]
[0041] in, Given a continuous smooth function, ξ d1 =ξ d For the robot's task space reference acceleration, This serves as a reference velocity for the robot's task space.
[0042] As a preferred technical solution, in step S2, the online updating of the second-order impedance model parameters based on the integral enhancement algorithm until the optimal parameters are obtained specifically involves:
[0043] The integral enhancement algorithm selected is as follows:
[0044] Strategy Evaluation:
[0045]
[0046] Strategy Update:
[0047] K i+1 =K r -1 B T P i
[0048] Where X(t) represents the value of the task space augmented system state X at time t, P i Let K represent the solution of the algorithm at the i-th iteration, T be the sampling time, τ be the auxiliary time variable, and K be defined. i Let K be the control gain matrix of the task space augmented system at the i-th iteration. i+1 Let B be the control gain matrix of the task space augmented system at the (i+1)th iteration, where B = [0, I]. n×n 0] T Augment the system matrix in the task space;
[0049] The above reinforcement learning algorithm is computed online in real time using the least squares method:
[0050]
[0051] in, For P i Transpose of an element-wise vector Let X(t) be the Kronecker product quadratic polynomial basis vector. As an auxiliary variable, For auxiliary matrix, For auxiliary matrix, For auxiliary matrix;
[0052] Substitute the initial value K0 that stabilizes the augmented system into the algorithm, and use the least squares method to calculate the solution at each step online. P was obtained later. i Substituting this into the policy update formula yields the control gain K. i+1 When ||K i+1 -K i When ||<δ, the optimal feedback gain K is obtained. * δ is a set error constant, which is usually taken as a small value;
[0053] At time t, the control gain K(t) of the task space augmentation system is:
[0054]
[0055] Based on the above relationships, by selecting a suitable M d From the (t) matrix, we can obtain the parameters K of the second-order impedance model for real-time human-computer interaction. d (t), B d (t), K f (t), when K(t) converges to K * At that time, the optimal second-order impedance model parameter K for human-computer interaction is obtained. d B d K f .
[0056] As a preferred technical solution, in step S3, the constant neural network weight W is specifically:
[0057]
[0058] Where t2>t1>T, and T is the convergence time.
[0059] In a second aspect, the present invention provides a robot, the robot comprising:
[0060] At least one processor; and,
[0061] A memory communicatively connected to the at least one processor; wherein,
[0062] The memory stores computer program instructions that can be executed by the at least one processor, which enables the at least one processor to execute the robot optimal human-machine interaction impedance control method based on joint learning.
[0063] Thirdly, the present invention provides a computer-readable storage medium storing a program, which, when executed by a processor, implements the aforementioned method for optimal human-machine interaction impedance control based on joint learning.
[0064] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0065] 1. This invention combines reinforcement learning to achieve online optimal selection of human-computer compliant interaction parameters under the condition that the kinematic characteristics of the interaction object are unknown in different human-computer interaction task scenarios, making the human-computer interaction control system more versatile;
[0066] 2. This invention achieves accurate identification of unknown nonlinear dynamics in the robot model during human-computer interaction. At the same time, it can reuse historical experience knowledge for similar human-computer interaction tasks to save computing resources and shorten adjustment time, making robot control more real-time.
[0067] 3. This invention combines the concepts of impedance control with reinforcement learning and deterministic learning theory to realize the characteristics of the desired impedance model, thereby improving the compliant control performance of human-computer interaction and enhancing the human-computer interaction experience. Attached Figure Description
[0068] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0069] Figure 1 This is a flowchart of a robot optimal human-machine interaction impedance control method based on joint learning, according to an embodiment of the present invention.
[0070] Figure 2 This is a schematic diagram of a dual-link robot according to an embodiment of the present invention.
[0071] Figure 3 This is a convergence diagram of the norm of the control gain matrix of the task space augmentation system in this embodiment.
[0072] Figure 4This is a diagram showing the impedance error curve of the human-machine interaction auxiliary robot system during the adaptive control phase of an embodiment of the present invention.
[0073] Figure 5 This is a convergence curve of the neural network weight norm of the robot system in the adaptive control stage of this invention.
[0074] Figure 6 This is a diagram illustrating the unknown dynamic effect of the neural network fitting system model of the robot system during the adaptive control stage of this invention.
[0075] Figure 7 This is a graph showing the change of control input signal in the task space of the robot system during the adaptive control phase of an embodiment of the present invention.
[0076] Figure 8 This is a force curve diagram of the interaction between the robot end effector and the operator in an embodiment of the present invention.
[0077] Figure 9 This is a graph showing the trajectory of the robot's end effector in this invention.
[0078] Figure 10 This is a graph showing the auxiliary impedance error variables of the robot system during the learning control phase in an embodiment of the present invention.
[0079] Figure 11 This is a schematic diagram of the robot according to an embodiment of the present invention. Detailed Implementation
[0080] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort are within the scope of protection of the present application.
[0081] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.
[0082] like Figure 1 As shown, this embodiment provides a robot optimal human-robot interaction impedance control method based on joint learning, selecting a two-link rigid robot as the model, and includes the following steps:
[0083] S1. Based on the robot's characteristics, establish corresponding kinematic and task space dynamic models, and construct task space reference regression trajectory, human-robot interaction second-order impedance model, and task space auxiliary trajectory:
[0084] Based on the characteristics of the robot, the mapping from joint space to task space of the robot is obtained as follows:
[0085] ξ=g(q)
[0086]
[0087] Where ξ is the position of the robot's end effector. Let be the robot's end effector velocity, g(·) be the mapping from the robot's angular displacement in joint space to its coordinates in task space, J be the Jacobian matrix of the robot system, and q = [q1, q2]. T Let q be the angular displacement of the robot in joint space. i (i = 1, 2) represents the angular displacement of the i-th joint. Let be the angular velocity of the robot in joint space. Let be the angular velocity of the i-th joint.
[0088] Based on the robot's structure, this embodiment selects the following forward kinematics model:
[0089]
[0090] Where x, y, and z represent the positions in the three directions of the task space, and l1 = 1m and l2 = 1m represent the lengths of link 1 and link 2, respectively. Figure 2 As shown.
[0091] The Jacobian matrix of the robot system is:
[0092]
[0093] The dynamic model of the robot in joint space is as follows:
[0094]
[0095] in, M represents the angular acceleration of the robot in joint space. q (q) represents the robot's inertia matrix in joint space. For the centripetal force matrix and G of the robot in joint space q (q) is the robot's gravity matrix in joint space, τ q For joint control torque, Let K be the angular acceleration of the i-th joint, f be the interaction force between the robot and the human operator, measured in real time by a force sensor, and K be the angular acceleration of the i-th joint. f (t) represents the human-computer interaction force gain at time t.
[0096] The robot dynamics model based on the task space is as follows:
[0097]
[0098] Where, τ f M(q) represents the control torque mapped from the joint torque to the robot's end effector, and M(q) represents the inertia matrix in the task space. Let G(q) be the centripetal force matrix in the task space, and G(q) be the gravity term in the task space. The conversion relationship between the robot's end-effector acceleration and the dynamic model parameters in joint space is as follows:
[0099] M(q)=J -T M q (q)J -1 ,
[0100]
[0101] G(q)=J -T G q (q).
[0102] The relevant parameters of the selected double-link rigid robot model in this embodiment are as follows:
[0103]
[0104] In this embodiment, m1 = 3kg and m2 = 3kg are selected as the masses of connecting rod 1 and connecting rod 2, respectively, and g = 9.8m / s². 2 This is the acceleration due to gravity.
[0105] The design task space reference regression trajectory is:
[0106]
[0107] in, Given a continuous smooth function, ξ d2 ξ is the reference velocity for the robot's task space. d1 This serves as the robot's task space reference position. In this embodiment, the selected task space reference trajectory is:
[0108] ξ d = [1 + 0.2sin(t), 1 - 0.2cos(t)] T
[0109] Design a second-order impedance model for human-computer interaction:
[0110]
[0111] Where t is time, Md (t) is the inertia matrix of the second-order impedance model at time t, B d (t) is the damping matrix of the second-order impedance model at time t, K d (t) is the stiffness matrix of the second-order impedance model at time t, K f (t) represents the human-computer interaction force gain at time t. For robot end-effector acceleration, Let ξ be the robot's end-effector velocity and ξ be the robot's end-effector position. For the robot's task space reference acceleration, ξ is the reference velocity for the robot's task space. d This serves as the reference position for the robot's task space. In this embodiment, M is selected. d (t) is a constant matrix.
[0112] Design the task space auxiliary trajectory:
[0113]
[0114] Where, ξ r1 ξ provides auxiliary positioning for the robot's task space. r2 To assist the robot's mission space speed.
[0115] S2. Establish a human-computer interaction task space augmentation system and its corresponding evaluation index function, and update the parameters of the second-order impedance model online based on the integral reinforcement algorithm until the optimal parameters are obtained:
[0116] Design a human-computer interaction task space augmentation system and its corresponding evaluation index function:
[0117]
[0118] U = KX
[0119]
[0120] in, To augment the system state in the human-computer interaction task space, ξ r To assist in positioning the robot within its task space. To assist in speeding up the task space, For task space assisted acceleration, k f1 k f2 k f3 Let U be the unknown human-computer interaction force characteristic parameters, K be the augmented system control input, V be the augmented system control gain matrix, t be the performance evaluation index function, and K be the time. q For a symmetric positive definite matrix, by designing K q Matrix elements allow for adjustments to the focus of human-computer interaction tasks, Kr Let be a symmetric positive definite matrix, and τ be an auxiliary time variable. In this embodiment, is selected as...
[0121] Based on the integral enhancement algorithm, the optimal control problem of the task space augmented system is solved:
[0122] The integral enhancement algorithm selected is as follows:
[0123] Strategy Evaluation:
[0124]
[0125] Strategy Update:
[0126] K i+1 =K r -1 B T P i
[0127] Where X(t) represents the value of the task space augmented system state X at time t, P i Let K represent the solution of the algorithm at the i-th iteration, T be the sampling time, τ be the auxiliary time variable, and K be defined. i Let K be the control gain matrix of the task space augmented system at the i-th iteration. i+1 Let B be the control gain matrix of the task space augmented system at the (i+1)th iteration, where B = [0, I]. n×n 0] T Augment the system matrix for the task space.
[0128] The above reinforcement learning algorithm is computed online in real time using the least squares method:
[0129]
[0130] in, For P i Transpose of an element-wise vector Let X(t) be the Kronecker product quadratic polynomial basis vector. As an auxiliary variable, For auxiliary matrix, For auxiliary matrix, Let N be the auxiliary matrix and N be the number of data samples. Substitute the initial value K0, which stabilizes the augmented system, into the algorithm, and use the least squares method to calculate the solution at each step online. P was obtained later. i Substituting this into the policy update formula yields the control gain K. i+1 When ||K i+1 -K i When ||<δ, the optimal feedback gain K is obtained. *δ is a set error constant, which is usually taken as a small value.
[0131] At time t, the control gain K(t) of the task space augmentation system is:
[0132]
[0133] Based on the above relationships, by selecting a suitable M d From the (t) matrix, we can obtain the parameters K of the second-order impedance model for real-time human-computer interaction. d (t), B d (t), K f (t), when K(t) converges to K * At that time, the optimal second-order impedance model parameter K for human-computer interaction can be obtained. d B d K f In this embodiment, a sampling time of T = 0.05 s is selected. δ = 0.1, N = 6.
[0134] S3. For the second-order impedance model, construct an adaptive neural network impedance controller. Based on deterministic learning theory, the weights of the trained and converged neural network are... Save as constant neural network weights
[0135] The impedance error is defined as:
[0136]
[0137] Design an adaptive neural network impedance controller:
[0138]
[0139] Where e is the auxiliary impedance error variable, and the convergence of e indicates the convergence of the impedance error ε, τ f This maps the robot's joint space control torque to the control input in the task space. This is the transpose of the neural network weight estimates. Let θ be the Gaussian radial basis function, M be the number of points in the neural network, and θ be the number of points in the network. k (k = i, 2, ..., M) is the center of the distribution point, ρ k (k = i, 2, ..., M) represents the neuron width. Among them, K e ξ is the gain matrix of the adaptive neural network controller. In this embodiment, ξ and The initial value of ξ is [0.8, 1]. T and Neural network weights initial value The neural network points are centered at [0.3, 0.3, 0.4, 0.3, 0.3, 0.4, 0.4, 0.4, 0.4, 0, 0]. T The width of the neural network neurons is [0.375, 0.375, 0.5, 0.375, 0.375, 0.5, 0.5, 0.5, 0.5, 0, 0]. T Adaptive neural network controller gain
[0140] Constructing neural network weight estimates The weight update law is:
[0141]
[0142] Where Γ is the gain term of the weight update law, and σ is the design constant of the weight update law. In this embodiment,
[0143] σ = 0.00001.
[0144] Using deterministic learning theory to calculate the weights of the converged neural network Save as constant weights Specifically:
[0145]
[0146] Where T < t1 < t2, and T is the convergence time. In this embodiment, T = 100s, t1 = 180s, and t2 = 200s.
[0147] S4. Using constant neural network weights Constructing a constant neural network impedance controller:
[0148]
[0149] in,
[0150] In this embodiment, the initial values and parameter settings of each state in the learning control phase are the same as those in the adaptive control phase.
[0151] Using the parameters in this embodiment, the following results can be obtained:
[0152] Figure 3 The convergence plot of the norm of the control gain matrix of the task space augmented system is shown in the figure. As can be seen from the figure, the norm of the actual gain matrix converges to the vicinity of the ideal optimal gain matrix after 4 iterations, with a time of 0.349s and a norm error of 0.16. This proves that the reinforcement learning algorithm can converge to obtain the optimal human-computer interaction impedance parameters in a short time. Figure 4The figure shows the auxiliary impedance error curve of the robot system during the adaptive control phase. It can be seen that after 100s, the auxiliary impedance error basically converges to near zero. This indicates that the adaptive neural network controller can basically achieve compliant human-machine interaction control, but its transient control performance is generally poor. Figure 5 This is a convergence curve of the neural network weight norm of the robot system during the adaptive control phase. Figure 6 The diagram shows the effect of the neural network fitting the unknown dynamics of the robot system model during the adaptive control stage. It can be seen that the neural network weights basically converged after 100s and achieved a good approximation of the unknown nonlinear dynamics inside the system. Figure 7 The curve of the control input signal variation in the task space of the robot system during the adaptive control phase shows that the control input signal is smooth and continuous with short transient vibration processes, which can ensure the stable and safe operation of the system. Figure 8 This is a force curve diagram showing the interaction between the robot's end effector and the operator. Figure 9 The figure shows the trajectory curve of the robot's end effector. As can be seen from the figure, during the human-machine interaction between the robot and the operator, the robot maintains good compliance characteristics and gradually converges to the reference trajectory as the human-machine interaction force decreases. Figure 10 The graph shows the impedance error variable curve of the robot system during the learning control phase. As can be seen from the graph, learning control greatly shortens the system adjustment time, improves control performance, saves computing resources, and achieves high-precision, compliant human-machine interaction control.
[0153] It should be noted that, for the sake of simplicity, the aforementioned method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously.
[0154] Please see Figure 11 In one embodiment, a robot based on a joint learning-based optimal human-machine interaction impedance control method is provided. The robot 100 may include a first processor 101, a first memory 102 and a bus, and may also include a computer program stored in the first memory 102 and executable on the first processor 101, such as a robot optimal human-machine interaction impedance control program 103.
[0155] The first memory 102 includes at least one type of readable storage medium, including flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the first memory 102 can be an internal storage unit of the robot 100, such as the robot 100's portable hard drive. In other embodiments, the first memory 102 can also be an external storage device of the robot 100, such as a plug-in portable hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the robot 100. Furthermore, the first memory 102 can include both internal storage units and external storage devices of the robot 100. The first memory 102 can be used not only to store application software and various types of data installed on the robot 100, such as the code of the robot's optimal human-machine interaction impedance control program 103, but also to temporarily store data that has been output or will be output.
[0156] In some embodiments, the first processor 101 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The first processor 101 is the control unit of the robot, connecting various components of the robot through various interfaces and lines. It executes programs or modules stored in the first memory 102 and calls data stored in the first memory 102 to perform various functions of the robot 100 and process data.
[0157] Figure 3 Only a robot with parts is shown; those skilled in the art will understand that... Figure 3 The structure shown does not constitute a limitation on the robot 100, and may include fewer or more parts than shown, or combine certain parts, or have different arrangements of parts.
[0158] The robot optimal human-machine interaction impedance control program 103 stored in the first memory 102 of the robot 100 is a combination of multiple instructions. When run in the first processor 101, it can achieve the following:
[0159] S1. Constructing a task space reference regression trajectory, a second-order impedance model for human-computer interaction, and a task space auxiliary trajectory based on robot characteristics:
[0160] The second-order impedance model for human-computer interaction is as follows:
[0161]
[0162] Where t is time, M d (t) is the inertia matrix of the second-order impedance model at time t, B d (t) is the damping matrix of the second-order impedance model at time t, K d (t) is the stiffness matrix of the second-order impedance model at time t, K f (t) represents the human-computer interaction force gain at time t. For robot end-effector acceleration, Let ξ be the robot's end-effector velocity and ξ be the robot's end-effector position. For the robot's task space reference acceleration, ξ is the reference velocity for the robot's task space. d Let f be the reference position in the robot's task space, and let f be the interaction force between the robot and the human operator.
[0163] The task space auxiliary trajectory is as follows:
[0164]
[0165] Where, ξ r1 ξ provides auxiliary positioning for the robot's task space. r2 To assist the robot's speed in the task space;
[0166] S2. Establish a human-computer interaction task space augmentation system and its corresponding evaluation index function, and update the parameters of the second-order impedance model online based on the integral reinforcement algorithm until the optimal parameters are obtained, as detailed below:
[0167] Design a human-computer interaction task space augmentation system and its corresponding evaluation index function:
[0168]
[0169] U = KX,
[0170]
[0171] in, To broaden the system state for human-computer interaction tasks. To assist in speeding up the task space, For task space assisted acceleration, k f1 k f2 k f3Let U be the unknown human-computer interaction force characteristic parameters, K be the augmented system control input, V be the augmented system control gain matrix, t be the performance evaluation index function, and K be the time. q For a symmetric positive definite matrix, by designing K q Matrix elements allow for adjustments to the focus of human-computer interaction tasks, K r It is a symmetric positive definite matrix, and τ is an auxiliary time variable;
[0172] S3. For the second-order impedance model, construct an adaptive neural network impedance controller. Based on deterministic learning theory, the weights of the trained and converged neural network are... Save as constant neural network weights Specifically as follows:
[0173] The impedance error is defined as:
[0174]
[0175] Design an adaptive neural network impedance controller:
[0176]
[0177] Where e is the auxiliary impedance error variable, and the convergence of e indicates the convergence of the impedance error ε, τ f This maps the robot's joint space control torque to the control input in the task space. This is the transpose of the neural network weight estimates. Let θ be a Gaussian radial basis function. k (k = i, 2, ..., N) is the center point of the distribution, ρ k (k = i, 2, ..., N) represents the width, and N represents the number of nodes in the neural network. Where q = [q1, q2, ..., q n ] T Let q be the angular displacement of the robot in joint space. i (i = 1, 2, ..., n) represents the angular displacement of the i-th joint, and n corresponds to the number of joints in the robot. Let be the angular velocity of the robot in joint space. Let K be the angular velocity of the i-th joint. e It is the gain matrix of the adaptive neural network controller;
[0178] Constructing neural network weight estimates The weight update law is:
[0179]
[0180] Where Γ is the gain term of the weight update law, and σ is the design constant of the weight update law;
[0181] S4. Using constant neural network weights Constructing a constant neural network impedance controller:
[0182]
[0183] Among them, K f It is the optimal human-computer interaction force gain. M d It is the inertia matrix of the second-order impedance model, B d It is the damping matrix of the optimal second-order impedance model, K d It is the stiffness matrix of the optimal second-order impedance model.
[0184] Furthermore, if the modules / units integrated into the robot 100 are implemented as software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).
[0185] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0186] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0187] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.
Claims
1. A method for optimal human-machine interaction impedance control of robots based on joint learning, characterized in that, Includes the following steps: S1. Constructing a task space reference regression trajectory, a second-order impedance model for human-computer interaction, and a task space auxiliary trajectory based on robot characteristics: The second-order impedance model for human-computer interaction is as follows: Where t is time, M d (t) is the inertia matrix of the second-order impedance model at time t, B d (t) is the damping matrix of the second-order impedance model at time t, K d (t) is the stiffness matrix of the second-order impedance model at time t, K f (t) represents the human-computer interaction force gain at time t. For robot end-effector acceleration, Let ξ be the robot's end-effector velocity and ξ be the robot's end-effector position. For the robot's task space reference acceleration, ξ is the reference velocity for the robot's task space. d Let f be the reference position in the robot's task space, and let f be the interaction force between the robot and the human operator. The task space auxiliary trajectory is as follows: Where, ξ r1 ξ provides auxiliary positioning for the robot's task space. r2 To assist the robot's task space speed; S2. Establish a human-computer interaction task space augmentation system and its corresponding evaluation index function, and update the parameters of the second-order impedance model online based on the integral reinforcement algorithm until the optimal parameters are obtained, as detailed below: Design a human-computer interaction task space augmentation system and its corresponding evaluation index function: U = KX, V=∫ t ∞ ( XK q X+KXK r KX)dτ, in, To broaden the system state for human-computer interaction tasks To assist in speeding up the task space, For task space auxiliary acceleration, k f1 k f2 k f3 Let U be the unknown human-computer interaction force characteristic parameters, K be the augmented system control input, V be the augmented system control gain matrix, t be the performance evaluation index function, and K be the time. q For a symmetric positive definite matrix, by designing K q Matrix elements allow for adjustments to the focus of human-computer interaction tasks, K r It is a symmetric positive definite matrix, and τ is an auxiliary time variable; S3. For the second-order impedance model, construct an adaptive neural network impedance controller. Based on deterministic learning theory, the weights of the trained and converged neural network are... Save as constant neural network weights Specifically as follows: The impedance error is defined as: Design an adaptive neural network impedance controller: Where e is the auxiliary impedance error variable, and the convergence of e indicates the convergence of the impedance error ε, τ f This maps the robot's joint space control torque to the control input in the task space. This is the transpose of the neural network weight estimates. Let θ be a Gaussian radial basis function. k Let ρ be the center point of the distribution, k = i, 2, ..., N. k Where N is the width and N is the number of nodes in the neural network. Where q = [q1, q2, ..., q n ] T Let q be the angular displacement of the robot in joint space. i Let be the angular displacement of the i-th joint, where i = 1, 2, ..., n, and n corresponds to the number of joints in the robot. Let be the angular velocity of the robot in joint space. Let K be the angular velocity of the i-th joint. e It is the gain matrix of the adaptive neural network controller; Constructing neural network weight estimates The weight update law is: Where Γ is the gain term of the weight update law, and σ is the design constant of the weight update law; S4. Using constant neural network weights Constructing a constant neural network impedance controller: Among them, K f It is the optimal human-computer interaction force gain. M d It is the inertia matrix of the second-order impedance model, B d It is the damping matrix of the optimal second-order impedance model, K d It is the stiffness matrix of the optimal second-order impedance model.
2. The robot optimal human-machine interaction impedance control method based on joint learning according to claim 1, characterized in that, The robot characteristics are determined by the robot model, which is set as an n-link rigid robotic arm model, specifically including: The robot's kinematic model is as follows: ξ = g(q), Where g(·) is the mapping of the robot from the joint space angular displacement to the task space coordinates, and J is the Jacobian matrix of the robot system; The robot joint space dynamics model is as follows: in, M represents the angular acceleration of the robot in joint space. q (q) represents the robot's inertia matrix in joint space. For the centripetal force matrix and G of the robot in joint space q (q) is the robot's gravity matrix in joint space, τ q For joint control torque, Let be the angular acceleration of the i-th joint, i = 1, 2, ..., n.
3. The robot optimal human-machine interaction impedance control method based on joint learning according to claim 1, characterized in that, In step S1, the task space reference regression trajectory is: in, Given a continuous smooth function, ξ d1 =ξ d For the robot's task space reference acceleration, This serves as a reference velocity for the robot's task space.
4. The robot optimal human-machine interaction impedance control method based on joint learning according to claim 1, characterized in that, In step S2, the second-order impedance model parameters are updated online based on the integral enhancement algorithm until the optimal parameters are obtained, specifically as follows: The integral enhancement algorithm selected is as follows: Strategy Evaluation: Strategy Update: K i+1 =K r -1 B T P i Where X(t) represents the value of the task space augmented system state X at time t, P i Let K represent the solution of the algorithm at the i-th iteration, T be the sampling time, τ be the auxiliary time variable, and K be defined. i Let K be the control gain matrix of the task space augmented system at the i-th iteration. i+1 Let B be the control gain matrix of the task space augmented system at the (i+1)th iteration, where B = [0, I]. n×n 0] T Augment the system matrix in the task space; The above-mentioned integral enhancement algorithm is calculated online in real time using the least squares method: in, For P i Transpose of an element-wise vector Let X(t) be the Kronecker product quadratic polynomial basis vector. As an auxiliary variable, For auxiliary matrix, For auxiliary matrix, For auxiliary matrix; Substitute the initial value K0 that stabilizes the augmented system into the algorithm, and use the least squares method to calculate the solution at each step online. P was obtained later. i Substituting this into the policy update formula yields the control gain K. i+1 When ||K i+1 -K i When ||<δ, the optimal feedback gain K is obtained. * δ is the set error constant; At time t, the control gain K(t) of the task space augmentation system is: Based on the above relationships, by selecting a suitable M d From the (t) matrix, we can obtain the parameters K of the second-order impedance model for real-time human-computer interaction. d (t), B d (t), K f (t), when K(t) converges to K * At that time, the optimal second-order impedance model parameter K for human-computer interaction is obtained. d B d K f .
5. The robot optimal human-machine interaction impedance control method based on joint learning according to claim 1, characterized in that, In step S3, the constant neural network weights Specifically: Where t2>t1>T, and T is the convergence time.
6. A robot, characterized in that, The robot includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores computer program instructions that can be executed by the at least one processor, which enables the at least one processor to perform the robot optimal human-machine interaction impedance control method based on joint learning as described in any one of claims 1-5.
7. A computer-readable storage medium storing a program, characterized in that, When the program is executed by the processor, it implements the robot optimal human-machine interaction impedance control method based on joint learning as described in any one of claims 1-5.
Citation Information
Patent Citations
Dual-module cooperative robot coordinated assembly system for 3C assembly and planning method
CN111522305A
Variable impedance control system and control method based on inverse reinforcement learning
CN115421387A