Robot-environment optimal interaction control method based on inverse reinforcement learning

Through the inverse reinforcement learning method, experts demonstrate data to recover the interactive performance function of unknown environments, design the optimal impedance control strategy, and solve the problem of optimal, safe and flexible control in the interaction between robots and environments, and is suitable for robot contact operation tasks in unstructured environments.

CN120406140APending Publication Date: 2025-08-01BEIHANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510531705.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The prior art cannot effectively achieve optimal, safe and flexible control when robots interact with the environment in unknown environments. Traditional methods rely on manual tuning and cannot guarantee interaction performance.

Method used

Using an inverse reinforcement learning method, the interaction performance function of unknown environments is restored using measurable expert demonstration data, and the optimal impedance control strategy is designed without the need for environmental dynamic model information. The optimal interaction between the robot and the environment is achieved through the inverse reinforcement learning algorithm.

Benefits of technology

It realizes the optimal, safe and flexible interaction of robots in unknown environments according to expert intentions, overcomes the limitations of traditional methods, and is suitable for contact operation tasks in unstructured environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120406140A_ABST
    Figure CN120406140A_ABST
Patent Text Reader

Abstract

The invention relates to a robot-environment optimal interaction control method based on inverse reinforcement learning, and the method comprises the steps: firstly building an interaction control model of a robot and an environment, and describing an environment position and an expected track through a linear expression; secondly, establishing an actual controlled system model based on a state space and giving an optimal impedance control strategy based on the model; then, constructing an expert demonstration system to generate expert demonstration data, and designing an estimation algorithm of expert state data and expert output feedback control gain; and finally, designing a robot optimal impedance control learning algorithm based on inverse reinforcement learning. Aiming at an interaction control task scene of a robot and an unknown environment, the designed method can learn an optimal expert control strategy and an unknown value function only by utilizing expert demonstration data, and the optimal impedance control problems of unknown contact environment, unknown environment position, unknown expected trajectory dynamics and unspecified interaction performance can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of robot intelligent compliant control, and particularly relates to a robot-environment optimal interaction control method based on inverse reinforcement learning. Background Art

[0002] Robots are an important symbol to measure a country's innovation ability and industrial competitiveness, and are known as the "pearl on the top of manufacturing". They are widely used in industrial automation, aerospace, home service, medical assistance and other fields. Performing complex contact operation tasks in a dynamic and unstructured environment is a core application scenario of robots, such as grinding and polishing, hole making, shaft-hole assembly, human-robot collaboration, etc. These scenarios usually face problems such as task changes, unknown environment models, and unmeasurable environment positions, which may cause large interaction forces at the moment of contact between the robot and the environment, and even damage the robot system. Therefore, controlling the dynamic interaction process between the robot and the environment and endowing the robot with the compliant interaction ability with the environment is a key technology to improve the safety and reliability of the robot system.

[0003] Impedance control is a classical robot-environment interaction control algorithm, which can adjust the amplitude of the interaction force by dynamically correcting the motion trajectory of the robot end. Traditional methods mostly adopt impedance control strategies with constant parameters and cannot handle scenarios such as unknown environment and task changes. Impedance control schemes based on iterative learning, environmental parameter estimation and adaptive control (CN119434368A, CN110202574A, CN115741668A) can effectively solve this problem and effectively improve the adaptability of the robot to dynamic and unstructured environments. However, these methods rely heavily on manual tuning or repeated operations and cannot guarantee the best interaction performance. Reinforcement learning technology can simultaneously ensure the adaptive, optimal and compliant interaction between the robot and the unknown environment, and is widely used in the field of robot compliant control (CN114851193A). However, this method requires pre-setting an interaction performance function, and in most cases, it cannot be directly quantified. Generally speaking, it is easier to demonstrate the interaction process by experts than to directly give the interaction performance function. How to learn the interaction performance function from the expert demonstration data and learn the optimal impedance control strategy is an urgent and challenging problem to be solved. Summary of the Invention

[0004] For the interactive control scenario between a robot and the environment, considering problems such as unknown environment models, unmeasurable environment positions, and unknown interaction performance functions, the present invention provides a robot-environment optimal interactive control method based on inverse reinforcement learning, which can recover unknown value functions and solve the optimal impedance control strategy only by using measurable expert demonstration data, without any information about the known environment dynamics, reference trajectory dynamics, and environment position, providing theoretical and technical support for the optimal interactive control between a robot and the environment in an unknown complex environment. This learning method can be executed in a model-free manner, and only by using measurable expert demonstration data, the expert interaction performance function and the expert impedance control strategy can be recovered, realizing the optimal, safe, and compliant interaction between the robot and the unknown environment according to the expert's expected intention, and can be used for the robot contact operation task in an unstructured environment.

[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] A robot-environment optimal interactive control method based on inverse reinforcement learning, comprising the following steps:

[0007] The first step is to establish the dynamic models of the interactive environment, impedance control, environment position, and expected trajectory;

[0008] The second step is to combine the established dynamic models into a state space form as the actual controlled system, and design the optimal impedance controller of the actual controlled system;

[0009] The third step is to apply the expert impedance control strategy to the actual controlled system to generate expert demonstration data, and design the expert state reconstruction and expert control strategy estimation algorithms;

[0010] The fourth step is to use the expert demonstration data to propose a robot optimal impedance control algorithm based on inverse reinforcement learning.

[0011] A computing device, comprising: at least one processor and a memory storing program instructions; when the program instructions are read and executed by the processor, the computing device is made to execute the robot-environment optimal interactive control method based on inverse reinforcement learning.

[0012] A readable storage medium storing program instructions, when the program instructions are read and executed by a computing device, the computing device is made to execute the robot-environment optimal interactive control method based on inverse reinforcement learning.

[0013] The beneficial effects of the present invention compared with the prior art are as follows:

[0014] The present invention takes into account problems such as unknown environmental models, unmeasurable environmental positions, unknown reference trajectory dynamics, and unknown interaction performance functions, and proposes a robot-environment optimal interaction control method based on inverse reinforcement learning. This method does not require any dynamic model information, and can recover the expert performance function only by using measurable expert demonstration data. It can be applied to different interaction tasks to learn optimal impedance parameters, overcoming the limitation of traditional reinforcement learning algorithms that require manual adjustment of value function parameters. It can be used in the scenario of robot contact operation tasks, and can ensure that the robot and the unknown environment perform optimal, safe, and compliant interactions according to the expert's intention. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a flow block diagram of a robot-environment optimal interaction control method based on inverse reinforcement learning;

[0016] Figure 2 It is a structural block diagram of the inverse reinforcement learning control algorithm proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.

[0018] Figure 1 It is a flow block diagram of a robot-environment optimal interaction control method based on inverse reinforcement learning. As Figure 1 shown, the method includes the following steps:

[0019] First, establish the dynamic models of the interaction environment, impedance control, environmental position, and desired trajectory;

[0020] The interaction environment model between the robot and the target is established as:

[0021] ,

[0022] where , , represent the unknown environmental inertia coefficient, damping coefficient, and stiffness coefficient, , , represent the position, velocity, and acceleration vectors of the robot end-effector respectively, is the interaction force between the environment and the robot, represents the unknown environmental position vector, which can be generated by the following environmental position generator:

[0023] ,

[0024] where and are the auxiliary state vectors of the environmental position generator, is the system matrix of the environmental position generator, is the output matrix of the environmental position generator;

[0025] The impedance control model adopted is:

[0026] ,

[0027] where represents the desired inertia coefficient, represents the desired damping coefficient, represents the desired stiffness coefficient, and represent the desired position and velocity vectors respectively, can be generated by the following desired trajectory generator:

[0028] ,

[0029] where and are the auxiliary state vectors of the desired trajectory generator, and are the system matrix and output matrix of the desired trajectory generator respectively;

[0030] In one embodiment, considering the interaction behavior in a single direction, the initial model parameters are set as follows: environmental parameters , , , impedance model parameters , environmental position generator , , desired trajectory generator , ;

[0031] Second step, merge the established dynamic model into the state - space form as the actual controlled system, and design the optimal impedance controller for the actual controlled system;

[0032] By constructing the augmented state vector , augmented output vector , then the following state - space model of the actual controlled system can be obtained:

[0033] ,

[0034] where represents the non-inertial interaction force, which is also the control strategy to be obtained, , , represent the system matrix, input matrix, and output matrix of the actual controlled system, represents the total inertia coefficient of the environment and impedance, and represent the zero matrix and identity matrix of the corresponding dimensions respectively. The superscript "T" represents the transpose of a matrix / vector, and the superscript "-1" represents the inverse of a matrix;

[0035] To quantitatively evaluate the robot-environment interaction performance, the value function is defined as:

[0036] ,

[0037] where and represent the weight coefficients of the state and input, represents the discount factor, is the kernel matrix to be solved, represents the current time, represents the integration time, is the natural exponential constant;

[0038] The expression of the optimal impedance control strategy is:

[0039] ,

[0040] where is the optimal impedance control gain to be solved, where the solution formula of is:

[0041] .

[0042] In the third step, apply the expert impedance control strategy to the actual controlled system to collect expert demonstration data, and design an expert state reconstruction and expert control strategy estimation algorithm;

[0043] Suppose there exists an optimal expert impedance control strategy applied to the actual controlled system , which generates the expert state and the expert output , then the system is written as:

[0044] ,

[0045] where the system is called the expert demonstration system; the optimal expert impedance control strategy and the expert output are collectively referred to as expert demonstration data, and the following expert state reconstruction method is introduced:

[0046] ,

[0047] where represents the state reconstruction matrix, is the collected expert output data set, represents the sampling period, is the dimension of the data set, and its value needs to ensure that the matrix is column full rank; based on this, the expert impedance control strategy can be expressed as:

[0048] ,

[0049] where is the expert impedance control gain based on state feedback, is the expert impedance control gain based on output feedback.

[0050] Estimate the expert impedance control gain by means of expert demonstration data , and the solution method is as follows:

[0051] ,

[0052] where is the collected expert input data set, is the collected expert output data set, represents the dimension of the data set. To ensure the accuracy of the solution, should be greater than in dimension;

[0053] In one embodiment, the sampling period , the dimension of the data set used to reconstruct the state is set to , and the dimension of the data set used to restore the expert output control strategy is set to .

[0054] Fourthly, using the expert demonstration data, propose an optimal impedance control algorithm for the robot based on inverse reinforcement learning;

[0055] To avoid using model parameters and known value function parameters, an optimal interaction control algorithm based on inverse reinforcement learning is proposed, and the optimal impedance control strategy can be obtained only by using the available expert demonstration data. The algorithm framework is as Figure 2 shown, and the algorithm flow is as follows:

[0056] a) Given the weight and the discount factor , set the iteration variable , , the initial value of the iteration weight is , the initial value of the inverse Hessian matrix is , the estimated value of the expert impedance control gain is ;

[0057] b) Optimal control iterative solution based on integral reinforcement learning:

[0058] b.1) Based on the expert demonstration data, use the following Bellman equation to iteratively solve the kernel matrix in the output feedback control and the control gain based on output feedback :

[0059] ,

[0060] where is the kernel matrix in the output feedback control, is the weight matrix of the output data set, represents the impedance control gain based on output feedback, the subscript "i" represents the iteration number of integral reinforcement learning, and the superscript "j" represents the iteration number of inverse reinforcement learning;

[0061] b.2) Check whether the condition is satisfied, where is a constant; if the condition is satisfied, let , , , and go to the next step; otherwise, let and go to b.1);

[0062] c) Improve the solution using the quasi-Newton algorithm and :

[0063] c.1) Calculate the control gain error scalar and the error vector :

[0064] ,

[0065] where is the operation operator that stacks an arbitrary matrix " " into a column vector by column;

[0066] c.2) Calculate the gradient vector , where the auxiliary variable ;

[0067] c.3) The search direction is ;

[0068] c.4) The vector form of the improved solution is , where is the search step size, is the kernel matrix in vector expression;

[0069] c.5) Converting to matrix form gives the improved solution of , and calculate the improved output feedback control gain ;

[0070] c.6) Update using the quasi - Newton algorithm formula;

[0071] d) Based on the expert demonstration data, update the weights and using the improved solutions :

[0072] ;

[0073] e) Check whether the conditions and are both satisfied, and are both constants; if the conditions are satisfied, stop the iteration and return the output - feedback - based impedance matrix , the kernel matrix , and the weights ; otherwise, let and repeat steps b) - e).

[0074] In one embodiment, and are set as fixed parameters, i.e., , , the search compensation is calculated in real - time using the Wolfe criterion, and the convergence threshold is set to , , The principles and execution processes of the quasi - Newton algorithm and the Wolfe criterion can be obtained from the following literature: N. Jorge and J. W. Stephen, Numerical optimization. Spinger, 2006 (N. Jorge, J.W. Stephen, Numerical Optimization, Springer, 2006).

[0075] According to an embodiment of the present invention, there is also provided a computing device, including: at least one processor and a memory storing program instructions; when the program instructions are read and executed by the processor, the computing device is caused to execute a robot-environment optimal interaction control method based on inverse reinforcement learning.

[0076] There is also provided a readable storage medium storing program instructions, when the program instructions are read and executed by a computing device, the computing device is caused to execute a robot-environment optimal interaction control method based on inverse reinforcement learning.

[0077] The content not detailedly described in the specification of the present invention belongs to the prior art well-known to those skilled in the art. It is easy for those skilled in the art to understand that the above are only preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A robot-environment optimal interaction control method based on inverse reinforcement learning, characterized in that It includes the following steps: In the first step, establish the dynamic models of the interaction environment, impedance control, environmental position, and desired trajectory; In the second step, merge the established dynamic models into the state-space form as the actual controlled system, and design the optimal impedance controller for the actual controlled system; In the third step, apply the expert impedance control strategy to the actual controlled system to collect expert demonstration data, and design the expert state reconstruction and expert control strategy estimation algorithms; In the fourth step, use the expert demonstration data to propose a robot optimal impedance control algorithm based on inverse reinforcement learning.

2. The robot-environment optimal interaction control method based on inverse reinforcement learning according to claim 1, characterized in that: The first step includes: The interaction environment model between the robot and the target is established as: , Among them , , represent the unknown environmental inertia coefficient, damping coefficient, and stiffness coefficient, , , respectively represent the position, velocity, and acceleration vectors of the robot end, is the interaction force between the environment and the robot, represents the unknown environmental position vector, which is generated by the following environmental position generator: , Among them and is the auxiliary state vector of the environmental position generator, is the system matrix of the environmental position generator, is the output matrix of the environmental position generator; The impedance control model adopted is: , wherein represents the desired inertia coefficient represents the desired damping coefficient represents the desired stiffness coefficient and represent the desired position and velocity vectors respectively generated by the following desired trajectory generator: , Among them and are the auxiliary state vectors of the desired trajectory generator, and are the system matrix and output matrix of the desired trajectory generator, respectively.

3. The robot-environment optimal interaction control method based on inverse reinforcement learning according to claim 2, characterized in that: , , , , , , , 。 4. The robot-environment optimal interaction control algorithm based on inverse reinforcement learning according to claim 2, characterized in that: The second step includes: By constructing an augmented state vector , an augmented output vector , the following state-space model of the actual controlled system is obtained: , wherein represents the non-inertial interaction force, which is also the control strategy to be obtained, , , represent the system matrix, input matrix, and output matrix of the actual controlled system, represents the total inertia coefficient of the environment and impedance, and represent the zero matrix and identity matrix of the corresponding dimensions respectively. The superscript "T" represents the transpose of a matrix or vector, and the superscript "-1" represents the inverse of a matrix; To quantitatively evaluate the interaction performance between the robot and the environment, the value function is defined as follows: , wherein and represent the state and the weight coefficient of the input, represents the discount factor, is the kernel matrix to be solved, represents the current time, represents the integration time, is the natural exponential constant; The expression of the optimal impedance control strategy is: , Among them is the optimal impedance control gain to be solved, where The solution formula for is as follows: 。 5. The robot-environment optimal interaction control method based on inverse reinforcement learning according to claim 3, characterized in that: The third step includes: Set the optimal expert impedance control strategy Applied to the actual controlled system , which generates the expert state and the expert output , then the system is written as: , Among them, the system is called the expert demonstration system; the optimal expert impedance control strategy and the expert output are collectively referred to as expert demonstration data, and the following state reconstruction method is introduced: , Among them represents the state reconstruction matrix, is the collected expert output data set, represents the sampling period, is the dimension of the data set, and its value needs to ensure that the matrix is column full rank; Based on this, the expert impedance control strategy is expressed as: , wherein is the expert impedance control gain based on state feedback, is the expert impedance control gain based on output feedback; Estimate the expert impedance control gain with the help of expert demonstration data , and the solution method is as follows: , Among them is the collected expert input data set, is the collected expert output data set, represents the dimension of the data set.

6. The robot-environment optimal interaction control method based on inverse reinforcement learning according to claim 5, wherein: Sampling period , the dimension of the dataset for reconstructing the state , the dimension of the dataset for restoring the expert output control strategy .

7. The robot-environment optimal interaction control method based on inverse reinforcement learning according to claim 5, characterized in that: The fourth step includes: Use the expert demonstration data to obtain the optimal impedance control strategy, and the algorithm flow is as follows: a) Given weights and discount factor , set the iteration variable , , the initial value of the iteration weight is , the initial value of the inverse Hessian matrix is , and the estimated value of the expert impedance control gain is ; b) Optimal control iterative solution based on integral reinforcement learning: b.1) Based on the expert demonstration data, the kernel matrix in the output feedback control is iteratively solved using the following Bellman equation and the control gain based on the output feedback : , wherein is the core matrix in the output feedback control, is the weight matrix of the output data set, represents the impedance control gain based on output feedback. The subscript "i" represents the iteration number of integral reinforcement learning, and the superscript "j" represents the iteration number of inverse reinforcement learning; b.2) Check the conditions to see if they are satisfied, where is a constant; if the conditions are satisfied, let , , , and go to the next step; otherwise, let and go to b.1); c) Improve the solution using the quasi-Newton algorithm and : c.1) Calculate the control gain error scalar and the error vector : , Among them is an operation operator that stacks any matrix " " by columns to form a column vector; c.2) Calculate the gradient vector , where the auxiliary variable ; c.3) The search direction is ; c.4) The vector form of the improved solution is , where is the search step size, is the kernel matrix vector expression; c.5) Convert to matrix form to obtain the improved solution , and calculate the improved output feedback control gain ; c.6) Update using the quasi-Newton algorithm ; d) Based on the expert demonstration data, use the improved solution and Update the weights : ; e) Check conditions and are both satisfied, and are both constants; if the conditions are satisfied, stop the iteration and return the impedance matrix based on the output feedback , the kernel matrix , and the weights ; otherwise, let and repeat steps b)-e).

8. The method for optimal interaction control of a robot-environment based on inverse reinforcement learning according to claim 4, characterized in that: , , search compensation Perform real-time calculation using the Wolfe criterion, , , .

9. A computing device, characterized in that, It includes: At least one processor and a memory storing program instructions; When the program instructions are read and executed by the processor, the computing device is caused to execute the robot-environment optimal interaction control method based on inverse reinforcement learning according to any one of claims 1-8.

10. A readable storage medium storing program instructions, characterized in that, When the program instructions are read and executed by the computing device, the computing device is caused to execute the robot-environment optimal interaction control method based on inverse reinforcement learning according to any one of claims 1-8.

Citation Information

Patent Citations

  • Intelligent flexible control method for contact process of space manipulator and unknown environment

    CN114851193A