Non-cooperative target intelligent compliant capture control method considering pose measurement error

By employing a reinforcement learning method based on pose measurement error, an impedance controller was designed and its parameters were updated online. This solved the problem of contact force and torque exceeding the spacecraft's tolerance in non-cooperative target acquisition missions, and enabled safe and reliable acquisition in harsh space environments.

CN117182927BActive Publication Date: 2026-03-27BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-27
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately measure the pose of non-cooperative targets in harsh space environments, leading to contact forces and torques exceeding the spacecraft's tolerance during space robot capture missions. This poses significant risks and uncertainties, and traditional impedance control methods are unable to handle unknown contact environments and changing operational tasks.

Method used

An optimal impedance control method based on pose measurement error is adopted. An impedance controller is designed by approximating the unknown contact environment through a linear model. The impedance control parameters are updated online using integral reinforcement learning to adapt to the unknown contact environment and changing contact tasks, thereby achieving a compliant contact effect.

Benefits of technology

This paper presents an intelligent compliant capture control method that adapts to the motion of non-cooperative targets and pose measurement errors. It can handle unknown contact environments and changing tasks, reduce risks, and improve the safety and reliability of capture tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117182927B_ABST
    Figure CN117182927B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of non-cooperative target intelligent compliant capture control method considering pose measurement error, first, linear model is used to approximate unknown contact environment, and impedance controller based on pose measurement value is designed for non-cooperative target capture process;Then, based on the state space expression of the contact process of space robot and target constructed by pose measurement error, the system is divided into nominal and disturbance two kinds of systems according to whether error exists;Finally, an integral reinforcement learning method is proposed for nominal system to update impedance control parameters online, and applied to disturbance system to achieve soft contact effect using measurement error.The present application uses limited measurement data to learn optimal impedance control parameters online, can adapt to unknown contact environment and achieve optimal interaction performance, realize space robot safe compliant capture non-cooperative target.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of intelligent control of space robots, and particularly relates to a non-cooperative target intelligent compliant capture control method considering pose measurement errors. BACKGROUND

[0002] With the vigorous development of space technology, the demand for exploration and development, and utilization of space is rapidly growing, and the number of failed satellites and space debris is increasing rapidly, which makes the space environment deteriorate day by day. These on-orbit targets usually have typical non-cooperative characteristics such as unknown structural characteristics and motion information, which seriously threaten the safety of some basic space projects such as space stations, communication satellites and navigation satellites. How to capture these targets by using space robots to provide a safe and reliable operating environment for on-orbit spacecraft is a direction that has been focused on in recent years.

[0003] In order to successfully capture non-cooperative targets, visual measurement technology is usually needed to obtain target motion information. However, due to the harsh space environment, measurement distance and sensor performance, the measurement accuracy will inevitably be reduced, and when applied to space capture tasks, it may produce contact forces / torques that exceed the bearing capacity of the spacecraft, bringing great risks and uncertainties to the spacecraft system. How to make the robot and the target interact compliantly to complete the capture task is a challenging task that needs to be solved urgently.

[0004] Impedance control algorithm is an effective means to deal with contact control problems, which can handle the compliant capture task of non-cooperative targets. Traditional impedance control methods rely heavily on manual parameter tuning and cannot handle unknown contact environments and changing operation tasks, so techniques such as adaptive variable impedance control (P. Xia, J. Luo, and M. Wang, “Adaptive compliant controller for space robot stabilization in post-capture phase,” Proceedings of the Institution of Mechanical Engineers, Part G: Journal of Aerospace Engineering, vol. 235, no. 8, pp. 937-948, 2021) and online estimation of environmental parameters (Y. Lin, Z. Chen, and B. Yao, “Unified motion / force / impedance control for manipulators in unknown contact environments based on robust model-reaching approach,” in IEEE / ASME Transactions on Mechatronics, vol. 26, no. 4, pp. 1905-1913, 2021) are needed. In recent years, reinforcement learning algorithms have been widely used in impedance control (CN112894809A). Compared with adaptive control and environmental parameter estimation techniques, this algorithm not only allows for autonomous learning of impedance parameters using experience data, but also enables optimal trade-offs between contact force and tracking accuracy by defining an interaction performance function.Most of the existing impedance control methods based on reinforcement learning need to discretize the contact model, such as the literature (X. Liu, S. S. Ge, F. Zhao, and X. Mei, "Optimized interaction control for robot manipulator interacting with flexible environment," IEEE / ASME Transactions on Mechatronics, vol. 26, no. 6, pp. 2888-2898, 2021); in order to avoid the discretization step, the algorithm designed is closer to the actual contact process, the literature (H. Wu, Q. Hu, Y. Shi, J. Zheng, K. Sun, and J. Wang, "Space manipulator optimal impedance control using integral reinforcement learning," Aerospace Science and Technology, 139, 108388, 2023) studies the reinforcement learning method based on continuous contact model, but this algorithm is mainly applied to the operation task of cooperative target, and the target pose is assumed to be constant, which cannot be directly used for non-cooperative target capture task. On this basis, considering the actual problems such as the motion characteristics of non-cooperative targets and the pose measurement error, how to develop reinforcement learning technology is a difficult problem to be solved and extremely challenging. Therefore, the present application designs a reinforcement learning optimal impedance control method based on pose measurement error, which can adapt to unknown contact environment and changing contact task, and provides theoretical and technical support for space robot capturing non-cooperative target task. SUMMARY

[0005] For the on-orbit operation task of space robot capturing non-cooperative target, considering the motion of non-cooperative target, inaccurate pose measurement, continuous and unknown contact model and other problems, the present application provides a non-cooperative target intelligent compliant capture control method considering pose measurement error. This method comes from impedance control and reinforcement learning technology, which can learn optimal impedance parameters online using limited measurement data, does not depend on any prior knowledge of the environment model, and realizes compliant contact effect according to the disturbance term generated by measurement error, which can handle target motion, inaccurate pose measurement, continuous and unknown contact model, optimal interaction between robot and moving target and other problems, and can be used for intelligent compliant capture control of non-cooperative target.

[0006] To achieve the above purpose, the technical scheme adopted by the present application is:

[0007] A non-cooperative target intelligent compliant capture control method considering pose measurement error, for the optimal impedance control problem of non-cooperative target motion, inaccurate pose measurement, continuous and unknown contact model, first, the linear model is used to approximate the unknown contact environment, and the impedance controller based on the pose measurement value is designed; Next, considering the pose measurement error, the state space expression of the contact process of the space robot and the target is constructed, and the system is divided into nominal and disturbance systems according to whether the error exists; Finally, an integral reinforcement learning method is proposed to update the impedance control parameters online for the nominal system, and then applied to the disturbance system to achieve the effect of compliant contact by using the measurement error; The specific implementation steps are as follows:

[0008] (1) The linear model is used to approximate the unknown contact environment, and the impedance controller based on the pose measurement value is designed;

[0009] The contact force generated in the contact process between the end of the space robot and the non-cooperative target can be expressed as:

[0010]

[0011] Where, D e represents the environmental damping coefficient matrix, G e represents the environmental stiffness coefficient matrix, x and respectively represent the pose and velocity vectors of the end tool of the space robot, x e and respectively represent the actual pose and actual velocity vectors of the non-cooperative target, F e represents the force generated in the contact process;

[0012] Since the actual pose of the non-cooperative target is difficult to accurately obtain, the visual measurement method is usually used to roughly obtain, and the pose measurement error is defined as:

[0013]

[0014] Where, represents the actual measured pose of the non-cooperative target, δx e represents the pose measurement error. Assuming δx e < 0 indicates that the space robot does not contact the target, so it is impossible to judge whether the capture task is successful at this time; Therefore, in order to ensure the success of the capture task, δx e ≥ 0 is required to enable the space robot to contact the target;

[0015] According to the measured pose, the expression of the impedance controller can be selected as:

[0016]

[0017] Where, M ddenotes the desired mass matrix, D d , G d denote the desired damping and stiffness matrices, respectively, to be further designed; is the second derivative of x with respect to time;

[0018] (2) Based on the pose measurement error, the state space expression of the contact process between the space robot and the target is constructed, and the system is divided into two types of nominal and disturbance according to whether the error exists;

[0019] Let e = x-x e denotes the error between the actual pose of the non-cooperative target and the pose of the space robot end, denotes the error between the measured pose of the non-cooperative target and the pose of the space robot end, and the pose measurement error can be further expressed as:

[0020] δx e = e'-e

[0021] According to the expression of the contact force and the impedance controller, the compliant contact dynamics between the space robot end and the non-cooperative target can be expressed in the form of error as:

[0022]

[0023] wherein, is the selected state vector, is the first time derivative of ξ, is the first time derivative of e'; the system matrix A and the input matrix B can be expressed as:

[0024]

[0025] wherein, 0 and I denote zero matrix and unit matrix of appropriate dimensions; it is noted that the matrix A contains the environmental parameters D e and G e , and therefore is completely unknown; while the M d in the matrix B is the given desired parameter, and therefore is known; d is the disturbance term generated by the measurement error, and the expression is:

[0026]

[0027] u denotes the control input vector, and the expression is:

[0028]

[0029] wherein, K = [D d G d ] denotes the control gain to be further designed;

[0030] System Π1 contains a disturbance term d, called the disturbed system; when the disturbance term d is absent, the nominal system Π2 is generated as follows:

[0031]

[0032] where the parameters ξ, A, B, u are defined as the same as system Π1.

[0033] (3) An integral reinforcement learning method is proposed for the nominal system to update the impedance control parameters online, and then applied to the disturbed system to achieve the compliant contact effect using measurement errors.

[0034] The following value function V is selected to describe the dynamic interaction performance of the space robot end and the non-cooperative target:

[0035]

[0036] where Q and R represent symmetric weight coefficient matrices, γ is the discount factor, t represents the current time, τ represents the integral variable, and P represents the kernel matrix to be solved.

[0037] For the nominal system Π2 where the model parameters A are completely unknown, an online policy iteration integral reinforcement learning method is designed to learn the optimal impedance control parameters as follows:

[0038] a) Initialization: Given an initial stable control policy u0.

[0039] b) Policy evaluation: The control policy u i For the disturbed system, record the historical input and output data; let Δt represent the sampling period, and the kernel matrix P at the ith × Δt time can be solved online by the following Bellman equation: i :

[0040]

[0041] where ξ(t) and ξ(t+Δt) represent the state vectors at t and t+Δt time, respectively; if the condition ||P i -P i-1 ||≤κ is met, the learning is completed, otherwise go to step c); where κ is a constant set to judge whether the matrix P i converges to a given threshold, and the symbol ||·|| represents the matrix norm.

[0042] c) Policy improvement: Update the control policy at the next time:

[0043] u i+1 =-R -1 B T P i ξ

[0044] Proceed to step b).

[0045] The advantages of this invention compared to existing technologies are as follows: This invention considers the influence of pose measurement errors of non-cooperative moving targets and provides a state-space model based on measurement errors for the dynamic compliant interaction process; by leveraging reinforcement learning, impedance control techniques, and disturbance control theory, an online policy iterative integral reinforcement learning method is designed to update the optimal impedance parameters in real time. The control methods involved in this invention are all designed for continuous contact models, requiring no prior information about the contact model, thus overcoming the limitations of traditional methods that do not consider target motion and pose measurement errors. This allows it to be used for compliant capture control of non-cooperative targets in real-world scenarios. Attached Figure Description

[0046] Figure 1 A flowchart of a non-cooperative target intelligent compliant capture control method that takes into account pose measurement errors;

[0047] Figure 2 A schematic diagram illustrating the process of a space robot compliantly capturing a non-cooperative target. Detailed Implementation

[0048] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0049] like Figure 1 As shown, the specific implementation steps of this invention are as follows:

[0050] The first step is to target Figure 2 The robot shown is tasked with capturing a non-cooperative target. A linear model is used to approximate the unknown contact environment, and an impedance controller based on pose measurement values ​​is designed.

[0051] The contact force generated during the contact process between the end effector of a space robot and a non-cooperative target can be expressed as:

[0052]

[0053] Among them, D e G represents the environmental damping coefficient matrix. e Represents the environmental stiffness coefficient matrix, x and Let x represent the pose and velocity vectors of the end effector of the space robot, respectively. e and Let F represent the actual pose and actual velocity vector of the non-cooperative target, respectively. e This indicates the force generated during the contact process;

[0054] Since the actual pose of non-cooperative targets is difficult to obtain precisely, it is usually obtained roughly using visual measurement methods. The pose measurement error is defined as:

[0055]

[0056] where, denotes the actual measured pose of the non-cooperative target, δx e denotes the pose measurement error. It is assumed that δx e < 0 indicates that the space robot does not contact with the target, and at this time it is impossible to judge whether the capture task is successful or not; therefore, in order to ensure the success of the capture task, it is necessary to make δx e ≥ 0, so that the space robot can contact with the target;

[0057] According to the measured pose, the expression of the impedance controller can be selected as:

[0058]

[0059] where, M d denotes the desired mass coefficient matrix, D d , G d denote the desired damping coefficient and stiffness coefficient matrices, respectively, which are to be further designed; is the second derivative of x with respect to time;

[0060] In this embodiment, it is assumed that the space robot end approaches the target along the x-axis direction of the inertial system, and the environmental parameters in this direction are taken as D e = 10 Ns / m, G e = 300 N / m, and the desired mass parameter is taken as M d = 1 kg; for the non-cooperative target floating freely in space, only its linear motion in the short term is considered, and its dynamics can be expressed as:

[0061]

[0062] where, M e denotes the mass of the non-cooperative target, taken as M e = 150 kg; denotes the actual acceleration vector of the non-cooperative target, and the initial values of the position and velocity of the non-cooperative target are taken as x e = 7.38 m, The pose measurement error of the non-cooperative target mainly considers constant error and variable error, and the expressions of the measured position and velocity are respectively t denotes the current simulation time.

[0063] Secondly, the state space expression of the contact process between the space robot and the target is constructed based on the pose measurement error, and the system is divided into nominal and disturbance systems according to whether the error exists or not;

[0064] Let e = x - x eerror between the actual pose of the non-cooperative target and the pose of the end-effector of the space robot, error between the measured pose of the non-cooperative target and the pose of the end-effector of the space robot, the pose measurement error can be further expressed as:

[0065] δx e = e' - e

[0066] According to the expression of the contact force and the impedance controller, the compliant contact dynamics between the end-effector of the space robot and the non-cooperative target can be expressed in the form of error as:

[0067]

[0068] where, is the selected state vector, is the first order time derivative of ξ, is the first order time derivative of e'; the system matrix A and the input matrix B can be expressed as:

[0069]

[0070] where, 0 and I represent the zero matrix and the identity matrix of appropriate dimensions; it is noted that the matrix A contains the environmental parameter D e and G e , which are completely unknown; while the matrix B contains M d , which is the given desired parameter, thus is known; d is the disturbance term caused by the measurement error, which is expressed as:

[0071]

[0072] u represents the control input vector, which is expressed as:

[0073]

[0074] where, K = [D d G d ] represents the control gain that needs to be further designed;

[0075] The system Π1 contains the disturbance term d, which is called the disturbed system; when the disturbance term d does not exist, the nominal system Π2 is generated as follows:

[0076]

[0077] where, the parameters ξ, A, B, u are defined the same as the system Π1;

[0078] Thirdly, an integral reinforcement learning method is proposed for the nominal system to update the impedance control parameters online, which is then applied to the disturbed system to achieve the compliant contact effect using the measurement error.

[0079] The following value function V is selected to describe the dynamic interaction performance of the space robot end and the non-cooperative target:

[0080]

[0081] Wherein, Q and R represent symmetric weight coefficient matrices, γ is a discount factor, t represents the current time, τ represents an integral variable, and P represents a kernel matrix to be solved.

[0082] For the nominal system Π2 in which the model parameter A is completely unknown, an online policy iteration integral reinforcement learning method is designed to learn the optimal impedance control parameter as follows:

[0083] a) Initialization: Given an initial stable control policy u0;

[0084] b) Policy evaluation: The control policy u i is used for the disturbed system to record historical input and output data; let Δt represent a sampling period, and the kernel matrix P at the ith×Δt moment can be solved online through the following Bellman equation i :

[0085]

[0086] Wherein, ξ(t) and ξ(t+Δt) represent state vectors at the t and t+Δt moments respectively; if the condition ||P i -P i-1 ||≤κ is met, the learning is completed, otherwise step c) is turned to; wherein κ is a constant set for judging whether the matrix P i converges to a given threshold, and the symbol ||·|| represents a matrix norm;

[0087] c) Policy improvement: The control policy at the next moment is updated as follows:

[0088] u i+1 =-R -1 B T P i ξ;

[0089] Step b) is turned to.

[0090] In the embodiment of the application, the system simulation step length is taken as Δt=2 ms, the weight coefficient matrix is taken as Q=diag([100;300000]), R=0.01, the discount factor is taken as γ=300, the convergence threshold is set as κ=0.01, the initial value of all elements in the matrix P is set as 0.1, and correspondingly, the initial value of the control gain K is K=[10 10]. Finally, the optimal impedance control gain solved is K=[94.166 667.0623].

[0091] The content described in the specification of the present application is the prior art known to those skilled in the art. Those skilled in the art can easily understand that the above description is only the preferred embodiment of the present application and is not intended to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for intelligent compliant capture control of a non-cooperative target considering pose measurement errors, characterized in that, The method comprises the following steps: In the first step, a linear model is used to approximate the unknown contact environment, and an impedance controller based on pose measurement is designed for the non-cooperative target capture process; In the second step, a state space expression of the contact process between the space robot and the target is constructed based on the pose measurement error, and the system is divided into two types of systems, i.e., a nominal system and a disturbance system, according to whether the error exists; In the third step, an integral reinforcement learning method is proposed to update the impedance controller parameters online for the nominal system, and is applied to the disturbance system to achieve a compliant contact effect by using the measurement error; The first step is specifically implemented as follows: The contact force generated in the contact process between the end of the space robot and the non-cooperative target is expressed as: wherein, denotes the environmental damping coefficient matrix, denotes the environmental stiffness coefficient matrix, and denote the pose and velocity vector of the end-effector of the space robot, respectively, and denote the actual pose and actual velocity vector of the non-cooperative target, respectively, denotes the contact force resulting from the contact process; Since the actual pose of the non-cooperative target is difficult to accurately obtain, a visual measurement method is usually used to roughly obtain the pose, and the pose measurement error is defined as: wherein, denotes the actual measured pose of the non-cooperative target, denotes the pose measurement error; it is assumed that denotes that the space robot does not have contact with the target, at this time it is not possible to determine whether the capture task is successful; therefore, in order to ensure the success of the capture task, it is necessary to make , so that the space robot can have contact with the target; According to the measured pose, the expression of the impedance controller is selected as: wherein, denotes a desired mass coefficient matrix, , denote a desired damping coefficient and stiffness coefficient matrix, respectively, to be further designed; is second derivative with respect to time.

2. The method of claim 1, wherein: The second step is specifically implemented as follows: Let denote the error between the pose of the end-effector of the space robot and the actual pose of the non-cooperative target, denote the error between the pose of the end-effector of the space robot and the measured pose of the non-cooperative target, the pose measurement error is further denoted as: According to the contact force and the expression of the impedance controller, the compliant contact dynamics of the end of the space robot and the non-cooperative target is expressed in the form of error as: : wherein is the state vector of the selected states, is the first time derivative of is the first time derivative of is the first time derivative of is the first time derivative of is the system matrix may be represented as: , where and denote zero and identity matrices of appropriate dimensions; note that the matrix contains the environmental parameters and and is therefore completely unknown; while the matrix contains the given desired parameters and is therefore known; is a perturbation term due to measurement errors and is expressed as: denotes the control input vector, which is expressed as: wherein denotes a control gain that needs to be further designed; System with a perturbation term , called the perturbed system; when the perturbation term is absent, then the nominal system results: : Wherein, the parameters , , , , are defined the same as the system .

3. The method of claim 1, wherein: The third step is specifically implemented as follows: The value function is defined as follows Describing the dynamic interaction performance of the space robot end-effector with a non-cooperative target: wherein, and denotes a symmetric weight coefficient matrix, is a discount factor, denotes the current time, denotes the integration variable, denotes a kernel matrix to be solved for; For model parameters Completely unknown nominal system To learn the optimal impedance controller parameters, an online policy iteration integrated reinforcement learning method is designed as follows: a) Initialization: Given an initial stable control policy ; b) policy evaluation: evaluate the control policy at the current time step For the perturbed system, record the history input-output data; let denote the sampling period, then the kernel matrix at the current time step can be solved on-line by the following Bellman equation ; in, , They represent , The state vector at time t; if the condition If the condition is met, the learning process is complete; otherwise, proceed to step c). A set constant is used to determine the matrix. Whether it converges to a given threshold, sign Represents the matrix norm; c) Strategy improvement: update the control strategy at the next time: Go to step b).

Citation Information

Patent Citations

  • Impedance controller design method and system based on reinforcement learning

    CN112894809A

  • Intelligent flexible control method for contact process of space manipulator and unknown environment

    CN114851193A