Mobile manipulator vision-haptic hybrid compliant control method based on security reinforcement learning
Patent Information
- Application Number
- CN202611192156.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-07
- Publication Date
- 2026-09-25
AI Technical Summary
[0066]本发明的移动机械臂视觉-力觉混合柔顺控制方法通过移动机械臂视觉图像空间的运动学与动力学建模,将视觉和力觉信息在图像空间进行统一的表征;本发明设计了基于力环动态参数与图像空间的视力混合导纳控制方法,能够及时响应环境刚度及阻尼的变化,保证控制器的稳定性和跟踪精度;本发明有效实现非结构化环境下移动机械臂接触作业的视觉-力觉混合跟踪控制,保证了接触作业的精度及柔顺性,实现了移动机械臂非结构化接触作业的高适应性柔顺控制。
Smart Images

Figure CN122807913A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot motion control, and specifically to a vision-force hybrid compliant control method for a mobile robotic arm based on safety reinforcement learning. Background Technology
[0002] In unstructured environments, the force and position control of mobile robotic arms in contact operations is highly susceptible to failure due to minute errors, potentially even damaging the robotic arm itself or the work object. Highly compliant force / position control technology is essential for mobile robotic arms to perform contact operations and ensure safe human-machine interaction. However, in unstructured contact operations, poor force tracking due to changes in target position and stiffness is a frequent challenge. Furthermore, single pose or force / torque sensors are easily affected by external interference and their own performance due to limited information. Therefore, integrating multi-sensor information to design an intelligent hybrid compliant control strategy is crucial for enabling mobile robotic arms to perform autonomous contact operations in unstructured environments. Visual sensors provide macroscopic environmental information and target localization, but are easily obstructed and cannot perceive minute contacts and precise interaction forces. While force / torque sensors provide contact information, they are passive and prone to localized adjustments or blind exploration.
[0003] Force / position hybrid control methods utilize position and force signals designed separately in free and constrained spaces, i.e., force and position decoupled control is performed separately in different degrees of freedom. However, numerous studies have pointed out the limitations of force / position hybrid control, such as its high dependence on complete contact dynamics constraints, making it difficult to adapt to dynamic changes in unstructured environments, or its impact on force / position tracking accuracy due to neglecting the dynamic coupling between the robotic arm and the operating environment. Impedance control methods, on the other hand, control contact forces by compensating for position and velocity errors based on the relationship between the end effector's position or velocity and the applied force. However, existing impedance control methods still suffer from problems such as force / position coupling and sensitivity to system parameter changes, requiring further research in impedance parameter selection and robust adaptive methods. Current vision-force fusion control research often uses visual detection of pose information to achieve force / position hybrid control in different spatial dimensions. However, the inability to accurately perceive constraint directions on unknown environmental surfaces leads to conflicts between force / position control targets and coupling instability. Furthermore, impedance control algorithms often assume constant stiffness and a single force control direction in variable impedance stability design, making them difficult to adapt to complex environments. Therefore, integrating vision-force / torque sensing information transforms robots from single-sensor systems into closed-loop multimodal sensing intelligent systems. For example, vision-force hybrid compliant control technology enables mobile robotic arms to "see" and "touch," helping them achieve precise and compliant autonomous operation in unstructured environments. Summary of the Invention
[0004] To address the limitations of existing technologies in adapting to dynamic changes in unstructured environments, neglecting the dynamic coupling between the robotic arm and the operating environment and thus affecting force / position tracking accuracy, and the continued technical problems of existing impedance control methods such as sensitivity to force / position coupling and system parameter changes, as well as difficulty in adapting to complex environments, this invention provides a vision-force hybrid compliant control method for mobile robotic arms based on safety reinforcement learning. The technical solution is as follows:
[0005] A vision-force hybrid compliant control method for a mobile robotic arm based on safety reinforcement learning includes the following steps:
[0006] Step 1: Design the kinematic and dynamic model of the mobile robotic arm in the visual image space to uniformly represent visual and force information in the image space;
[0007] Step 2: Design visual feature vectors to represent the motion pose of the end effector of the mobile robotic arm; design a highly stable trajectory planning method to represent the end-effector task in the visual image space.
[0008] Step 3: Design a vision-force hybrid admittance control method for a mobile robotic arm based on force loop dynamic parameters and image space, and construct a state-independent stability constraint model for the variable admittance control system.
[0009] Step 4: Design a safety reinforcement learning method for self-tuning admittance control parameters, and finally obtain the corrected joint commands of the mobile robotic arm to perform highly adaptive and compliant control of the mobile robotic arm for unstructured contact operations.
[0010] Furthermore, step 1 defines a Cartesian world coordinate system. Mobile platform coordinate system Camera coordinate system With end coordinate system ;based on Compared to The position, orientation, and joint angles of the manipulator define the generalized coordinate vector of the mobile robotic arm. The kinematic model of the mobile robotic arm in the visual image space is as follows:
[0011] ,
[0012] in, Represents the image feature vector; It is the image Jacobian matrix; the dynamic model of the mobile robotic arm's visual image space is:
[0013] ,
[0014] in, The inertia matrix represents a symmetric positive definite inertia matrix; The input torque representing the generalized joint space of the mobile robotic arm; This is the matrix of centrifugal force and Coriolis force coefficients; Represents the gravitational moment vector; For generalized unknown torque vectors; The transformation matrix represents the camera velocity. This represents the inverse mapping of the inertia matrix of the mobile robotic arm to the camera space; The nominal inertia matrix representing the mobile robotic arm in camera space; The camera velocity is represented from the coordinate system. arrive The transformation matrix; Represents the external forces / torques involved in operations that come into contact with the environment, expressed in a coordinate system. Inside.
[0015] Furthermore, the image Jacobian matrix for:
[0016] ,
[0017] in, represent In coordinate system The internal expression form, It is the movement of the robotic arm relative to the coordinate system. The Jacobian matrix; the formula for the inverse mapping of the mobile robotic arm's inertia matrix to camera space is:
[0018] .
[0019] Furthermore, in step 2, the perspective projection of visual feature points onto the camera image plane of the mobile robotic arm is used. The camera motion speed is mapped to the motion speed of visual features in the image space, and the visual feature vector is designed. for:
[0020] ,
[0021] in, ;
[0022] ;
[0023] ;
[0024] ;
[0025] ;
[0026] ,
[0027] By substituting its first-order time derivative with the motion velocity of the image spatial visual features, the camera velocity representation transformation matrix is calculated. .
[0028] Furthermore, in step 2, let Given the desired trajectory of visual features representing the image space of the end effector of a mobile robotic arm, design a planning algorithm that minimizes the rate of change of trajectory jerk to solve for the desired trajectory of each element of the visual feature vector. ,Right now:
[0029] ,
[0030] in, and These represent the expected trajectories of visual features at the initial time. and the final moment The The expected value of the second time derivative; The integral representing the square of the rate of change of the trajectory jerks; the first The expected trajectory of a spatial visual feature of an image is defined as a multinomial model:
[0031] ,
[0032] in, Represents the order of a polynomial; Represents trajectory parameters.
[0033] Furthermore, in step 3, the vision-force hybrid admittance control method for the mobile robotic arm based on force loop dynamic parameters and image space specifically involves: designing an image-based vision hybrid admittance controller, whose output control law dynamically modifies the desired trajectory of non-contact visual features. The desired trajectory of visual features for highly compliant contact operations is obtained. The formula is:
[0034] ,
[0035] in, , , These are positive definite diagonal matrices representing the mass, damping, and stiffness coefficients of the admittance control expectation, respectively. Desired trajectory of non-contact visual features The amount of correction; The tracking error represents the virtual force projected onto the visual image space by the interactive force at the end effector of the mobile robotic arm; a high-bandwidth force loop controller is designed, with the following formula:
[0036] ,
[0037] in, The tracking error represents the interaction force at the end effector of the mobile robotic arm. In coordinate system The expectancy of internal representation; , , These represent the positive definite proportional and integral gain diagonal matrices of the controller, respectively. and It is a positive odd number and satisfies .
[0038] Furthermore, the tracking error of the virtual force projected onto the visual image space by the interactive force at the end of the mobile robotic arm. It possesses the following dynamic characteristics:
[0039] ,
[0040] in, It is a time-varying positive definite diagonal coefficient matrix; and It maps the visual feature trajectory correction amount and its rate of change to a continuously differentiable time-varying diagonal matrix of virtual force error dynamics.
[0041] Furthermore, in step 3, the real-time constraint admittance controller parameters and force loop dynamic parameters of the Lyapunov function are constructed, as shown in the following formula:
[0042] ,
[0043] in, It is a positive definite diagonal matrix of virtual force tracking error weights; This is the weight coefficient matrix. It is a mass diagonal matrix. , This is the diagonal matrix for damping and stiffness; For undetermined constants, , For operation matrix The minimum and maximum eigenvalues; let Design a matrix block to represent the identity matrix:
[0044] ,
[0045] ,
[0046] ,
[0047] ,
[0048] ,
[0049] ,
[0050] The state-independent stability constraint model of the variable admittance control system is as follows:
[0051] .
[0052] Furthermore, in step 4, the safety reinforcement learning for self-tuning the admittance control parameters is defined as a constrained Markov decision process:
[0053] ,
[0054] in, and This represents the state space and action space of the reinforcement learning algorithm; Discount factor representing the reward; Represents the reward function; It is the probability of system state transition; It is the initial value of the system state; and These represent the safety constraint cost function and its upper bound on the expected cumulative cost, respectively; the input state of the safety reinforcement learning is defined as:
[0055] ,
[0056] in, The visual feature tracking error is represented by: The instantaneous reward function for the security reinforcement learning is:
[0057] ,
[0058] in, Represents the weighting coefficient of each reward item; It is the expected value of the visual feature representing the distance between the camera projection plane and the surface of the interactive environment; the instantaneous cost function of the security reinforcement learning is:
[0059] .
[0060] Furthermore, the constrained Markov decision process optimization problem is as follows:
[0061] ,
[0062] in, The optimal strategy model; Represents the time span of the exploration; The policy function is modeled using a feedforward neural network, with the hidden layer employing the leaky ReLU activation function and the output layer using an exponential activation function. The classic proximal policy optimization algorithm, PPO, is used to solve the problem. The gradient descent algorithm is used to optimize the solution of the Lagrange operator. The vision servo controller for the mobile robotic arm is designed as follows:
[0063] ,
[0064] in, The desired velocity command representing the joint space of the mobile robotic arm; It is the image Jacobian matrix The false reversal; It is a positive definite diagonal matrix of gain coefficients; Represents visual feature tracking error. The actual location representing the visual feature.
[0065] Beneficial effects
[0066] The present invention provides a vision-force hybrid compliant control method for mobile robotic arms. By modeling the kinematics and dynamics of the mobile robotic arm's visual image space, it unifies the representation of visual and force information in the image space. The present invention designs a vision-force hybrid admittance control method based on force loop dynamic parameters and image space, which can respond promptly to changes in environmental stiffness and damping, ensuring the stability and tracking accuracy of the controller. The present invention effectively realizes vision-force hybrid tracking control for contact operations of mobile robotic arms in unstructured environments, ensuring the accuracy and compliance of contact operations, and achieving highly adaptive compliant control for unstructured contact operations of mobile robotic arms. Attached Figure Description
[0067] Figure 1 Flowchart of a vision-force hybrid compliant control method for a mobile robotic arm based on safety reinforcement learning;
[0068] Figure 2 A schematic diagram of the reinforcement learning results for the visual admittance coefficients of a mobile robotic arm in contact with a variable stiffness / variable damping environment.
[0069] Figure 3 This is a schematic diagram illustrating the force tracking effect of a mobile robotic arm in contact with a variable stiffness / variable damping environment. Detailed Implementation
[0070] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0071] This invention first constructs a kinematic and dynamic model of the visual image space of a mobile robotic arm, and represents visual and force information in a unified manner in the image space. Then, it designs a vision-hybrid admittance control method based on force loop dynamic parameters and image space, as well as a safety reinforcement learning algorithm for self-tuning admittance parameters-force loop dynamic parameters, and finally achieves highly adaptive compliant control of the mobile robotic arm for contact operations in unstructured environments.
[0072] like Figure 1 As shown, the specific process of the vision-force hybrid compliant control method for mobile robotic arms based on safety reinforcement learning of the present invention is as follows:
[0073] Step 1: Design the kinematic and dynamic model of the mobile robotic arm in the visual image space to uniformly represent visual and force information in the image space.
[0074] make Represents the Cartesian world coordinate system. Define the coordinate system representing the mobile platform with its origin at the centroid, and define the camera coordinate system. and end coordinate system Having the same orientation; let represent Compared to Position and direction Let the joint angles of the manipulator represent the coordinates of the moving robotic arm. Then, the generalized coordinate vector of the moving robotic arm is defined as follows: The kinematic model for the visual spatial representation of the mobile robotic arm is shown in the following formula (1):
[0075] ,
[0076] In formula (1), This represents the image feature vector, which is calculated by projecting visual feature points in the environment onto the camera image plane. It is the image Jacobian matrix. The camera velocity is represented from the coordinate system. arrive The transformation matrix, represent In coordinate system The internal expression form, It is the movement of the robotic arm relative to the coordinate system. Jacobian matrix, The transformation matrix representing the camera velocity is full rank.
[0077] Based on this, the dynamic model of the mobile robotic arm in the visual image space is designed as shown in the following formula (2):
[0078] ,
[0079] In formula (2), The inertia matrix represents a symmetric positive definite inertia matrix; It is the matrix of centrifugal force and Coriolis force coefficients; Represents the gravitational moment vector; It is a generalized unknown torque vector; Represents the external forces / torques involved in operations that come into contact with the environment, expressed in a coordinate system. Inside; The input torque representing the generalized joint space of the mobile robotic arm; It inversely maps the inertia matrix of the mobile robotic arm to the camera space. The nominal inertia matrix representing the mobile robotic arm in camera space is a positive definite constant diagonal matrix of inertia.
[0080] The last term on the right side of formula (2) will move the robotic arm in the coordinate system. End-effector interaction forces of internal characterization It is also mapped to the visual image space, and it no longer depends on the inertia matrix of the moving robotic arm. And Jacobi matrix Therefore, changes in the joint position of the mobile robotic arm no longer cause changes in the representation of its end effector force in the image space, which is more conducive to the design of vision-force hybrid control algorithms.
[0081] Step 2: Design visual feature vectors to represent the motion pose of the end effector of the mobile robotic arm; design a highly stable trajectory planning method to represent the end-effector task in the visual image space.
[0082] make Representing the The perspective projection of a visual feature point onto the image plane of the moving robotic arm camera is given by the following formula (3): The mapping relationship between the camera's motion speed and the motion speed of the visual features in the image space is as follows:
[0083] ,
[0084] In formula (3), Representing visual feature points in the camera coordinate system Depth information within.
[0085] In order to observe and control the six-dimensional pose of the visual camera in the task space, the visual feature vector is designed as shown in the following formula (4):
[0086] ,
[0087] In formula (4),
[0088] ;
[0089] ;
[0090] ;
[0091] ;
[0092] ;
[0093] ;
[0094] The translational motion of the camera in three-dimensional space is determined by visual features. Controlled rotational motion is achieved through visual features. Control. Calculate the first-order time derivative of formula (4), and substitute the feature point velocities calculated in formula (3) into it to obtain the transformation matrix of the camera velocity. .
[0095] The vision-force hybrid admittance control algorithm aims to make the end effector of the mobile robotic arm follow the desired motion trajectory while maintaining the desired interaction force with the environment. Since the pose of the end effector of the mobile robotic arm can be defined by the visual feature vector of formula (4), Completely determined, therefore the desired trajectory of the end effector can be obtained by planning the desired trajectory of the image space visual features. Let Given the desired trajectory representing visual features, design a planning algorithm that minimizes the rate of change of trajectory jerks to solve for the desired trajectory of each element in the visual feature vector. As shown in the following formula (5):
[0096] ,
[0097] In formula (5), , No. The expected trajectory of each visual feature is defined as a multinomial model:
[0098] ,
[0099] in, Represents the order of a polynomial; Represents trajectory parameters; and These represent the expected trajectories of visual features at the initial time. and the final moment The The expected value of the second time derivative, that is, the expected values of position, velocity, and acceleration respectively in the above formula; This represents the integral of the square of the rate of change of trajectory jerk. This image-space trajectory planning method not only determines the desired trajectory for the smooth operation of the mobile robotic arm's end effector, but also retains the advantages of image-based visual servoing methods in robot task execution control.
[0100] Step 3: Design a vision-force hybrid admittance control method for a mobile robotic arm based on force loop dynamic parameters and image space, and construct a state-independent stability constraint model for the variable admittance control system.
[0101] To enable the mobile robotic arm to track the desired interactive force during contact operations, an image-based vision-hybrid admittance controller is designed as shown in Equation (6). Its output control law dynamically modifies the desired trajectory of the non-contact visual features. This allows for the acquisition of the desired visual trajectory for highly compliant contact operations. To enable tracking of interactive forces.
[0102] ,
[0103] In formula (6), , , These are positive definite diagonal matrices representing the mass, damping, and stiffness coefficients of the admittance control expectation, respectively. Desired trajectory of non-contact visual features The amount of correction; Tracking error representing the virtual force projected onto the visual image space by the interaction force at the end of the mobile robotic arm; the magnitude of the interaction force can be adjusted by modifying the motion trajectory of the visual features.
[0104] Due to factors such as force sensor measurement delay, insufficient controller bandwidth, and limited actuator response, the virtual force tracking error in visual space cannot respond instantaneously, but instead exhibits the dynamic characteristics shown in the following formula (7).
[0105] ,
[0106] In formula (7), It is a time-varying positive definite diagonal coefficient matrix; and It maps the visual feature trajectory correction amount and its rate of change to a continuously differentiable time-varying diagonal matrix of virtual force error dynamics.
[0107] Based on the dynamic model of the mobile robotic arm in the visual image space, a high-bandwidth force loop controller is designed as shown in the following formula (8):
[0108] ,
[0109] In formula (8), The tracking error represents the interaction force at the end effector of the mobile robotic arm. In coordinate system The expectancy of internal representation; , , These represent the positive definite proportional and integral gain diagonal matrices of the controller, respectively. and It is a positive odd number and satisfies This fractional term can further accelerate the convergence of the force tracking error near the origin.
[0110] To ensure system stability during the adaptive adjustment of admittance parameters, real-time constraints on admittance controller parameters and force loop dynamic parameters need to be designed. The Lyapunov function for this purpose is constructed as shown in the following formula (9):
[0111] ,
[0112] In formula (9), These are undetermined constants; It is a positive definite diagonal matrix of virtual force tracking error weights; construct the weight coefficient matrix. As shown in the following formula (10):
[0113] ,
[0114] For image-based vision-hybrid admittance control systems for mobile robotic arms It is a time-invariant positive definite mass diagonal matrix, damping and stiffness The diagonal matrix is time-varying, positive definite, and continuously differentiable; the virtual force tracking error is dynamically determined by the force loop parameters. , and Determine the design of undetermined constants. The values are shown in the following formula (11):
[0115] ,
[0116] In formula (11), and Representing operation matrices respectively The minimum and maximum eigenvalues of ; let . Representing the identity matrix, design the following matrix block.
[0117] ,
[0118] ,
[0119] ,
[0120] ,
[0121] ,
[0122] ,
[0123] To ensure that the vision-hybrid admittance control system of the mobile robotic arm is strictly and progressively stable, it is necessary to... At that time, block matrix Since it is negative definite, the state-independent stability constraint model of the above-mentioned mobile robotic arm variable admittance control system is designed as shown in the following formula (12):
[0124] ,
[0125] If the time-varying coefficient matrix If it is negative definite, then It always holds true, and the coefficient matrix is also true. Is positive definite such that the Lyapunov function If non-negativity is satisfied, then the rate of change of the correction amount of the expected trajectory of visual features Force tracking error All will gradually converge to the origin.
[0126] Step 4: Design a safety reinforcement learning method for self-tuning admittance control parameters to achieve highly adaptive and compliant control of the mobile robotic arm in unstructured contact operations.
[0127] The safety reinforcement learning problem of the hybrid admittance controller parameters of a mobile robotic arm's vision is defined as a constrained Markov decision process:
[0128] ,
[0129] in, and This represents the state space and action space of the reinforcement learning algorithm; Discount factor representing the reward; Represents the reward function; It is the probability of system state transition; It is the initial value of the system state; and These represent the upper bounds of the safety constraint cost function and its expected cumulative cost, respectively.
[0130] The input state for admittance control parameter safety reinforcement learning is defined as follows: This refers to the image spatial projection, which includes visual feature motion speed and tracking error, visual feature expected trajectory correction, and end-effector interaction force tracking error. This represents the visual feature tracking error. To explore the dynamic relationship between the end effector force and visual features during unstructured contact operations of a mobile robotic arm, and to maintain the continuity of the control law, the instantaneous reward function for safety reinforcement learning is designed as shown in formula (13):
[0131] ,
[0132] In formula (13), Represents the weighting coefficient of each reward item; It is the expected value of the visual feature representing the distance between the camera projection plane and the surface of the interactive environment. It can be calculated by the planning algorithm. The related reward is intended to ensure effective contact between the end effector of the mobile robotic arm and the surface of the working environment. The last reward is used to limit the abrupt changes in the policy output and ensure the smoothness of the vision hybrid admittance control law.
[0133] Based on the stability condition formula (12) of the vision-hybrid admittance control system of the mobile robotic arm, in order to ensure the safety and stability of the control system during the reinforcement learning adjustment of the admittance parameters, the instantaneous cost function is designed as shown in the following formula (14):
[0134] ,
[0135] The main objective of the aforementioned vision-based hybrid admittance controller parameter safety reinforcement learning algorithm design is to solve for the optimal policy model. To maximize the expected cumulative reward while ensuring that the cumulative safety cost satisfies the constraints, the corresponding constrained Markov decision process optimization problem can be designed as shown in the following formula (15):
[0136] ,
[0137] In formula (15), Represents the time span of the exploration. In the safety reinforcement learning process of the vision-hybrid admittance controller parameters of the mobile robotic arm, the policy function... A feedforward neural network is used for modeling, and the hidden layer uses the leaky ReLU activation function. Meanwhile, to ensure the coefficient matrix... , , To ensure positive definiteness, an exponential activation function is used in the output layer. Furthermore, the classic Proximal Policy Optimization (PPO) algorithm is employed to solve for the optimal policy. The gradient descent algorithm is used to optimize the solution of the Lagrange operator. This enables a rapid solution to the aforementioned optimization problem.
[0138] Based on this, in order to track the visual features of the contact operation and the desired trajectory To achieve tracking of interactive forces, the visual servo controller for the mobile robotic arm is designed as shown in the following formula (16):
[0139] ,
[0140] In formula (16), The desired velocity command representing the joint space of the mobile robotic arm; It is the image Jacobian matrix The false reversal; It is a positive definite diagonal matrix of gain coefficients; Represents visual feature tracking error. The actual position representing the visual features is calculated by the camera. Based on this, the method corrects the positional error of the visual tracking, ultimately achieving highly adaptive and compliant control of the mobile robotic arm for unstructured contact operations.
[0141] Example
[0142] The flowchart of the vision-force hybrid compliance control method for mobile robotic arms based on safety reinforcement learning in this embodiment is as follows: Figure 1 As shown, the specific implementation object is a mobile robotic arm consisting of an omnidirectional moving platform and a six-degree-of-freedom robotic arm. The actual working force at the end effector is calculated from the environmental stiffness and damping parameters. Simultaneously, the focal length of the vision system is defined as 800mm, and the center coordinates of the image plane are... By combining pixels with the end effector of the mobile robotic arm and the camera pose, the visual feature trajectory corresponding to the visual feature point can be calculated, and finally online feedback of the visual trajectory can be achieved.
[0143] In the implementation of this embodiment, the relevant parameters of the vision-force hybrid compliance control method for the mobile robotic arm are set as follows:
[0144] , , , , , , , , , In the security reinforcement learning process, the policy neural network is designed with two hidden layers, with 64 and 32 neurons respectively. The number of training iterations is set to 500, and the reward discount factor is... The pruning threshold and entropy coefficient of the PPO algorithm are set to 0.2 and 0.02, respectively, and the learning rate of the policy neural network is set to 0.001 to meet the long-term exploration requirements of continuous control tasks and ensure stable learning and updating.
[0145] Make the planar environment parallel to the world coordinate system of The plane, the motion trajectory of the end effector of the mobile robotic arm on it. The definition is as follows: the working plane is located at... At the location, the end effector of the mobile robotic arm is , Both axes move sinusoidally upwards; simultaneously , Centered on, with Four visual feature points are defined on the working plane to represent the side length. Furthermore, the expected value of the normal contact force between the end effector of the mobile robotic arm and the working plane is set to... .
[0146] In the implementation of this embodiment, the contact stiffness and damping between the mobile robotic arm and the planar environment are respectively set as follows: and 2 To further verify the adaptability of the designed self-learning vision hybrid compliant control algorithm, when time is... At that time, the contact stiffness and damping of the environment will both increase by 20%; while when time is in At that time, the contact stiffness and damping of the environment were both reduced by 20%, thereby obtaining the visual hybrid compliance control results for the mobile robotic arm in contact with a variable stiffness / variable damping environment, as shown in the figure. Figure 2 , Figure 3 As shown. Figure 2 This demonstrates that the designed safety reinforcement learning algorithm can adaptively adjust the stiffness coefficient of the vision-hybrid admittance controller and respond promptly to changes in environmental stiffness and damping, thereby ensuring the stability and tracking accuracy of the controller; simultaneously, Figure 3 The proposed method has been demonstrated to effectively achieve vision-force hybrid tracking control for contact operations of mobile robotic arms in unstructured environments, ensuring the accuracy and compliance of the contact operations.
[0147] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A vision-force hybrid compliant control method for a mobile robotic arm based on safety reinforcement learning, characterized in that, Includes the following steps: Step 1: Design the kinematic and dynamic model of the mobile robotic arm in the visual image space to uniformly represent visual and force information in the image space; Step 2: Design visual feature vectors to represent the motion pose of the end effector of the mobile robotic arm; design a highly stable trajectory planning method to represent the end-effector task in the visual image space. Step 3: Design a vision-force hybrid admittance control method for a mobile robotic arm based on force loop dynamic parameters and image space, and construct a state-independent stability constraint model for the variable admittance control system. Step 4: Design a safety reinforcement learning method for self-tuning admittance control parameters, and finally obtain the corrected joint commands of the mobile robotic arm to perform highly adaptive and compliant control of the mobile robotic arm for unstructured contact operations.
2. The vision-force hybrid compliant control method for a mobile robotic arm based on safety reinforcement learning as described in claim 1, characterized in that: Step 1 defines the Cartesian world coordinate system. Mobile platform coordinate system Camera coordinate system With end coordinate system ;based on Compared to The position, orientation, and joint angles of the manipulator define the generalized coordinate vector of the mobile robotic arm. The kinematic model of the mobile robotic arm in the visual image space is as follows: , in, Represents the image feature vector; It is the image Jacobian matrix; the dynamic model of the visual image space of the mobile robotic arm is: , in, The inertia matrix represents a symmetric positive definite inertia matrix; The input torque representing the generalized joint space of the mobile robotic arm; This is the matrix of centrifugal force and Coriolis force coefficients; Represents the gravitational moment vector; For generalized unknown torque vectors; The transformation matrix represents the camera velocity. This represents the inverse mapping of the inertia matrix of the mobile robotic arm to the camera space; The nominal inertia matrix representing the mobile robotic arm in camera space; The camera velocity is represented from the coordinate system. arrive The transformation matrix; Represents the external forces / torques involved in operations that come into contact with the environment, expressed in a coordinate system. Inside.
3. The vision-force hybrid compliant control method for a mobile robotic arm based on safety reinforcement learning as described in claim 2, characterized in that: The image Jacobian matrix for: , in, represent In coordinate system The internal expression form, It is the moving robotic arm relative to the coordinate system The Jacobian matrix; the formula for the inverse mapping of the mobile robotic arm's inertia matrix to camera space is: 。 4. The vision-force hybrid compliant control method for a mobile robotic arm based on safety reinforcement learning as described in claim 3, characterized in that: In step 2, the perspective projection of visual feature points onto the image plane of the mobile robotic arm's camera is used. The camera motion speed is mapped to the motion speed of visual features in the image space, and the visual feature vector is designed. for: , in, ; ; ; ; ; , By substituting its first-order time derivative with the motion velocity of the image spatial visual features, the camera velocity representation transformation matrix is calculated. .
5. The vision-force hybrid compliant control method for a mobile robotic arm based on safety reinforcement learning as described in claim 4, characterized in that: In step 2, let Given the desired trajectory of visual features representing the image space of the end effector of a mobile robotic arm, design a planning algorithm that minimizes the rate of change of trajectory jerk to solve for the desired trajectory of each element of the visual feature vector. ,Right now: , in, and These represent the expected trajectories of visual features at the initial time. and the final moment The The expected value of the second time derivative; The integral representing the square of the rate of change of the trajectory jerks; the first The expected trajectory of a spatial visual feature of an image is defined as a multinomial model: , in, Represents the order of a polynomial; Represents trajectory parameters.
6. The vision-force hybrid compliant control method for a mobile robotic arm based on safety reinforcement learning as described in claim 5, characterized in that: In step 3, the vision-force hybrid admittance control method for the mobile robotic arm based on force loop dynamic parameters and image space specifically involves: designing an image-based vision hybrid admittance controller, whose output control law dynamically modifies the desired trajectory of non-contact visual features. The desired trajectory of visual features for highly compliant contact operations is obtained. The formula is: , in, , , These are positive definite diagonal matrices representing the mass, damping, and stiffness coefficients of the admittance control expectation, respectively. Desired trajectory of non-contact visual features The amount of correction; The tracking error represents the virtual force projected onto the visual image space by the interactive force at the end effector of the mobile robotic arm; a high-bandwidth force loop controller is designed, with the following formula: , in, The tracking error represents the interaction force at the end effector of the mobile robotic arm. In coordinate system The expectancy of internal representation; , , These represent the positive definite proportional and integral gain diagonal matrices of the controller, respectively. and It is a positive odd number and satisfies .
7. The vision-force hybrid compliant control method for a mobile robotic arm based on safety reinforcement learning as described in claim 6, characterized in that: Tracking error of the virtual force projected onto the visual image space by the interactive force at the end of the mobile robotic arm. It possesses the following dynamic characteristics: , in, It is a time-varying positive definite diagonal coefficient matrix; and It maps the visual feature trajectory correction amount and its rate of change to a continuously differentiable time-varying diagonal matrix of virtual force error dynamics.
8. The vision-force hybrid compliant control method for a mobile robotic arm based on safety reinforcement learning as described in claim 7, characterized in that: In step 3, the parameters of the Lyapunov function real-time constraint admittance controller and the dynamic parameters of the force loop are constructed, as shown in the following formula: , in, It is a positive definite diagonal matrix of virtual force tracking error weights; This is the weight coefficient matrix. It is a mass diagonal matrix. , This is the diagonal matrix for damping and stiffness; For undetermined constants, , For operation matrix The minimum and maximum eigenvalues; let Design a matrix block to represent the identity matrix: , , , , , , The state-independent stability constraint model of the variable admittance control system is as follows: 。 9. The vision-force hybrid compliant control method for a mobile robotic arm based on safety reinforcement learning as described in claim 8, characterized in that: In step 4, the safety reinforcement learning for self-tuning the admittance control parameters is defined as a constrained Markov decision process: , in, and This represents the state space and action space of the reinforcement learning algorithm; Discount factor representing the reward; Represents the reward function; It is the probability of system state transition; It is the initial value of the system state; and These represent the safety constraint cost function and its upper bound on the expected cumulative cost, respectively; the input state of the safety reinforcement learning is defined as: , in, The visual feature tracking error is represented by: The instantaneous reward function for the security reinforcement learning is: , in, Represents the weighting coefficient of each reward item; It is the expected value of the visual feature representing the distance between the camera projection plane and the surface of the interactive environment; the instantaneous cost function of the security reinforcement learning is: 。 10. The vision-force hybrid compliant control method for a mobile robotic arm based on safety reinforcement learning as described in claim 9, characterized in that: The constrained Markov decision process optimization problem is as follows: , in, This is the optimal strategy model; Represents the time span of the exploration; The policy function is modeled using a feedforward neural network, with the hidden layer employing the leaky ReLU activation function and the output layer using an exponential function-like activation function. The classic proximal policy optimization algorithm, PPO, is used to solve the problem. The gradient descent algorithm is used to optimize the solution of the Lagrange operator. The vision servo controller for the mobile robotic arm is designed as follows: , in, The desired velocity command representing the joint space of the mobile robotic arm; It is the image Jacobian matrix The false reversal; It is a positive definite diagonal matrix of gain coefficients; Represents visual feature tracking error. The actual location representing the visual feature.