Space double-arm robot control method based on machine learning and LQR control

By combining machine learning and LQR control in spatial two-arm robot control, the problems of dynamic modeling complexity and dynamic changes in the prior art are solved, achieving more efficient online adaptive control and more precise tracking performance.

CN119973980AActive Publication Date: 2025-05-13HEFEI UNIV OF TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510035507.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-13
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

When the prior art deals with the complex nonlinear dynamic characteristics of spatial two-arm robots, the dynamic complexity of dynamics and dynamic changes and noise influence of the task operating environment make it difficult to realize online adaptive control, and the control delay problem has not been fully solved.

Method used

Using a method based on machine learning and LQR control, a dynamic model of a space double-arm robot is constructed, and the prediction error is obtained through following error and deep learning model is designed to design the optimal LQR control to achieve effective control of a space double-arm robot.

Benefits of technology

It significantly improves the adaptability of the dynamic system to external changes, improves the tracking accuracy of the controller, and systematically improves the control performance, which is suitable for complex track task scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119973980A_ABST
    Figure CN119973980A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of robot control, in particular to a space double-arm robot control method based on machine learning and LQR control. Comprising the following steps: constructing a dynamic model of the space double-arm robot in a space environment; observing the track of the space double-arm robot, and constructing a following error; the following error is combined with a deep learning model to obtain a prediction error of space double-arm robot control; and designing LQR optimal control according to the prediction error. Different from a traditional controller only based on current state feedback, the method carries out control compensation in advance by predicting future deviation, and the adaptability of a dynamic system to external changes is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robot control technology, and in particular to a space dual-arm robot control method based on machine learning and LQR control. Background Art

[0002] As an important tool for orbital operation and maintenance tasks, the space dual-arm robot has complex nonlinear dynamic characteristics in a free-floating environment. Existing technologies are mainly based on traditional control methods based on precise modeling, such as LQR and robust control, which can handle some known uncertainties. However, there are still shortcomings in the face of the following challenges: the high-dimensional complexity of dynamic modeling makes it difficult for model parameterization methods to adapt to complex nonlinear scenarios; the dynamic changes and noise influence of the mission operation environment make it difficult for traditional methods to effectively implement online adaptive control; and the control delay problem has not been fully solved. Summary of the invention

[0003] The present invention discloses a spatial dual-arm robot control method based on machine learning and LQR control, and the specific method is as follows:

[0004] Construct a dynamic model of a space dual-arm robot in a space environment;

[0005] Observe the trajectory of the dual-arm robot in space and construct the following error;

[0006] The following error is combined with the deep learning model to obtain the prediction error of the spatial dual-arm robot control;

[0007] According to the prediction error, LQR optimal control is designed to control the spatial dual-arm robot.

[0008] Furthermore, the dynamic model of the space dual-arm robot in the space environment includes the base inertia term, the coupling term between the manipulators, the joint dynamic term, and the end external force term. The specific formula of the dynamic model is constructed as follows:

[0009] The Lagrange equation is:

[0010]

[0011] Among them, L=KP is the Lagrangian function, K is kinetic energy, P is potential energy, and τ is the generalized force.

[0012] For a space dual-arm robot, define the generalized velocity variable in, To describe the linear velocity of the base With angular velocity The generalized variable of and are the angular velocities of the two robotic arm joints;

[0013] Then the kinetic energy and potential energy of the generalized coordinate x are:

[0014]

[0015] In a microgravity environment, G(x) can be ignored; substituting L=KP into the Lagrange equation and introducing the joint force / external force term, the dynamic equation of the dual-arm robot is obtained:

[0016]

[0017] Where: M is the inertia matrix of the space dual-arm robot; C is the nonlinear term; To control the control force and torque of the robot arm; J e is the Jacobian matrix of the end of the spatial dual-arm robot, which maps the velocity in the joint space to the velocity in the operation space; is the contact force and moment on the ends of the two arms, and f e It is accurately measured by the force sensor installed at the end of the robot arm;

[0018] The inertia matrix M(z) is in the form of:

[0019]

[0020] M 12 (z1,x2) and M 21 (x1, x2) is the coupling inertia characteristic between the two manipulators; M b is the inertia of the floating base; M b,i 、M i,b is the coupling inertia between the base and the robotic arm; M1(x1) and M2(x2) are the inertia matrices of the first robotic arm and the second robotic arm, respectively, which are specifically expressed as:

[0021]

[0022] Among them, J i (x j ) Jacobian matrix of the jth link; I j The inertia tensor of the jth connecting rod; The Jacobian matrix of the jth link's center of mass; m j The mass of the jth connecting rod;

[0023] Terminal Jacobian matrix J e Link joint speed and terminal motion speed The expression is:

[0024]

[0025] The terminal Jacobian matrix Je is a block matrix in the form of:

[0026]

[0027] J e1 (x1) and J e2 (x2) is the end Jacobian matrix of the first and second robotic arms:

[0028]

[0029] J v (x i ) and J ω (x i ) is the Jacobian matrix of linear velocity and angular velocity at the end;

[0030] According to the dynamic equation of the space dual-arm robot, the state equation of the system can be expressed as:

[0031]

[0032] in,

[0033] is the state matrix;

[0034] is the control input matrix;

[0035] is the external force input matrix.

[0036] Furthermore, the following error is constructed, and the specific method is as follows:

[0037] The actual state x of the space dual-arm robot is measured by various sensors;

[0038] Actual state using generalized coordinates The expected state is

[0039] The following error formula is:

[0040] e(t)=x(t)-x d (t)

[0041] If the expected following error is within the target range, then

[0042] ∥e(t)∥≤∈

[0043] Where ∈ is the error tolerance of the space dual-arm robot.

[0044] Furthermore, the prediction error of the spatial dual-arm robot control is obtained, and the specific method is as follows:

[0045] Collect the spatial dual-arm robot system status, the input force at the end of the robot arm, and the following error data to construct the training data set D:

[0046]

[0047] in, is the system state variable; f ei Input force to the end of the robot arm; is the following error; N is the total number of data points;

[0048] Train a nonlinear mapping using the training data set D make:

[0049]

[0050] Where θ is the model parameter and the goal is to minimize the difference between the predicted error and the true error;

[0051] During the training process, the loss function is optimized by minimizing the loss function. The loss function is constructed as follows:

[0052]

[0053] During the training process, when the actual error e i and prediction error When the deviation is large, the model parameters are adjusted and updated using gradient descent:

[0054]

[0055] Where η is the learning rate, is the gradient of the loss function with respect to the parameters.

[0056] Training completes nonlinear mapping After that, the prediction error of the spatial dual-arm robot control is The specific formula is as follows:

[0057] .

[0059] Furthermore, the LQR optimal control is designed, and the specific method is as follows:

[0060] The prediction error is introduced into the state equation to construct the extended state equation;

[0061] Design the objective function of the LQR controller;

[0062] Solve the Richcat i equation and calculate the extended feedback gain matrix and the optimal control law.

[0063] Furthermore, the extended state equation is constructed, and the specific method is as follows:

[0064] The prediction error Introduced into the state equation, expanded the original system dynamics, expanded state variables:

[0065]

[0066] but

[0067]

[0068] The expanded equation is integrated into the state description:

[0069]

[0070] in

[0071]

[0072] In the extended system, the controller is designed through LQR optimization to minimize the state deviation and the control input energy consumption. The objective function is designed as:

[0073]

[0074] Where Q is the extended state weight matrix, including the weights for position, velocity and prediction error; R is the input weight matrix, constraining the control energy.

[0075] Furthermore, the Riccati equation is solved to calculate the extended feedback gain matrix and the optimal control law. The specific method is as follows:

[0076] Solve the continuous Riccati equation and use the solution P of the Riccati equation to calculate the extended feedback gain matrix K ext :

[0077]

[0078] Where P is a symmetric positive definite matrix describing the weights of the extended state;

[0079] Based on P, calculate the extended feedback gain matrix:

[0080]

[0081] The final optimal control law is:

[0082] u(t)=-K ext z ext (t)

[0083] The expanded form is:

[0084]

[0085] Among them, K1 is the feedback gain of the controller to the state variable z(t); K2 is the feedback gain of the controller to the prediction error The feedback gain.

[0086] Due to the adoption of the above technical solution, the present invention has the following beneficial effects:

[0087] 1. Different from the traditional controller which is based only on current state feedback, the present invention predicts future deviations and performs control compensation in advance, which significantly improves the adaptability of the dynamic system to external changes.

[0088] 2. The present invention combines the extended state modeling of the prediction error and expands the static feedback into a control scheme including dynamic prediction.

[0089] 3. The present invention adds a feedforward correction term based on machine learning prediction to the classic LQR control law, thereby improving the tracking accuracy of the controller.

[0090] 4. The present invention systematically improves the control performance from dynamic modeling, error prediction to controller design, and is suitable for complex orbital mission scenarios.

[0091] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] The accompanying drawings of the present invention are as follows.

[0093] Figure 1 This is a schematic diagram of the structure of a space dual-arm robot.

[0094] Figure 2 This is the control principle block diagram of the space dual-arm robot.

[0095] Figure 3 The figure is a schematic diagram of the overall process. DETAILED DESCRIPTION

[0096] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0097] A space robot is a special robot that can perform various tasks in a space environment. It is generally composed of a base spacecraft and two six-degree-of-freedom robotic arms. Figure 1The satellite base is equipped with two robotic arms, each with only one end effector. Considering that the robotic arm base is always in a free-floating state during the capture of non-cooperative targets, the space multi-arm robot system is in an under-actuated state. According to the Lagrangian method, the dynamic model of the system can be established.

[0098] A control method for a spatial dual-arm robot based on machine learning and LQR control, such as Figure 2 and Figure 3 As shown, the specific steps are as follows:

[0099] S1. Construct a dynamic model of a space dual-arm robot in a space environment.

[0100] In step S1, the dynamic model adopts the Lagrange equation:

[0101]

[0102] Where L=KP is the Lagrangian function, K is the kinetic energy, P is the potential energy, and τ is the generalized force.

[0103] For a space dual-arm robot, define the generalized velocity variable in To describe the linear velocity of the base With angular velocity The generalized variable. and Then are the angular velocities of the two robot joints. Then the kinetic energy and potential energy of the generalized coordinate x are:

[0104]

[0105] In a microgravity environment, G(x) can be ignored. Substituting L=KP into the Lagrange equation and introducing the joint force / external force term, the dynamic equation of the dual-arm robot is obtained:

[0106]

[0107] Where: M is the inertia matrix of the space dual-arm robot; C is the nonlinear term; To control the control force and torque of the robot arm; J e is the Jacobian matrix of the end of the spatial dual-arm robot, which maps the velocity in the joint space to the velocity in the operation space; is the contact force and moment on the ends of the two arms, and assuming that f e It can be accurately measured by a force sensor installed at the end of the robotic arm.

[0108] The inertia matrix M(x) is in the form of:

[0109]

[0110] M 12 (x1,x2) and M 21 (x1, x2) is the coupling inertia characteristic between the two manipulators; M b is the inertia of the floating base; M b,i 、M i,b is the coupling inertia between the base and the robot; M1(x1) and M2(x2) are the inertia matrices of robot 1 and robot 2, respectively, which are specifically expressed as:

[0111]

[0112] Among them, J i (x j ) Jacobian matrix of the jth link; I j The inertia tensor of the jth connecting rod; The Jacobian matrix of the jth link's center of mass; m j The mass of the jth connecting rod.

[0113] Terminal Jacobian matrix J e Link joint speed and terminal motion speed Including linear velocity and angular velocity, the expression is:

[0114]

[0115] The terminal Jacobian matrix J e is a block matrix in the form of:

[0116]

[0117] J e1 (x1) and J e2 (x2) is the end Jacobian matrix of robot arm 1 and robot arm 2:

[0118]

[0119] J v (x i ) and J ω (x i ) is the Jacobian matrix of the linear velocity and angular velocity at the end.

[0120] According to the dynamic equation of the space dual-arm robot, the state equation of the system can be expressed as:

[0121]

[0122] in,

[0123] is the state matrix;

[0124] is the control input matrix;

[0125] is the external force input matrix.

[0126] S2. Observe the trajectory of the spatial dual-arm robot and construct the following error.

[0127] For a space dual-arm robot, the actual state x of the system is measured by various sensors. If the actual state is expressed in generalized coordinates, description, and the desired state is The error is expressed as:

[0128] e(t)=x(t)-x d (t)

[0129] For a tracking control task, the expected error e(t) converges to zero, or varies within the allowed range:

[0130] ∥e(t)∥≤∈

[0131] where ∈ is the error tolerance of the spatial dual-arm robot system.

[0132] S3. Following error is combined with deep learning model to obtain the prediction error of spatial dual-arm robot control.

[0133] In actual systems, there are external disturbances and uncertainties that may not be predicted in advance, making it difficult for traditional optimization solutions to be calculated in real time. Using machine learning, the following error of the system can be predicted. Machine learning can use the learning model to train data offline, achieve online fast prediction, and can adapt to unmodeled uncertainties.

[0134] In step S3, a deep learning model is constructed. The specific method is as follows:

[0135] S31. Training data construction:

[0136] Collect the spatial dual-arm robot system status, end-arm input force and following error data to cover as many task conditions and workspaces as possible. According to the system status and end-arm input force, predict the system following error and define the training data set D:

[0137]

[0138] in:

[0139] System state variables;

[0140] f ei: Input force at the end of the robot arm;

[0141] Systematic errors;

[0142] N: The total number of data points.

[0143] S32. Machine Learning Model:

[0144] Train a nonlinear mapping using the collected data

[0145]

[0146] Where θ is the model parameter and the goal is to minimize the difference between the model prediction error and the true error, which is optimized by minimizing the loss function:

[0147]

[0148] In order to facilitate the subsequent LQR controller design, the error can be used as a feedback variable. For continuous time dynamics, the error change rate relationship can be regressed through the sequence prediction results. It can be described by the dynamic equation:

[0149]

[0150] where f err is the mapping relationship extracted from the training data.

[0151] S33, Error feedback and model adaptation:

[0152] When the actual error e(t) is equal to the predicted error When the deviation is large, the model parameters are adjusted and updated using gradient descent:

[0153]

[0154] Where η is the learning rate, is the gradient of the loss function with respect to the parameters.

[0155] S34, model training stop condition:

[0156] During the training process, if the model's loss function value L(θ) no longer decreases significantly after several consecutive iterations, the model training is considered to have converged. At this point, the training can be stopped.

[0157]

[0158] S4. Based on the prediction error, design LQR optimal control to control the spatial dual-arm robot.

[0159] In step S, the LQR optimal controller is designed. The specific steps are as follows:

[0160] S41. Constructing the extended state equation

[0161] The error of machine learning prediction Introduced into the state equation, expanded the original system dynamics, expanded state variables:

[0162]

[0163] but

[0164]

[0165] The expanded equation is integrated into the state description:

[0166]

[0167] in

[0168]

[0169] S42. Cost function definition

[0170] In the extended system, the controller is designed through LQR optimization to minimize the state deviation and the control input energy consumption. The objective function is:

[0171]

[0172] Where Q is the extended state weight matrix, including the weights for position, velocity and prediction error; R is the input weight matrix, constraining the control energy.

[0173] S43, Riccati equation solution

[0174] To design the LQR control law, we need to first solve the continuous Riccati equation and use the solution P of the Riccati equation to calculate the extended feedback gain matrix K ext :

[0175]

[0176] Where P is a symmetric positive definite matrix describing the weights of the extended state.

[0177] S44, feedback gain matrix

[0178] Based on P, calculate the extended feedback gain matrix:

[0179]

[0180] S45, Calculation of control law

[0181] The final optimal control law is:

[0182] u(t)=-K ext z ext (t)

[0183] The expanded form is:

[0184]

[0185] Among them, K1 is the feedback gain of the controller to the state variable z(t); K2 is the feedback gain of the controller to the prediction error The feedback gain.

[0186] The traditional LQR feedback law - K1z(t) dominates the dynamic adjustment of the current state to ensure the convergence and stability of the system. Realize feedforward correction of future deviations and further improve control accuracy.

[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A spatial dual-arm robot control method based on machine learning and LQR control, characterized in that: The specific method is as follows: Construct a dynamic model of a space dual-arm robot in a space environment; Observe the trajectory of the dual-arm robot in space and construct the following error; The following error is combined with the deep learning model to obtain the prediction error of the spatial dual-arm robot control; According to the prediction error, LQR optimal control is designed to control the spatial dual-arm robot.

2. The spatial dual-arm robot control method based on machine learning and LQR control as claimed in claim 1, characterized in that: The dynamic model of the space dual-arm robot in the space environment includes the base inertia term, the coupling term between the manipulators, the joint dynamic term, and the end external force term. The specific formula of the dynamic model is constructed as follows: The Lagrange equation is: Among them, L=KP is the Lagrangian function, K is kinetic energy, P is potential energy, and τ is the generalized force. For a space dual-arm robot, define the generalized velocity variable in, To describe the linear velocity of the base With angular velocity The generalized variable of and are the angular velocities of the two robotic arm joints; Then the kinetic energy and potential energy of the generalized coordinate x are: In a microgravity environment, G(x) can be ignored; substituting L=KP into the Lagrange equation and introducing the joint force / external force term, the dynamic equation of the dual-arm robot is obtained: Where: M is the inertia matrix of the space dual-arm robot; C is the nonlinear term; To control the control force and torque of the robot arm; J e is the Jacobian matrix of the end of the spatial dual-arm robot, which maps the velocity in the joint space to the velocity in the operation space; is the contact force and moment on the ends of the two arms, and f e It is accurately measured by the force sensor installed at the end of the robot arm; The inertia matrix M(x) is in the form of: M 12 (x1,x2) and M 21 (x1, x2) is the coupling inertia characteristic between the two manipulators; M b is the inertia of the floating base; M b,i 、M i,b is the coupling inertia between the base and the robotic arm; M1(x1) and M2(x2) are the inertia matrices of the first robotic arm and the second robotic arm, respectively, which are specifically expressed as: Among them, J i (x j ) Jacobian matrix of the jth link; I j The inertia tensor of the jth connecting rod; The Jacobian matrix of the jth link's center of mass; m j The mass of the jth connecting rod; Terminal Jacobian matrix J e Link joint speed and terminal motion speed The expression is: The terminal Jacobian matrix J e is a block matrix in the form of: J e1 (x1) and J e2 (x2) is the end Jacobian matrix of the first and second robotic arms: J v (x i ) and J ω (x i ) is the Jacobian matrix of the linear velocity and angular velocity at the end; According to the dynamic equation of the space dual-arm robot, the state equation of the system can be expressed as: in, is the state matrix; is the control input matrix; is the external force input matrix.

3. The spatial dual-arm robot control method based on machine learning and LQR control as claimed in claim 2, characterized in that: Construct the following error as follows: The actual state x of the space dual-arm robot is measured by various sensors; Actual state using generalized coordinates The expected state is The following error formula is: e(t)=x(t)-x d (t) If the expected following error is within the target range, then ∥e(t)∥≤∈ Where ∈ is the error tolerance of the space dual-arm robot.

4. The spatial dual-arm robot control method based on machine learning and LQR control as claimed in claim 3, characterized in that: Obtain the prediction error of the spatial dual-arm robot control. The specific method is as follows: Collect the spatial dual-arm robot system status, the input force at the end of the robot arm, and the following error data to construct the training data set D: in, is the system state variable; f ei Input force to the end of the robot arm; is the following error; N is the total number of data points; Train a nonlinear mapping using the training data set D make: Where θ is the model parameter and the goal is to minimize the difference between the predicted error and the true error; During the training process, the loss function is optimized by minimizing the loss function. The loss function is constructed as follows: During the training process, when the actual error e i and prediction error When the deviation is large, the model parameters are adjusted and updated using gradient descent: Where η is the learning rate, is the gradient of the loss function with respect to the parameters. Training completes nonlinear mapping After that, the prediction error of the spatial dual-arm robot control is The specific formula is as follows:

5. The spatial dual-arm robot control method based on machine learning and LQR control as claimed in claim 4, characterized in that: Design LQR optimal control. The specific method is as follows: The prediction error is introduced into the state equation to construct the extended state equation; Design the objective function of the LQR controller; Solve the Riccati equation and calculate the extended feedback gain matrix and the optimal control law.

6. The spatial dual-arm robot control method based on machine learning and LQR control as claimed in claim 5, characterized in that: Construct the extended state equation as follows: The prediction error Introduced into the state equation, expanded the original system dynamics, expanded state variables: but The expanded equation is integrated into the state description: in In the extended system, the controller is designed through LQR optimization to minimize the state deviation and the control input energy consumption. The objective function is designed as: Where Q is the extended state weight matrix, including the weights for position, velocity and prediction error; R is the input weight matrix, constraining the control energy.

7. The spatial dual-arm robot control method based on machine learning and LQR control as claimed in claim 6, characterized in that: Solve the Riccati equation, calculate the extended feedback gain matrix and the optimal control law. The specific method is as follows: Solve the continuous Riccati equation and use the solution P of the Riccati equation to calculate the extended feedback gain matrix K ext : Where P is a symmetric positive definite matrix describing the weights of the extended state; Based on P, calculate the extended feedback gain matrix: The final optimal control law is: u(t)=-K ext z ext (t) The expanded form is: Among them, K1 is the feedback gain of the controller to the state variable z(t); K2 is the feedback gain of the controller to the prediction error The feedback gain.

Citation Information

Patent Citations

  • Spatial robot prediction control method based on quantum particle swarm optimization algorithm

    CN107662211A

  • Space robot arresting control system, reinforce learning method and dynamics modeling method

    CN109605365A

  • Robot offline reinforcement learning control method based on model

    CN116460860A

  • Six-degree-of-freedom parallel robot pose calibration method based on deep learning

    CN116638524A

  • Intelligent cooperative control method and system for space multi-arm robot to capture non-cooperative target

    CN116834014A