A variable impedance control method for shaft-hole assembly of a space manipulator based on reinforcement learning

By adopting a variable impedance control method based on reinforcement learning on the space manipulator and updating the impedance parameters in real time, the problem of the space manipulator being difficult to achieve stable contact force control in complex environments is solved, and higher control accuracy and stability are achieved.

CN115256401BActive Publication Date: 2025-09-19NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211038250.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-29
Publication Date
2025-09-19
Estimated Expiration
2042-08-29

AI Technical Summary

Technical Problem

When a space robot performs shaft-hole assembly in a complex environment, it is difficult to achieve stable contact force control. Especially when the environmental geometry and stiffness parameters are uncertain, the traditional constant impedance control method is difficult to meet the mission requirements.

Method used

A variable impedance control method based on reinforcement learning is adopted. By constructing an impedance controller, using a neural network to train the impedance parameters, and updating the parameters of the impedance controller in real time, the ideal dynamic relationship between the end position and contact force of the robotic arm can be achieved.

Benefits of technology

It realizes the compliant control of the space robot arm in complex environments, improves the control accuracy and stability of assembly operations, can quickly respond to and accurately track dynamic forces, and weakens the influence of uncertain factors in the environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115256401B_ABST
    Figure CN115256401B_ABST
Patent Text Reader

Abstract

The present invention discloses a reinforcement learning-based variable impedance control method for the shaft-hole assembly of a space manipulator. First, a space manipulator model and a conversion model of the manipulator joint angle state and the end position are constructed separately. Then, a binocular camera is used to collect the position information of the assembly hole, and an impedance controller based on reinforcement learning is constructed. The impedance controller is trained using a neural network. Then, real-time information of the end of the manipulator is input, the impedance parameters of the impedance controller are updated, and the position correction value of the end of the manipulator is output to complete the variable impedance control of the shaft-hole assembly of the space manipulator. The solution of the present invention is based on reinforcement learning to perform variable impedance control on the shaft-hole assembly of the space manipulator. The control can track dynamic forces, and the dynamic error is smaller than that of traditional fixed impedance control, and the response speed is also faster. It can effectively weaken the influence of uncertain factors in the environment and has better tracking accuracy than traditional fixed impedance control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of space manipulator control, and in particular relates to a variable impedance control method for shaft-hole assembly of a space manipulator based on reinforcement learning. Background Art

[0002] With the advancement and development of space technology, the use of spacecraft and space stations has significantly impacted human production and daily life. Due to the vacuum and weightlessness of the space environment, a large amount of space debris and garbage floats in the space surrounding the Earth, posing a serious threat to the safety of in-orbit spacecraft and space stations. Furthermore, as space facilities age, they inevitably face issues such as aging and malfunctions, necessitating their maintenance.

[0003] When performing on-orbit assembly and other service tasks, space manipulators inevitably generate force contact with the external environment. This places high demands on contact force control. Furthermore, the space environment is subject to various external disturbances, such as gravity gradient torque and friction, which must be overcome. Compliant control can adapt the manipulator's movements to changes in the external environment, effectively improving the control accuracy and stability of assembly operations.

[0004] To coordinate the contact force between the manipulator and the environment, Hogen N pioneered impedance control. This approach achieves compliant contact between the robot and the environment by establishing an ideal dynamic relationship between the contact force at the manipulator's end and the deviation between the desired and actual trajectories. However, constant impedance control struggles to maintain a stable contact force when the environment's geometry and stiffness parameters are uncertain. The environments in which space manipulators perform their tasks are complex and ever-changing, making accurate environmental information difficult to identify. Furthermore, due to the presence of nonlinear, time-varying factors in the target environment, impedance control with fixed parameters is difficult to achieve the target task. If the impedance control parameters can be dynamically adjusted in real time based on changes in the task and environment, the control performance will be improved. Summary of the Invention

[0005] Based on the above-mentioned problems, the purpose of the present invention is to provide a variable impedance control method for the shaft-hole assembly of a space robot arm based on reinforcement learning, which can update the impedance controller parameters in the interaction with a complex environment, ensure the rapidity of static force response and the accuracy of dynamic force tracking, and realize the smooth control of the space robot arm assembly operation.

[0006] The technical solution for achieving the purpose of the present invention is:

[0007] A variable impedance control method for shaft-hole assembly of a space manipulator based on reinforcement learning comprises the following steps:

[0008] Step 1: Construct a spatial manipulator model based on the DH parameter method;

[0009] Step 2: Construct a transformation model of the joint angle state and end position of the spatial manipulator based on the forward and inverse kinematics algorithm;

[0010] Step 3: Initialize the internal and external parameters of the binocular camera, and use the binocular camera to capture images to obtain the position information of the assembly holes;

[0011] Step 4: Build an impedance controller based on reinforcement learning and set the impedance parameter action table, reward function, and termination conditions during training according to the expected goals.

[0012] Step 5: training the impedance controller based on the neural network;

[0013] Step 6: Input the real-time information of the end of the manipulator, update the impedance parameters of the impedance controller, output the position correction value of the end of the manipulator, and complete the variable impedance control of the spatial manipulator shaft-hole assembly.

[0014] Compared with the prior art, the present invention has the following significant advantages:

[0015] (1) The technical solution of the present invention is to perform variable impedance control on the shaft-hole assembly of a space manipulator based on reinforcement learning. The control can track the dynamic force, and the dynamic error is smaller than that of the traditional fixed impedance control, and the response speed is also faster, and the tracking accuracy is better than that of the traditional fixed impedance control.

[0016] (2) The technical solution of the present invention realizes variable impedance control in the shaft-hole assembly of the manipulator based on reinforcement learning, which can effectively weaken the influence of uncertain factors in the environment and improve the accuracy and speed of the end force control of the space manipulator.

[0017] The present invention is further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a flow chart of the steps of the variable impedance control method for the shaft hole assembly of a space robot arm based on reinforcement learning of the present invention.

[0019] Figure 2 Schematic diagram of the structure of the impedance controller based on reinforcement learning of the present invention.

[0020] Figure 3 This is a flow chart of the neural network-based impedance controller training process of the present invention.

[0021] Figure 4 Schematic diagram of the fully connected neural network structure of the present invention.

[0022] Figure 5 This is a schematic diagram of the assembly of a space robot arm in an embodiment of the present invention.

[0023] Figure 6 Schematic diagram of a simulation of a space robot arm in an embodiment of the present invention.

[0024] Figure 7 4 is a diagram of the simulated impedance parameter trajectory in an embodiment of the present invention.

[0025] Figure 8 Schematic diagram of the simulation of the position trajectory of the end of the space robot arm in an embodiment of the present invention.

[0026] Figure 9 Schematic diagram of the velocity trajectory simulation of the end of the space manipulator in an embodiment of the present invention.

[0027] Figure 10 Schematic diagram of static force tracking trajectory simulation of a space manipulator in an embodiment of the present invention.

[0028] Figure 11 Schematic diagram of dynamic force tracking trajectory simulation of a space manipulator in an embodiment of the present invention. DETAILED DESCRIPTION

[0029] A variable impedance control method for shaft-hole assembly of a space manipulator based on reinforcement learning comprises the following steps:

[0030] Step 1: Construct a spatial manipulator model based on the DH parameter method;

[0031] Step 2: Construct a transformation model of the joint angle state and end position of the spatial manipulator based on the forward and inverse kinematics algorithm;

[0032] Step 3: Initialize the internal and external parameters of the binocular camera, and use the binocular camera to capture images to obtain the position information of the assembly holes;

[0033] Step 4: Build an impedance controller based on reinforcement learning and set the impedance parameter action table, reward function, and termination conditions during training according to the expected goals. Specifically:

[0034] Step 4-1: Build an impedance controller:

[0035] The goal of the impedance control strategy is to achieve an ideal dynamic relationship between the end position and the end contact force of the space robot. This application simplifies the relationship between the end fixture of the manipulator and the assembly plane into a spring-mass-damper model, the mathematical model of which is:

[0036]

[0037] in, , They represent the actual motion trajectory and the expected motion trajectory of the end of the space manipulator, Represents the force between the end of the robotic arm and the external environment, They correspond to the desired inertia matrix, desired stiffness matrix and desired damping matrix of the impedance controller respectively; 、 They represent the actual acceleration, expected acceleration, actual speed and expected speed of the end of the space manipulator respectively. The impedance controller is selected As a control quantity, Set to a constant value of 1;

[0038] Step 4-2: The control objective of the impedance controller in this application is to quickly track the desired force, so that the velocity of the robot end quickly approaches zero, while optimizing the overshoot during the static force tracking process (overshoot refers to the deviation between the maximum actual force of the system and the desired force, that is, the deviation between the peak and the desired force);

[0039] To this end, it is necessary to give corresponding rewards and penalties to the state of the end of the robotic arm during the training process. When the state of the end of the robotic arm reaches the desired goal, a corresponding positive reward is given, and the optimal control parameters are found and the reward function is set:

[0040]

[0041]

[0042] Where T represents the duration of a single training session, and v represents the speed of the end of the space manipulator;

[0043] In the above formula is the error between the expected force and the current force, T is the current training simulation duration, and the speed is expected to be quickly close to the range of 0-0.2. Set the reward function as above.

[0044] This function gives larger rewards for smaller force steady-state errors and larger penalties as the velocity deviates further from 0.

[0045] Step 4-3: Considering that if the single change of the impedance parameter is too small, the impedance control of the end position of the manipulator will not achieve a significant effect, and if the impedance parameter changes too much, the stability of the impedance control of the end position of the manipulator will be reduced, the impedance parameter action table for reinforcement learning is set:

[0046]

[0047] in, is the delta correction value set, is the stiffness coefficient transformation, In order to damp the transformation, the corresponding action is selected in each sampling period, and the optimal action strategy is obtained after multiple trainings.

[0048] In addition, the training termination condition is set as: the number of training times reaches a set threshold.

[0049] Alternatively, if the error between the expected force and the current force during training is greater than the set threshold, or if the error between the system's maximum current force and the expected force during training exceeds the set threshold, this indicates that the strategy for this training is diverging, and the central setting parameters should be returned to for retraining.

[0050] Step 5: Train the impedance controller based on the neural network, specifically:

[0051] The Q-learning algorithm is essentially a Markov decision process that takes actions in the current state to obtain the reward value of the next state, and continuously updates the Q table. The specific formula is as follows:

[0052]

[0053] The traditional Q-learning method updates the current state based on the Q-value of the next state. This method relies on the Q-value table, and too many system states will waste a lot of memory space. DDQN uses a fully connected neural network "policy network" to predict the Q-value of the current state. The state information of the end of the robotic arm is input into the "policy network" to obtain the Q-value at that moment, and the "target network" is introduced to predict the state at the next moment. The mean square error of the difference between the prediction results of the two neural networks is used as the loss function of the model, as shown in the following formula. The back-propagation network parameters finally realize the update of the "policy network".

[0054] Specifically: first, set the total number of training times, collect the experience table of the space robot arm in a single training, and place it in the experience pool (that is, the queue has a maximum storage length. Once the maximum length is exceeded, the experience with poor performance will be ejected). The experience with higher rewards in the experience pool will also be input into the policy network at intervals together with the experience randomly extracted from the experience pool. The policy network is updated by the residual between the predicted value in the policy network and the target network. Set the update time. Once the time is exceeded, the target network is replaced by the policy network to achieve the target network update. Finally, the target network outputs the action with the highest score through feedback from the environment. Repeat in sequence until the final set total number of training times is greater than the set value, and the training ends.

[0055] Furthermore, the policy network predicts the Q value of the current moment in the impedance controller based on reinforcement learning, and predicts the Q value of the next moment in the impedance controller based on reinforcement learning based on the target network, and the mean square error of the difference between the two moments is used as the loss function:

[0056]

[0057] in, represents the mean square error, ) represents the Q value at time t, represents the decay rate during the learning process, Represents the learning rate of the model.

[0058] Furthermore, the strategy network adopts a full neural network structure, takes the position, velocity, acceleration and force error information of the end of the robotic arm as the network input, sets the number of hidden layer neurons to 400, selects the ReLU function as the activation function, and outputs the Q value of each action at the current moment.

[0059] The target network adopts a full neural network structure, takes the position, velocity, acceleration and force error information of the end of the robotic arm as the network input, sets the number of hidden layer neurons to 400, selects the ReLU function as the activation function, and outputs the Q value of each action at the next moment.

[0060] The present invention improves the updating process of the experience pool, marks and stores the best state information generated during the training process, and inputs the best state information into the experience pool every training cycle. High-yield state information can improve the rapid convergence of the DDQN model.

[0061] Step 6: Input the real-time information of the end of the manipulator, update the impedance parameters of the impedance controller, output the position correction value of the end of the manipulator, and complete the variable impedance control of the spatial manipulator shaft-hole assembly.

[0062] A variable impedance control system for shaft-hole assembly of a space manipulator based on reinforcement learning, including the following modules:

[0063] Space manipulator model construction module: used to construct the space manipulator model based on the DH parameter method;

[0064] End-position conversion model construction module: used to construct the conversion model of the joint angle state and end-position of the spatial manipulator based on the forward and inverse kinematics algorithm;

[0065] Assembly hole position information acquisition module: used to initialize the internal and external parameters of the binocular camera, and use the binocular camera to capture images to obtain the position information of the assembly holes;

[0066] Impedance controller building module: used to build an impedance controller based on reinforcement learning and set the impedance parameter action table, reward function, and termination conditions during training according to the expected goals;

[0067] Training module: training impedance controller based on neural network;

[0068] Variable impedance control module for shaft-hole assembly of space manipulator: used to input real-time information of the end of the manipulator, update the impedance parameters of the impedance controller, output the position correction value of the end of the manipulator, and complete the variable impedance control of the shaft-hole assembly of the space manipulator.

[0069] A computer device comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the following steps are implemented:

[0070] Step 1: Construct a spatial manipulator model based on the DH parameter method;

[0071] Step 2: Construct a transformation model of the joint angle state and end position of the spatial manipulator based on the forward and inverse kinematics algorithm;

[0072] Step 3: Initialize the internal and external parameters of the binocular camera, and use the binocular camera to capture images to obtain the position information of the assembly holes;

[0073] Step 4: Build an impedance controller based on reinforcement learning and set the impedance parameter action table, reward function, and termination conditions during training according to the expected goals.

[0074] Step 5: training the impedance controller based on the neural network;

[0075] Step 6: Input the real-time information of the end of the manipulator, update the impedance parameters of the impedance controller, output the position correction value of the end of the manipulator, and complete the variable impedance control of the spatial manipulator shaft-hole assembly.

[0076] A computer storable medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the computer program implements the following steps:

[0077] Step 1: Construct a spatial manipulator model based on the DH parameter method;

[0078] Step 2: Construct a transformation model of the joint angle state and end position of the spatial manipulator based on the forward and inverse kinematics algorithm;

[0079] Step 3: Initialize the internal and external parameters of the binocular camera, and use the binocular camera to capture images to obtain the position information of the assembly holes;

[0080] Step 4: Build an impedance controller based on reinforcement learning and set the impedance parameter action table, reward function, and termination conditions during training according to the expected goals.

[0081] Step 5: training the impedance controller based on the neural network;

[0082] Step 6: Input the real-time information of the end of the manipulator, update the impedance parameters of the impedance controller, output the position correction value of the end of the manipulator, and complete the variable impedance control of the spatial manipulator shaft-hole assembly.

[0083] The present invention will be further described below with reference to the embodiments.

[0084] Example

[0085] Combine Figure 1 A variable impedance control method for shaft-hole assembly of a space manipulator based on reinforcement learning, comprising the following steps:

[0086] Step 1: Construct a spatial manipulator model based on the DH parameter method;

[0087] Step 2: Construct a transformation model of the joint angle state and end position of the spatial manipulator based on the forward and inverse kinematics algorithm;

[0088] Step 3: Initialize the internal and external parameters of the binocular camera, and use the binocular camera to capture images to obtain the position information of the assembly holes;

[0089] Step 4: Build an impedance controller based on reinforcement learning and set the impedance parameter action table, reward function, and termination conditions during training according to the expected goals. Specifically:

[0090] Step 4-1: Build an impedance controller:

[0091] Combine Figure 2 and Figure 3 The goal of the impedance control strategy is to achieve an ideal dynamic relationship between the end position and the end contact force of the space robot. This application simplifies the relationship between the end fixture of the manipulator and the assembly plane into a spring-mass-damper model, the mathematical model of which is:

[0092]

[0093] in, , They represent the actual motion trajectory and the expected motion trajectory of the end of the space manipulator, Represents the force between the end of the robotic arm and the external environment, They correspond to the desired inertia matrix, desired stiffness matrix and desired damping matrix of the impedance controller respectively; 、 They represent the actual acceleration, expected acceleration, actual speed and expected speed of the end of the space manipulator respectively. The impedance controller is selected As the control quantity, Set to a constant value of 1;

[0094] Step 4-2: The control objective of the impedance controller in this application is to quickly track the desired force, causing the end velocity of the manipulator to quickly approach zero, while optimizing the overshoot during the static force tracking process (overshoot refers to the deviation between the maximum actual force of the system and the desired force, i.e., the deviation between the peak and the desired force).

[0095] To this end, it is necessary to give corresponding rewards and penalties to the state of the end of the robotic arm during the training process. When the state of the end of the robotic arm reaches the desired goal, a corresponding positive reward is given, and the optimal control parameters are found and the reward function is set:

[0096]

[0097]

[0098] Among them, T represents the duration of a single training session. is the error between the expected force and the current force, T is the current simulation duration, and the speed is expected to be quickly close to the range of 0-0.2. Set the reward function as above.

[0099] This function gives larger rewards for smaller force steady-state errors and larger penalties as the velocity deviates further from 0.

[0100] Step 4-3: Considering that if the single change of the impedance parameter is too small, the impedance control of the end position of the manipulator will not achieve a significant effect, and if the impedance parameter changes too much, the stability of the impedance control of the end position of the manipulator will be reduced, the impedance parameter action table for reinforcement learning is set:

[0101]

[0102] in, The delta correction is set, and the corresponding action is selected in each sampling period. The optimal action strategy is obtained after multiple trainings.

[0103] In addition, the training termination condition is set as: the number of training times reaches a set threshold.

[0104] Alternatively, if the error between the expected force and the current force during training is greater than the set threshold, or if the error between the system's maximum current force and the expected force during training exceeds the set threshold, this indicates that the strategy for this training is diverging, and the central setting parameters should be returned to for retraining.

[0105] Step 5: Train the impedance controller based on the neural network, specifically:

[0106] The Q-learning algorithm is essentially a Markov decision process that takes actions in the current state to obtain the reward value of the next state, and continuously updates the Q table. The specific formula is as follows:

[0107]

[0108] The traditional Q-learning method updates the current state based on the Q-value of the next state. This method relies on the Q-value table, and too many system states will waste a lot of memory space. DDQN uses a fully connected neural network "policy network" to predict the Q-value of the current state. The state information of the end of the robotic arm is input into the "policy network" to obtain the Q-value at that moment, and the "target network" is introduced to predict the state at the next moment. The mean square error of the difference between the prediction results of the two neural networks is used as the loss function of the model, as shown in the following formula. The back-propagation network parameters finally realize the update of the "policy network".

[0109] Specifically: first, set the total number of training times, collect the experience table of the space robot arm in a single training, and place it in the experience pool (that is, the queue has a maximum storage length. Once the maximum length is exceeded, the experience with poor performance will be ejected). The experience with higher rewards in the experience pool will also be input into the policy network at intervals together with the experience randomly extracted from the experience pool. The policy network is updated by the residual between the predicted value in the policy network and the target network. Set the update time. Once the time is exceeded, the target network is replaced by the policy network to achieve the target network update. Finally, the target network outputs the action with the highest score through feedback from the environment. Repeat in sequence until the final set total number of training times is greater than the set value, and the training ends.

[0110] Furthermore, the policy network predicts the Q value of the current moment in the impedance controller based on reinforcement learning, and predicts the Q value of the next moment in the impedance controller based on reinforcement learning based on the target network, and the mean square error of the difference between the two moments is used as the loss function:

[0111]

[0112] in, represents the mean square error, ) represents the Q value at time t, represents the decay rate during the learning process, Represents the learning rate of the model.

[0113] Further, combined Figure 4The strategy network adopts a full neural network structure, takes the position, velocity, acceleration and force error information of the end of the robotic arm as the network input, sets the number of hidden layer neurons to 400, selects the ReLU function as the activation function, and outputs the Q value of each action at the current moment.

[0114] The target network adopts a full neural network structure, takes the position, velocity, acceleration and force error information of the end of the robotic arm as the network input, sets the number of hidden layer neurons to 400, selects the ReLU function as the activation function, and outputs the Q value of each action at the next moment.

[0115] The present invention improves the updating process of the experience pool, marks and stores the best state information generated during the training process, and inputs the best state information into the experience pool every training cycle. High-yield state information can improve the rapid convergence of the DDQN model.

[0116] Step 6: Input the real-time information of the end of the manipulator, update the impedance parameters of the impedance controller, output the position correction value of the end of the manipulator, and complete the variable impedance control of the spatial manipulator shaft-hole assembly.

[0117] Common space robot arm assembly diagram Figure 5 As shown, in this embodiment, the RoboticToolbox in MATLAB and Python tensorflow2.0 are combined to realize the simulation of the impedance control of the end of the robot arm.

[0118] Created using the Robotic Toolbox Figure 6 UR5 robotic arm simulation environment.

[0119] When the robot arm is hindered by the environment while moving on the desired trajectory, the robot arm will generate an interaction force with the external environment because the environment is generally rigid. , the force / position relationship between the robot arm and the environment can be regarded as a spring model, as shown in the following equation.

[0120] (6)

[0121] in represents the environmental stiffness, Indicates the environmental position offset. It is 500N / m.

[0122] In the simulation, set the robot arm to move downward along the Z axis and set the desired force to , , ]=[0,0,15N], that is, only the force information in the Z-axis direction is considered, and the initial state of the end of the robot arm is set to [x,v,a]=[0,-0.5m / s,0].

[0123] Set simulation time , simulation cycle , expected position is 0.2m, the final simulation results are as follows Figure 7 As shown, Figure 7 It is the optimal parameter table of impedance control selected after reinforcement learning. The simulation is divided into three stages, such as Figure 7 As shown,

[0124] 1) In the first stage, there is a large error between the end of the robot arm and the desired position, so a high stiffness and high damping strategy is selected to make the impedance controller respond quickly.

[0125] 2) In the second stage, after the robot arm reaches the target plane, it adopts a strategy of reducing stiffness, and through this method, the force error of the system is gradually reduced.

[0126] 3) In the third stage, the overshoot at the end of the robotic arm is 0. It can be seen that at this time, due to the low position and speed of the robotic arm end, a low stiffness and low damping strategy is adopted, which makes the static force error of the system approach 0.

[0127] Figure 8 、 Figure 10 This is a control effect diagram of the end position and static force of the robot arm. By comparing the two, it can be found that the variable impedance method proposed in the present invention can respond faster to the force error at the initial moment, and at the same time can track the target force while ensuring a small overshoot. The static error is smaller than that of traditional fixed impedance control.

[0128] Figure 9 This is a tracking simulation diagram of the speed of the end of the robot arm. In the simulation results, the speed of the end of the robot arm can be well stabilized within the set threshold, that is, |v|<0.2. At the same time, the speed of the invented method can reach 0 faster than the traditional method.

[0129] Figure 11 This is a tracking curve diagram of the dynamic force of the end of the robotic arm. The variable impedance control proposed in the present invention has a smaller dynamic error and a faster response speed for tracking dynamic force than traditional impedance control, and has better tracking accuracy than traditional fixed impedance control.

[0130] The above embodiments illustrate and describe the basic principles and main features of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention, and such changes and modifications fall within the scope of the invention as claimed.

Claims

1. A variable impedance control method for shaft-hole assembly of a space manipulator based on reinforcement learning, characterized in that: The following steps are involved: Step 1: Construct a spatial manipulator model based on the DH parameter method; Step 2: Construct a transformation model of the joint angle state and end position of the spatial manipulator based on the forward and inverse kinematics algorithm; Step 3: Initialize the internal and external parameters of the binocular camera, and use the binocular camera to capture images to obtain the position information of the assembly holes; Step 4: Build an impedance controller based on reinforcement learning and set the impedance parameter action table, reward function, and termination conditions during training according to the expected goals: Step 4-1. Build an impedance controller: ; in, , They represent the actual motion trajectory and the expected motion trajectory of the end of the space manipulator, Represents the force between the end of the robotic arm and the external environment, They correspond to the desired inertia matrix, desired stiffness matrix and desired damping matrix of the impedance controller respectively; 、 They represent the actual acceleration, expected acceleration, actual speed and expected speed of the end of the space manipulator respectively. The impedance controller is selected As a control quantity; Step 4-2, set the reward function: ; ; Among them, T represents the duration of a single training session. is the error between the expected force and the force at the current moment; Step 4-3: Set the impedance parameter action table for reinforcement learning: ; is the correction amount set; Step 5: Train the impedance controller based on the neural network: First, the total number of trainings is set. In a single training run, the experience table of the space manipulator is collected and placed in the experience pool. The policy network is updated by the residual between the predicted value in the policy network and the target network. The update time is set. Once the update time is exceeded, the target network is replaced with the policy network to achieve the target network update. Finally, the target network updates the feedback in the environment and outputs the action with the highest score. This cycle is repeated until the total number of trainings is greater than the set value, and the training ends. The policy network predicts the Q value of the current moment in the impedance controller based on reinforcement learning, and predicts the Q value of the next moment in the impedance controller based on reinforcement learning based on the target network, and uses the mean square error of the difference between the two moments as the loss function: ; in, represents the mean square error, ) represents the Q value at time t, represents the decay rate during the learning process; The strategy network adopts a full neural network structure, takes the position, velocity, acceleration and force error value of the end of the robot arm as the network input, sets the number of hidden layer neurons to 400, selects the ReLU function as the activation function, and outputs the Q value of each action at the current moment; The target network adopts a full neural network structure, takes the position, velocity, acceleration and force error information of the end of the manipulator as the network input, sets the number of hidden layer neurons to 400, selects the ReLU function as the activation function, and outputs the Q value of each action at the next moment; Step 6: Input the real-time information of the end of the manipulator, update the impedance parameters of the impedance controller, output the position correction value of the end of the manipulator, and complete the variable impedance control of the spatial manipulator shaft-hole assembly.

2. The variable impedance control method for shaft-hole assembly of a space manipulator based on reinforcement learning according to claim 1 is characterized in that: The termination condition is set to: The number of training times reaches the set threshold.

3. A variable impedance control system for shaft-hole assembly of a space manipulator based on reinforcement learning, used to execute the method of claim 1, characterized in that: Includes the following modules: Space manipulator model construction module: used to construct the space manipulator model based on the DH parameter method; End-position conversion model construction module: used to construct the conversion model of the joint angle state and end-position of the spatial manipulator based on the forward and inverse kinematics algorithm; Assembly hole position information acquisition module: used to initialize the internal and external parameters of the binocular camera, and use the binocular camera to capture images to obtain the position information of the assembly holes; Impedance controller building module: used to build an impedance controller based on reinforcement learning and set the impedance parameter action table, reward function, and termination conditions during training according to the expected goals; Training module: training impedance controller based on neural network; Variable impedance control module for shaft-hole assembly of space manipulator: used to input real-time information of the end of the manipulator, update the impedance parameters of the impedance controller, output the position correction value of the end of the manipulator, and complete the variable impedance control of the shaft-hole assembly of the space manipulator.

4. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 2 are implemented.

5. A computer storable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 2 are implemented.