Space rope-driven mechanical arm control method based on model-free reinforcement learning
Through the model-free reinforcement learning algorithm, the simulation environment is constructed and the dense reward function and SAC algorithm are combined, which solves the generalization and floating base stability problems of traditional control methods in the space rope-driven robot arm, achieving efficient and complex task execution.
Patent Information
- Application Number
- CN202510663798.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-09-02
AI Technical Summary
Traditional control methods have weak generalization in space rope drive robot arms, rope tension changes are difficult to quantify, and floating bases in weightless environments increases modeling difficulty, resulting in a decrease in control effect.
A model-free reinforcement learning algorithm is adopted to construct a multi-joint contact dynamics simulation environment, combining dense reward function and SAC algorithm, and combining cost function estimation and target strategy smoothing, to achieve control of the space rope drive robot arm.
It improves the generalization of the strategy, enhances the ability of the agent to perform complex tasks, stabilizes the position of the floating base, and weakens the impact of rope tension changes on control.
Smart Images

Figure CN120572518A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robotic arms, and in particular to a control method for a space rope-driven robotic arm based on model-free reinforcement learning. Background Art
[0002] With the advancement of science and technology, aerospace and aviation have become important areas of demonstrating a nation's strength. The number of spacecraft in space has increased dramatically, driven by initiatives such as China's Beidou satellite navigation system and SpaceX's Starlink communications system. Regular inspection and maintenance are essential to ensuring the smooth operation of spacecraft, and the removal of failed spacecraft is crucial to ensuring a safe space environment and the overall performance of the Starlink network. Because spacecraft often have complex structures and limited space for operations, conventional manipulators often struggle to perform these tasks due to their limited physical limitations. A tethered space manipulator arm is a multi-jointed rigid robot that is actuated by cable tension rather than direct motor control of joint angles. This robot is actuated by servo motors, which control the rotation angles of the joints by controlling the tension in the tethers. Tethered space manipulator arms offer advantages such as high flexibility, low cost, lightweight construction, and the ability to carry large payloads, making them suitable for space missions such as satellite servicing and space debris removal.
[0003] Over the past decade, China has systematically designed the mechanical structure of rope-driven manipulators, evolving from discrete joints to coupled linkage within the manipulator segments, and evolving from aerospace applications to civilian applications such as charging stations. The relevant kinematic and dynamic modeling has also been completed. Building on the mechanical prototype, researchers have continued their research on motion planning and control methods, primarily including traditional control methods such as proportional-derivative control, dynamic feedforward control, and variable structure sliding membrane control, and developed a corresponding teleoperation system. While these control methods offer high precision and stable control, they also present challenges. First, because traditional control methods are often based on the kinematic and dynamic models of rope-driven manipulators, changes in the mechanical structure can significantly alter model parameters. Control strategies must be adjusted accordingly, resulting in limited generalization. Second, because ropes are non-rigid objects, deformation during operation causes changes in their internal tension. Furthermore, during motion, elastic and frictional forces are generated between the rope and the mechanical structure, affecting the rope's pulling action. These effects are difficult to quantify, leading to errors in the kinematic and dynamic models and, consequently, reduced control effectiveness. Third, since the space robot is located in a weightless environment, the floating base increases the difficulty of modeling the system. When the complexity of the task increases, the control difficulty increases significantly, resulting in performance degradation.
[0004] The disclosure of the above background technology content is only used to assist in understanding the concept and technical solution of the present invention. It does not necessarily belong to the prior art of this patent application. In the absence of clear evidence that the above content has been disclosed on the filing date of this patent application, the above background technology should not be used to evaluate the novelty and creativity of this application. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a control method for a spatial rope-driven robotic arm based on model-free reinforcement learning, which can improve the generalization of the strategy and enhance the ability of the intelligent agent to perform complex tasks.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention discloses a control method for a space rope-driven manipulator based on model-free reinforcement learning, comprising the following steps:
[0008] S1: establishing a simulation environment for the space rope-driven manipulator using multi-joint contact dynamics, wherein the space rope-driven manipulator includes a floating base and a rope-driven manipulator arm, a first end of the rope-driven manipulator arm is connected to the floating base, and the other end of the rope-driven manipulator arm is an end point of the space rope-driven manipulator arm;
[0009] S2: Construct a reinforcement learning framework based on a dense reward function and design a reward function to control the space rope-driven manipulator, wherein the reward function is related to the distance between the end point of the space rope-driven manipulator and the target point at each moment.
[0010] Preferably, the rope-driven robotic arm comprises a plurality of universal joints, and each of the universal joints is driven by four tendon modules.
[0011] Preferably, when establishing a simulation environment for the space rope-driven manipulator using multi-joint contact dynamics, the space rope-driven manipulator is defined through an XML file, and the space rope-driven manipulator is encapsulated as a simulation environment through a Python function package.
[0012] Preferably, the reinforcement learning framework based on dense reward function is the SAC algorithm.
[0013] Preferably, the reinforcement learning framework based on dense reward function is constructed as follows:
[0014]
[0015] In the formula, r(s t , a t ) represents the original reward of reinforcement learning, Represents the state s at the next moment t+1, H(·) represents the entropy term, π(·) represents the strategy, α represents the temperature coefficient, r soft (s t , a t ) represents the reinforcement learning reward with the addition of entropy;
[0016] Among them, the loss function of the policy network in the reinforcement learning framework is:
[0017]
[0018] Where, target entropy represents the action space, Indicates that the experience replay pool Extract the state s at time t t , a t ~π is a t ~π φ (·|s t ), a t ~π φ (·|s t ) represents action a at time t t The satisfied state is s t When φ is the parameter of strategy π φ The probability distribution of the output.
[0019] Preferably, the reward function is:
[0020] r t =r l,t +r small,t +r episode,t ·
[0021] Where r l,t =-k l l t is the distance bonus item, k l is the distance weight, l t represents the distance between the end point of the space rope-driven manipulator and the target point at time t; r small,t =-k s ln(l t +∈) is the logarithmic distance reward term, k s is the logarithmic distance weight, ∈ is a numerical parameter; r episode,t This is a round distance bonus.
[0022] Preferably, the round distance reward item r episode,t The value of is:
[0023]
[0024] In the formula, done = True indicates the number of steps n that the agent performs the task e,t Less than the maximum number of steps n in the specified round max , at this time the round distance reward item is positive, where k e is the round weight; in other cases, the round distance bonus is 0.
[0025] Preferably, step S2 further comprises: combining a reinforcement learning framework with cost function estimation to control the spatial rope-driven manipulator.
[0026] Preferably, the reinforcement learning framework is combined with the cost function estimation to obtain the constraint reinforcement learning optimization problem as follows:
[0027]
[0028] Where, J r (π φ ) represents the original optimization problem of reinforcement learning based on the reward function, π φ represents the strategy with φ as parameter, d i represents the cost threshold, C represents the number of constraints;
[0029] Among them, the cost function c i The constraint definition is as follows:
[0030]
[0031] In the formula, i represents different safety constraints, T represents the number of termination steps, and J Ci Denotes the cost function c i The expectation and d i Representative J Ci The cost threshold that needs to be met, the inequality formed by the two represents the cost function c i Constraints, s0 represents the state at the initial moment, a t ~π is a t ~π φ (·|s t ), a t ~π φ (·|s t ) represents action a at time t t The satisfied state is s t When φ is the parameter of strategy π φ The probability distribution of the output, s t+1 ~p represents the next moment state that conforms to the probability distribution p, s0~u represents the initial state that conforms to the initial state distribution u, It represents expectation.
[0032] The Lagrange multiplier method is used to transform the optimization problem into the corresponding dual problem, as follows:
[0033]
[0034] Where λ=[λ1…λ C ] T represents the Lagrange multiplier corresponding to the constraint.
[0035] Preferably, step S2 further comprises: smoothly combining a reinforcement learning framework with a target strategy to control the space rope-driven manipulator.
[0036] Preferably, the reinforcement learning framework is combined with the target policy smoothing as follows:
[0037]
[0038] Where a L , a H represent the lower and upper limits of the action space respectively, ∈ is the clipping noise; the clipping function clip(·) is defined as follows:
[0039]
[0040] Where a represents the action before target policy smoothing, a' represents the action after target policy smoothing, and s' represents the state at time t+1.
[0041] In a second aspect, the present invention discloses a computer-readable storage medium storing a computer program, wherein the computer program is configured to be executable by a processor to execute the control method of the spatial rope-driven robotic arm based on model-free reinforcement learning as described in the first aspect.
[0042] Compared with the existing technology, the beneficial effects of the present invention are: the control method of the space rope-driven manipulator based on model-free reinforcement learning proposed in the present invention avoids the traditional control method of system kinematic and dynamic modeling based on the model-free reinforcement learning algorithm; and the model-free reinforcement learning algorithm has little correlation with the structural parameters of the space rope-driven manipulator and has strong generalization; thereby, it can improve the generalization of the strategy and enhance the ability of the intelligent agent to perform complex tasks.
[0043] In a further embodiment, the present invention also has the following beneficial effects:
[0044] (1) The present invention uses MuJoCo to construct a simulation environment and uses the tendon module to simulate the rope, avoiding rope modeling and achieving good simulation results.
[0045] (2) The present invention combines SAC with cost function estimation, which effectively reduces the posture change of the floating base during the rope-driven manipulator's task execution.
[0046] (3) The present invention combines SAC with target strategy smoothing to remove abnormal values of the control quantity and further improve the stability of the floating base. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 1 is an overall flow chart of a method for controlling a space rope-driven manipulator based on model-free reinforcement learning in an embodiment of the present invention;
[0048] Figure 2 is a schematic diagram of a simulation environment in an embodiment of the present invention;
[0049] Figure 3 2. This is a schematic diagram of the posture change of the floating base of the space rope-driven manipulator in an embodiment of the present invention;
[0050] Figure 4 2 is a comparison chart of cost function values when different frameworks are applied in an embodiment of the present invention. DETAILED DESCRIPTION
[0051] The following is a detailed description of the embodiments of the present invention. It should be emphasized that the following description is only exemplary and is not intended to limit the scope of the present invention and its application.
[0052] It should be noted that when an element is referred to as being "fixed to" or "disposed on" another element, it can be directly on the other element or indirectly on the other element. When an element is referred to as being "connected to" another element, it can be directly connected to the other element or indirectly connected to the other element. In addition, connection can be used for both fixing and circuit / signal communication.
[0053] It should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the embodiments of the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limiting the present invention.
[0054] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0055] In response to the problems existing in the prior art, the present invention introduces RL (Reinforcement Learning) methods into the control of a space rope-driven manipulator, hoping to make up for the shortcomings of traditional control methods through artificial intelligence methods. Reinforcement learning is a type of machine learning, and its goal is to enable an intelligent agent to learn the optimal action execution strategy by interacting with the environment, so as to maximize the expected cumulative reward obtained during the interaction process. Reinforcement learning is an intelligent decision-making method that is applicable to MDP (Markov Decision Process). In particular, after the introduction of the deep reinforcement learning algorithm DQN (Deep Q-Network), its performance has been greatly improved.
[0056] This paper applies a model-free reinforcement learning algorithm to the control problem of a rope-driven space manipulator. To address the challenges of traditional control methods, such as the difficulty in modeling, the difficulty in quantifying rope tension changes, and the susceptibility of model parameters to perturbations, this paper builds a simulation environment within the simulation engine that does not include explicit kinematic and dynamic models. This algorithm, combined with a model-free reinforcement learning algorithm, achieves a "black box" effect, avoiding the complex system dynamics modeling and unstable rope modeling required for the rope-driven space manipulator. Through extensive training through the interaction of a large number of intelligent agents with the environment, the generalization of the strategy is improved, enhancing the agents' ability to perform complex tasks.
[0057] like Figure 1 As shown, the preferred embodiment of the present invention discloses a method for controlling a space rope-driven manipulator based on model-free reinforcement learning, comprising the following steps:
[0058] S1: A simulation environment was established using MuJoCo (Multi-Joint dynamics with Contact). MuJoCo defines models using XML files and encapsulates these models into a simulation environment using the Python package mujoco-py. The space-based tethered manipulator consists of a floating base and a tethered manipulator with three universal joints.
[0059] In the embodiment of the present invention, MuJoCo is used to establish a simulation environment. Figure 2As shown, the floating base 10 floats in the space environment, and the rope-driven manipulator 20 connected to the floating base 10 gradually approaches the target point 30 under the drive of the rope 21. MuJoCo can define the model through an XML file, and encapsulate the model into a simulation environment through mujoco-py for the intelligent agent to interact with it. The unique tendon module (tendon) in MuJoCo has properties such as stiffness coefficient and damping coefficient, which respectively determine the spring force and damping force acting on the tendon module. Its properties are similar to those of a rope and can be used to simulate the rope. The model of the space rope-driven manipulator includes a floating base 10 and a rope-driven manipulator 20, wherein the rope-driven manipulator 20 includes three universal joints, each of which is driven by 4 tendon modules. The structural diagram is shown as follows Figure 3 As shown in the figure, the floating base 10 is freely movable and has 6 degrees of freedom. Each universal joint has 2 degrees of freedom, so the model has a total of 12 degrees of freedom. The end of the three universal joints is provided with an end effector 22. Driven by the three universal joints 22, the end effector 22 gradually approaches the target point 30. The end effector 22 corresponds to the end point of the space rope-driven manipulator.
[0060] S2: Build a reinforcement learning framework based on dense reward function and design reward function;
[0061] Among them, the reward function of the reinforcement learning framework based on dense reward function is related to the distance between the end point of the spatial rope-driven robot arm and the target point at each moment, and the performance of the reinforcement learning agent is improved by shaping the reward function.
[0062] In order to accurately and comprehensively describe the kinematic performance of the space rope-driven manipulator, the motion of the floating base and the end effector are considered simultaneously in the state space.
[0063] Among them, the position and velocity information of the floating base are shown in formula (1),
[0064]
[0065] Where, and Represent the position and velocity information of the floating base, P b , V b Represent the position and velocity vector of the floating base mass center, H b ,Ω b Represent the attitude quaternion and angular velocity vector of the floating base respectively.
[0066] The state vector s of the space rope-driven manipulator at time t t As shown in formula (2):
[0067]
[0068] Where, P e , V e Represent the position and velocity vector of the end effector (end), P target Indicates the target point position; l indicates the distance between the end effector and the target point,
[0069] The motion vector a of the space rope-driven manipulator at time t t As shown in formula (3):
[0070]
[0071] Where, X i,t Represents the control quantities of the four cables corresponding to the i-th universal joint.
[0072] In order to improve the convergence speed of the algorithm and the learning effect of the agent, the reward function r t The settings are shown in formula (4):
[0073] r t =r l,t +r small,t +r episode,t (4)
[0074] Where r l,t =-k l l t is the distance bonus item, k l Represents the distance weight, guiding the end effector of the rope-driven manipulator to approach the target point, l t represents the distance between the end effector and the target point at time t; r small,t =-k s ln(l t +∈) is the logarithmic distance reward term, k s is the logarithmic distance weight, ∈ is a numerical parameter used to prevent infinite errors. When the distance is small, due to r l,t If it is too small, it will not be able to play a guiding role. After taking the logarithm, the reward value will increase and continue to promote the reduction of distance. episode,t It is a round distance reward item that guides the agent to complete the task in fewer time steps.
[0075] Specifically, when the number of steps n that the agent performs the task e,t Less than the maximum number of steps n in the specified round max When done=True, the round distance reward item is positive, where k e is the round weight; in other cases, the round distance reward is 0, as shown in formula (5):
[0076]
[0077] In this invention, DDPG (Deep Deterministic Policy Gradient) and SAC (Soft Actor-Critic) are selected to verify the success rate respectively, and SAC is finally selected as the reinforcement learning framework for the main application.
[0078] SAC maximizes both the reward value and the entropy of the action distribution, enhancing the randomness of the agent's exploration, as shown in formula (6):
[0079]
[0080] In the formula, r(s t , a t ) represents the original reward of reinforcement learning, Represents the state s at the next moment t+1 , H(·) represents the entropy term, π(·) represents the strategy, α represents the temperature coefficient, r soft (s t , a t ) represents the reinforcement learning reward with the addition of entropy. Therefore, the loss function of its policy network can be expressed as shown in formula (7):
[0081]
[0082] Where, represents the action-value function, Indicates that the experience replay pool Extract the state s at time t t , a t ~π φ (·|s t ) represents action a at time t t The satisfied state is s t When φ is the parameter of strategy π φ The probability distribution of the output can be simplified to a t ~π.
[0083] At the same time, SAC can automatically adjust the entropy regularization term coefficient, and the loss function is shown in Formula (8) to balance the value enhancement and the randomness of exploration.
[0084]
[0085] Where, target entropy Represents the action space.
[0086] In addition to the dense reward reinforcement learning framework, this paper integrates SAC (Soft Actor-Critic) and HER (Hindsight Experience Replay) to construct a sparse reward reinforcement learning framework (SAC-HER). The two frameworks are compared and the best one is selected as the framework for subsequent content.
[0087] Among them, the reward function of the reinforcement learning method based on sparse reward function is only related to whether the task is completed. This embodiment combines SAC with HER to improve the efficiency of experience utilization.
[0088] HER is a classic GoRL (Goal-oriented Reinforcement Learning). GoRL is suitable for tasks where the goal is constantly changing, such as goal approach tasks. It is applicable to augmented tuples. To define MDP (Markov Decision Process). Among them, represents the state space, represents the action space, P represents the state transition function, represents the reward function corresponding to the given target g, represents the target space, ψ(·) represents the transition from the state space To the target space The mapping function can obtain the corresponding achieved target through the state s, that is
[0089] Specifically, for the target reaching task, g = P target ,ψ(s)=P e The reward function of GoRL is a sparse reward function, which is composed of the current state s t 、Action a t and the established goal g, the sparse reward function used in the embodiment of the present invention is shown in formula (9):
[0090]
[0091] Where, δ g represents the threshold distance under the current target, and ||·||2 is used to solve the reached target ψ(s t+1 ) is the distance from the predetermined target g.
[0092] Similar to the reward function, GoRL’s action-value function Q(s, a, g) and policy π(a|s, g) are also based on the goal. In general, GoRL transforms the state vector s into t Expanded to vector (s t, g), and repeat the original interaction and update process in the original reinforcement learning framework.
[0093] HER is a classic GoRL algorithm that can be combined with frameworks such as DDPG and SAC to improve the learning effect of the GoRL algorithm.
[0094] Specifically, for the interactive trajectory {s1, s2, ..., s T} And ψ(s1), ψ(s2),…, ψ(s T )≠g, although each state has not reached the established goal g, but has reached its own goals ψ(s1), ψ(s2),…, ψ(s T ). By replacing the original established target g with the new target g′=ψ(s″) and recalculating the corresponding reward value, (s t , a t , s t+1 , r t,g , g) is corrected to (s t , a t , s t+1 , r t,ψ(s″) ,ψ(s″)), the strategy can learn valuable information from the failure experience and accelerate the training process. Regarding the selection of s″, this embodiment selects t On the same trajectory and at time s t A certain state after that is taken as s″.
[0095] In this embodiment of the present invention, SAC and SAC-HER are respectively applied to perform the target point arrival task, and the task success rates of the two are compared. Finally, the original SAC algorithm with a higher success rate is selected as the framework of SSAC.
[0096] S3: Combine the framework selected in S2 with the cost function estimation to weaken the position change of the floating base; then combine it with the target strategy smoothing to remove outliers in the control volume.
[0097] The cost function is based on the base's pose changes. This introduces constraints into the original reinforcement learning optimization problem, limiting the base's pose changes. Furthermore, a clipping function prevents the control variable from being too large or too small, thereby eliminating outliers in the control variable.
[0098] The embodiment of the present invention combines SAC with cost function estimation to weaken the attitude change of the free-floating base, stabilize the floating base, and construct CMDP (Constrained Markov Decision Process, constrained Markov decision process), that is, in, Represents the upper limit set of the cost function, that is, the cost function value μ(·) represents the initial state distribution; γ∈(0,1) represents the discount coefficient. For the cost function c i The constraint definition of is shown in formula (10):
[0099]
[0100] Where i represents different security constraints, d i represents the cost threshold, T represents the number of termination steps, C represents the number of constraints, and J Ci Denotes the cost function c i The expectation and d i Representative J Ci The cost threshold that needs to be met, the inequality formed by the two represents the cost function c i Constraints, s0 represents the state at the initial moment, a t ~π is a t ~π φ (·|s t ), a t ~π φ (·|s t ) represents action a at time t t The satisfied state is s t When φ is the parameter of strategy π φ The probability distribution of the output, s t+1 ~p represents the next moment state that conforms to the probability distribution p, s0~u represents the initial state that conforms to the initial state distribution u, It represents expectation.
[0101] Based on this, the constraint-based reinforcement learning optimization problem can be obtained as shown in formula (11):
[0102]
[0103] Where, J r (π φ ) represents the original reward-based optimization problem of reinforcement learning.
[0104] The Lagrange multiplier method is used to transform the optimization problem into the corresponding dual problem, as shown in formula (12):
[0105]
[0106] Where λ=[λ1…λ C ] T represents the Lagrange multiplier corresponding to the constraint.
[0107] Since the floating base of the space rope-driven manipulator can float freely, the position of the floating base will change during the space rope-driven manipulator performs tasks, such as Figure 3 To ensure the stability of the floating base, the cost function is defined as the change in the floating base posture vector as shown in formula (13):
[0108]
[0109] Where k c is the cost weight, represents the position of the floating base at time t, Indicates the position of the floating base at the initial moment.
[0110] The key to solving the dual problem (12) is to update the Lagrange multiplier during the optimization process. The embodiment of the present invention adopts the PID (Proportional Integral Differential) update method as shown in (14):
[0111]
[0112] Where k P , k I , k D Represent the proportional, integral and differential weights respectively; (·) + ReLU (Rectified Linear Unit) function, which is used to limit the growth of the cost function value without limiting its decrease; d represents the cost threshold; Represents the time t with φ t PID is a parameterized strategy. Specifically, the proportional term damps the response to damped oscillations; the differential term prevents overshoot of the cost value and limits the cost growth rate to within the feasible region; and the integral term eliminates steady-state disturbances and promotes convergence. Overall, the PID method is powerful and easy to implement.
[0113] The embodiment of the present invention applies target policy smoothing in the exploration phase and the update phase to prevent the policy network from utilizing erroneous spikes from the value network, as shown in formula (15):
[0114]
[0115] Where a L ,a H Represent the upper and lower limits of the action space, and ∈ is the clipping noise. The clipping function clip(·) is defined as shown in formula (16):
[0116]
[0117] Where a represents the action before target policy smoothing, a' represents the action after target policy smoothing, and s' represents the state at time t+1.
[0118] Subsequent validity verification work showed that the control effect of SAC is stronger than that of SAC-HER. Therefore, the original SAC is selected as the framework in the present invention, and combined with cost function estimation and target policy smoothing, the specific algorithm of the embodiment of the present invention, SSAC (Smooth Soft Actor-Critic) algorithm, is obtained to realize the control of the space rope-driven manipulator based on model-free reinforcement learning.
[0119] The preferred embodiment of the present invention combines the SAC (Soft Actor-Critic algorithm) with cost function estimation and target strategy smoothing to form the SSAC (Smooth Soft Actor-Critic algorithm). This effectively solves the base floating problem and control variable outlier problem of the space rope-driven manipulator without considering the dynamic characteristics of the floating base.
[0120] The following verifies the effectiveness of the spatial rope-driven manipulator control method based on model-free reinforcement learning proposed in the preferred embodiment of the present invention.
[0121] The control task completed by the embodiment of the present invention is to drive the rope to change the arm shape of the rope-driven manipulator so that the end effector of the rope-driven manipulator reaches the target point. First, the task success rates of different dense reward functions are compared. In order to compare the control performance of DDPG and SAC and design the optimal dense reward function, the embodiment of the present invention adjusts the reward term r l,t and r small,t The weight k l and k s Several comparative experiments were conducted (due to episode,t The weight k e The adjustment of has little effect on the success rate, so it is not explained here. If the distance between the agent and the target point is less than the threshold within the maximum number of rounds defined by the environment, the task is considered successful. During the training phase, 500 points are taken from the workspace of the space rope-driven manipulator as the target point set. In each round, one point is randomly selected from this set as the target point for the experiment. Each training session interacts with the environment for 1×10 6 step; in the testing phase, each time a point is randomly selected from the workspace as the target point, a total of 100 rounds of testing are performed, and the success rate is calculated.
[0122] Table 1 is a comparison table of the success rates of target reaching tasks performed by reinforcement learning methods using different forms of dense reward functions in an embodiment of the present invention.
[0123] Table 1
[0124]
[0125]
[0126] As can be seen from Table 1, the success rate of SAC is higher than that of DDPG under most weight combinations. l =10,k s When ∈ R=0.25, SAC achieves approximately 6 times the success rate of DDPG in the test phase. This result shows that the action distribution entropy introduced by SAC improves the randomness of the strategy, enhances the exploration ability, and has stronger generalization.
[0127] In order to select the better one between SAC and SAC-HER as the framework of SSAC, a comparison experiment of the success rate of sparse reward reinforcement learning and dense reward reinforcement learning was conducted. The weight combination of DDPG and SAC was selected as k l =10,k s =0.25. Table 2 shows a comparison of the success rates of target reaching tasks using the reinforcement learning method with dense reward function and sparse reward function in an embodiment of the present invention.
[0128] Table 2
[0129]
[0130] From Table 2, we can see that no matter the number of training steps is 1×10 6 or 2×10 6 , SAC achieved higher success rates than SAC-HER in both training and testing phases. This result demonstrates that dense reward functions contain richer information and offer superior control performance compared to sparse reward functions. Therefore, this embodiment of the present invention selected SAC as the framework for the final SSAC algorithm.
[0131] Secondly, the cost function value changes of different methods during the training phase are compared. In order to explore the optimal method to reduce the change of the floating base posture, the embodiment of the present invention compares DDPG, constrained SAC ( Figure 4 SAC) and SSAC ( Figure 4 The change of the cost function value during the training of Smooth SAC (DDPG), in which DDPG does not use the cost function for constraint, while SAC and Smooth SAC both use the cost function for constraint. Figure 4 As shown, when the number of training steps is 0.3×10 6 When the training step is 0.3×10 6When the training step is 0.1×10 6 When , the cost function value of Smooth SAC begins to drop rapidly and remains at a low level. This result shows that although the cost function constraint can suppress the growth of the cost value, the sudden change and spike of the tension in the rope will aggravate the instability of the floating base, resulting in a rebound of the cost value in the later stage of training (such as Figure 4 (As shown in the SAC curve in the figure), applying the target policy smoothing suppresses abnormal movements, causing the cost to decrease earlier and remain low, thus suppressing the base change of the space rope-driven manipulator. This comparison demonstrates the effectiveness of the proposed SSAC algorithm, thereby establishing a control method for a space rope-driven manipulator based on model-free reinforcement learning.
[0132] Another preferred embodiment of the present invention discloses a computer-readable storage medium, which stores a computer program, wherein the computer program is configured to be executable by a processor to execute the control method of the spatial rope-driven robotic arm based on model-free reinforcement learning in the above preferred embodiment.
[0133] Optionally, the storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store computer programs.
[0134] The background section of the present invention may contain background information about the problem or environment of the present invention rather than describing prior art by others. Therefore, the inclusion of content in the background section is not an admission by the applicant that the prior art is available.
[0135] The above description further details the present invention in conjunction with specific / preferred embodiments, and the specific implementation of the present invention should not be construed as being limited to these descriptions. Persons skilled in the art will appreciate that, without departing from the spirit of the present invention, they may make various substitutions or modifications to the described embodiments, and these substitutions or modifications should be considered to fall within the scope of protection of the present invention. Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "preferred embodiments," "examples," "specific examples," or "some examples" indicates that the specific features, structures, materials, or characteristics described in conjunction with such embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. Furthermore, those skilled in the art may combine and assemble the different embodiments or examples described in this specification, as well as features from different embodiments or examples, without conflicting opinions. Although the embodiments of the present invention and their advantages have been described in detail, it should be understood that various changes, substitutions, and modifications may be made herein without departing from the scope of the appended claims.
Claims
1. A control method for a space rope-driven manipulator based on model-free reinforcement learning, characterized in that: The following steps are involved: S1: establishing a simulation environment for the space rope-driven manipulator using multi-joint contact dynamics, wherein the space rope-driven manipulator includes a floating base and a rope-driven manipulator arm, a first end of the rope-driven manipulator arm is connected to the floating base, and the other end of the rope-driven manipulator arm is an end point of the space rope-driven manipulator arm; S2: Construct a reinforcement learning framework based on a dense reward function and design a reward function to control the space rope-driven manipulator, wherein the reward function is related to the distance between the end point of the space rope-driven manipulator and the target point at each moment.
2. The control method of a space rope-driven manipulator based on model-free reinforcement learning according to claim 1 is characterized in that: The rope-driven robotic arm includes a plurality of universal joints, and each of the universal joints is driven by four tendon modules.
3. The control method of a space rope-driven manipulator based on model-free reinforcement learning according to claim 1, characterized in that: When a simulation environment is established for the space rope-driven manipulator using multi-joint contact dynamics, the space rope-driven manipulator is defined through an XML file, and the space rope-driven manipulator is encapsulated as a simulation environment through a Python function package.
4. The control method of a space rope-driven manipulator based on model-free reinforcement learning according to claim 1, characterized in that: The reinforcement learning framework based on dense reward function is the SAC algorithm.
5. The control method of a space rope-driven manipulator based on model-free reinforcement learning according to claim 4 is characterized in that: The specific steps to build a reinforcement learning framework based on dense reward function are: In the formula, r(s t , a t ) represents the original reward of reinforcement learning, Represents the state s at the next moment t+1 , H(·) represents the entropy term, π(·) represents the strategy, α represents the temperature coefficient, r soft (s t , a t ) represents the reinforcement learning reward with the addition of entropy; Among them, the loss function of the policy network in the reinforcement learning framework is: Where, target entropy represents the action space, Indicates that the experience replay pool Extract the state s at time t t , a t ~π is a t ~π φ (·|s t ), a t ~π φ (·|s t ) represents action a at time t t The satisfied state is s t When φ is the parameter of strategy π φ The probability distribution of the output.
6. The control method of a space rope-driven manipulator based on model-free reinforcement learning according to claim 1, characterized in that: The reward function is: r t =r l,t +r small,t +r episode,t Where r l,t =-k l l t is the distance bonus item, k l is the distance weight, l t represents the distance between the end point of the space rope-driven manipulator and the target point at time t; r small,t =-k s ln(l t +∈) is the logarithmic distance reward term, k s is the logarithmic distance weight, ∈ is a numerical parameter; r episode,t This is a round distance bonus.
7. The control method of a space rope-driven manipulator based on model-free reinforcement learning according to claim 6, characterized in that: Round distance bonus item r episode,t The value of is: In the formula, done = True indicates the number of steps n that the agent performs the task e,t Less than the maximum number of steps n in the specified round max , at this time the round distance reward item is positive, where k e is the round weight; in other cases, the round distance bonus is 0.
8. The control method of a space rope-driven manipulator based on model-free reinforcement learning according to claim 1, characterized in that: Step S2 further includes: combining a reinforcement learning framework with a cost function estimation to control the spatial rope-driven manipulator; Preferably, the reinforcement learning framework is combined with the cost function estimation to obtain the constraint reinforcement learning optimization problem as follows: Where, J r (π φ ) represents the original optimization problem of reinforcement learning based on the reward function, π φ represents the strategy with φ as parameter, d i represents the cost threshold, C represents the number of constraints; Among them, the cost function c i The constraint definition is as follows: In the formula, i represents different safety constraints, T represents the number of termination steps, and J Ci Denotes the cost function c i The expectation and d i Representative J Ci The cost threshold that needs to be met, the inequality formed by the two represents the cost function c i Constraints, s0 represents the state at the initial moment, a t ~π is a t ~π φ (·|s t ), a t ~π φ (·|s t ) represents action a at time t t The satisfied state is s t When φ is the parameter of strategy π φ The probability distribution of the output, s t+1 ~p represents the next moment state that conforms to the probability distribution p, s0~u represents the initial state that conforms to the initial state distribution u, It represents expectation. The Lagrange multiplier method is used to transform the optimization problem into the corresponding dual problem, as follows: Where λ=[λ1…λ C ] T represents the Lagrange multiplier corresponding to the constraint.
9. The control method of a space rope-driven manipulator based on model-free reinforcement learning according to claim 1, characterized in that: Step S2 further includes: smoothly combining a reinforcement learning framework with a target strategy to control the space rope-driven manipulator; Preferably, the reinforcement learning framework is combined with the target policy smoothing as follows: Where a L , a H represent the lower and upper limits of the action space respectively, ∈ is the clipping noise; the clipping function clip(·) is defined as follows: Where a represents the action before target policy smoothing, a' represents the action after target policy smoothing, and s' represents the state at time t+1.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program is configured to be executable by a processor to execute the control method of a space rope-driven robotic arm based on model-free reinforcement learning according to any one of claims 1 to 9.
Citation Information
Cited By
Control method of space continuous type rope driving arm
CN120791797A