Space robot trajectory planning method based on deep reinforcement learning
By using a trajectory planning method based on deep reinforcement learning, the problems of high computational cost and dynamic singularity in space robot trajectory planning are solved, spacecraft attitude disturbances are reduced, robot accuracy is guaranteed, and it is suitable for spacecraft with limited computing resources.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG INST OF BUSINESS & TECH
- Filing Date
- 2026-03-12
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies for space robot trajectory planning suffer from problems such as high computational load, difficulty in escaping local optima, and dynamic singularities. Furthermore, traditional methods sacrifice the attitude accuracy of the spacecraft itself or are limited by computational resources.
A trajectory planning method based on deep reinforcement learning is adopted. By designing a reward function to consider the spacecraft's attitude perturbation factors, and combining intelligent decision network and traditional trapezoidal planning, the motion process is decomposed to generate a trajectory from the initial pose to the desired pose.
It reduces the attitude disturbance of the spacecraft, avoids large computational loads, handles dynamic singularity problems, and ensures the accuracy of the space robot's end effector.
Smart Images

Figure CN121870773A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of space robot technology, specifically relating to a space robot trajectory planning method based on deep reinforcement learning. Background Technology
[0002] For free-floating space robot systems, there is a dynamic coupling between the robot and the spacecraft body; that is, the robot's motion changes the spacecraft's attitude. However, some space missions (such as communication and observation) have certain requirements for the spacecraft's attitude. Therefore, trajectory planning should minimize the disturbance to the spacecraft's attitude caused by the robot's motion. Although trajectory planning methods based on traditional optimization algorithms can obtain relatively accurate and ideal results, they face problems such as huge computational costs and difficulty in escaping local optima, making them difficult to apply to spacecraft with severely limited computational resources.
[0003] Furthermore, trajectory planning in mission space faces the problem of dynamic singularities. While existing damped minimum variance methods are simple and easy to implement, they come at the cost of sacrificing the accuracy of the space robot. Therefore, there is an urgent need for a space robot trajectory planning algorithm that reduces spacecraft attitude disturbances, avoids massive computational loads, handles dynamic singularities, and ensures the accuracy of the space robot. Summary of the Invention
[0004] To overcome the problems in the prior art, this invention proposes a spatial robot trajectory planning method based on deep reinforcement learning.
[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: This invention provides a spatial robot trajectory planning method based on deep reinforcement learning, comprising the following steps: Build a space robot simulation environment; The initial pose and desired pose of the space robot are set randomly; Based on deep reinforcement learning algorithms, a reward function is designed considering the attitude disturbance factors of the spacecraft itself. A stage target pose decision model is constructed and trained until the reward function converges, and the trained stage target pose decision model is obtained. The trained stage target pose decision model is loaded into the space robot system. Using the trained stage target pose decision model, the corresponding stage target pose is obtained according to the current state of the space robot. The trapezoidal programming method is used to form the trajectory from the current pose to the stage target pose. Repeat the above steps until the space robot reaches the desired pose, and finally generate the trajectory from the initial pose to the desired pose.
[0006] Furthermore, the construction of the space robot simulation environment includes: constructing a kinematic model of a free-floating space robot, wherein the mathematical model of the free-floating space robot includes a kinematic model and attitude relationships.
[0007] Furthermore, the construction of the space robot simulation environment also includes: building a trapezoidal programming method module based on the damped minimum variance method, which is used to generate the trajectory of the space robot from its current pose to the target pose of the stage.
[0008] Furthermore, the initial pose of the space robot is randomized, including: The arbitrary arm shape and spacecraft body posture within the workspace of the space robot are taken as the initial nominal state, and the randomized range of each joint angle is set with the initial nominal value as the center. If the initial arm shape after randomization is in a non-singular state, it is used as the initial pose of the space robot; otherwise, it is re-randomized until a non-singular initial pose is obtained.
[0009] Furthermore, based on deep reinforcement learning algorithms, a phase target pose decision model is constructed and trained, including: The design phase training of the target pose decision model includes the state space, action space, and reward function; the state space includes the current state of the space robot, the expected end-effector pose, and singular analysis; the action space is the target pose information of the space robot in future stages.
[0010] Furthermore, the reward function includes spacecraft attitude disturbance reward, dynamic singularity reward, space robot process accuracy reward, and space robot final accuracy reward.
[0011] Compared with the prior art, the present invention has the following technical effects: (1) This invention designs a space robot trajectory planning method based on deep reinforcement learning, which incorporates spacecraft body attitude perturbation factors as part of the reward function design, thereby reducing changes in spacecraft body attitude; (2) This invention adopts the idea of decomposing the motion process, combined with intelligent decision network model and traditional trapezoidal programming, and does not have a huge amount of computation, making it suitable for spacecraft with limited computing resources; (3) The trajectory planning method designed in this invention, compared with the traditional damped minimum variance method, can ensure the accuracy of the end effector of the space robot while dealing with dynamic singularities. Attached Figure Description
[0012] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart of a spatial robot trajectory planning method based on deep reinforcement learning; Figure 2 Composition of a space robot system; Figure 3 Design for action decision networks; Figure 4 This refers to the change in the reward function during the training process of this invention; Figure 5 This is a comparison chart of the errors of the present invention and the conventional method. Detailed Implementation
[0014] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the specific implementation methods, structures, features, and effects of the technical solutions proposed according to the present invention are described in detail below with reference to the accompanying drawings and preferred embodiments. Specific features, structures, or characteristics in one or more embodiments may be combined in any suitable form. Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0015] To address the trajectory planning problem of free-floating space robots, which involves spacecraft attitude perturbations and dynamic singularities, a deep reinforcement learning-based trajectory planning method for space robots is proposed. (1) The trajectory planned by this method can be described as the trajectory of the space robot end effector from its initial pose. Successively pass through the target pose of each stage Finally, the desired pose is reached. .
[0016] Based on deep reinforcement learning algorithms, a stage-specific target pose decision model is constructed, with a focus on spacecraft attitude perturbation when designing the reward function. This target pose decision model determines the target pose based on the current state of the space robot system (end-effector pose). By doing so, the target pose for the future stage can be obtained. In this embodiment, , .
[0017] (2) Initial pose of each stage Transfer to target pose At that time, a trapezoidal programming method based on the damped minimum variance method is adopted. Therefore, when the set of decisions for all process target pose sequences is obtained... At that time, the planned trajectory can be determined.
[0018] In this embodiment, refer to Figures 1-5 This paper presents a spatial robot trajectory planning method based on deep reinforcement learning, which includes the following steps: Build a space robot simulation environment; The initial pose and desired pose of the space robot are set randomly; Based on deep reinforcement learning algorithms, a reward function is designed considering the attitude disturbance factors of the spacecraft itself. A stage target pose decision model is constructed and trained until the reward function converges, and the trained stage target pose decision model is obtained. The trained stage target pose decision model is loaded into the space robot system. Using the trained stage target pose decision model, the corresponding stage target pose is obtained according to the current state of the space robot. The trapezoidal programming method is used to form the trajectory from the current pose to the stage target pose. Repeat the above steps until the space robot reaches the desired pose, and finally generate the trajectory from the initial pose to the desired pose.
[0019] The following is a detailed explanation of each of the above steps: Step 100: Build a space robot simulation environment, including constructing a kinematic model of a free-floating space robot. The mathematical model of the free-floating space robot includes a kinematic model and attitude relationships.
[0020] This invention is mainly aimed at the spacecraft body and n This invention relates to a space robot system composed of robots with multiple degrees of freedom. When the space robot operates in a free-floating mode, it satisfies the conservation of linear and angular momentum. Based on the characteristics of various posture description methods, this invention utilizes unit quaternions to design a trajectory planning algorithm, employing Euler angles to describe the problem and present the results. The kinematic model of the free-floating space robot constructed accordingly is an important component of the subsequent deep reinforcement learning training environment.
[0021] Step 110: The kinematic model of the space robot is as follows: ; The above formula is denoted as Formula 1, where, Represents the position vector of the end effector of a space robot. The derivative with respect to time t; This represents the angular velocity vector of the end effector of the space robot. The Jacobian matrix represents the motion of the spacecraft itself. The linear velocity vector of the spacecraft body, i.e. the time derivative of the spacecraft body's position, describes the speed and direction of the spacecraft body's movement in space. This represents the angular velocity vector of the spacecraft body, reflecting the rate and direction of the spacecraft body's attitude change; Represents the Jacobian matrix associated with robot motion; Represents the joint angle vector of a space robot. , Indicates the space robot's first i The joint angles of each joint. , Indicates the number of joints.
[0022] For a space robot in free-floating mode, the linear momentum and angular momentum of the system are conserved when there are no external forces or torques. Assuming the initial linear momentum and angular momentum of the system are zero, the velocity and angular velocity of the spacecraft body can be expressed as joint angular velocities. This means we get Equation 2: ; In the above formula, The Jacobian matrix represents the spacecraft body—the robot.
[0023] Substituting Equation 2 into Equation 1, we obtain the kinematic equations for a freely floating space robot: ; In the above formula, The attitude of the spacecraft itself; Indicates the first i The mass of each link; Indicates the first i Moment of inertia of each link; The generalized Jacobian matrix, known as the space robot matrix, is related not only to the kinematic parameters of the space robot but also to its dynamic parameters. Therefore, the singularity of the space robot is called "dynamic singularity".
[0024] Step 120: Use unit quaternions for trajectory planning, and use Euler angles to describe the problem and present the results.
[0025] Euler angles are intuitive and easy to understand when describing attitude, but they suffer from singularity issues. The unit quaternion method avoids singularity, is linear, computationally efficient, and has simple equations. Therefore, unit quaternions are primarily used in trajectory planning algorithm research, while Euler angles are used to describe the problem and present the results.
[0026] The attitude is represented by a unit quaternion, i.e.: ; In the above formula, A unit quaternion representing attitude; It is a rotation axis, also known as an Euler axis; It represents the angle rotated around the axis of rotation. This is the scalar part of the quaternion. This is the vector part of the quaternion. (Constraint) , At this point, the quaternion is unique. The relationship between the derivative of the quaternion with respect to time and the angular velocity is: ; In the above formula, , They represent respectively and The derivative with respect to time represents the rate of change of a quaternion over time; express Transpose of; Represents the cross product matrix; Represents the identity matrix; This represents the angular velocity vector.
[0027] Euler angles in rotational order zyx The conversion relationship between quaternions and unit quaternions is shown below: ; ; In this embodiment, a kinematic model of a free-floating space robot is built using 6 degrees of freedom as an example. .
[0028] Step 200: Build a space robot simulation environment, including constructing a trapezoidal programming method module based on the damped minimum variance method to generate the trajectory of the space robot from its current pose to the target pose of the stage.
[0029] As an example, this step includes: Step 210: Plan the trajectory of the space robot, including its position, attitude, velocity, and angular velocity.
[0030] The current attitude of the space robot's end effector, using the inertial coordinate system as the reference frame. for: ; In the above formula, , , This represents the three Euler angles at the current moment, with zyx as the rotation order.
[0031] Target Phase Pose for: ; In the above formula, , , These represent the three Euler angles corresponding to the target pose of the stage, with zyx as the rotation order.
[0032] Based on the space robot from the current position Current posture Continuous and smooth movement to the target stage position Target phase pose Corresponding rotation matrix: ; In the above formula, This represents the rotation matrix corresponding to the current pose; This represents the rotation matrix corresponding to the target stage pose; Indicates circling Axis rotation The rotation matrix corresponding to the angle; Indicates circling Axis rotation The rotation matrix corresponding to the angle; Indicates circling Axis rotation The rotation matrix corresponding to the angle; Indicates circling Axis rotation The rotation matrix corresponding to the angle; Indicates circling Axis rotation The rotation matrix corresponding to the angle; Indicates circling Axis rotation The rotation matrix corresponding to the angle; Representation matrix The first column; Representation matrix The second column; Representation matrix The third column; express The first column; express The second column; express The third column.
[0033] Calculate the pointing deviation of a space robot based on the column vectors of the rotation matrix. : ; in, ; The above formula represents the coordinate system of a space robot around... Rotation Afterwards, the posture changed from Adjust to .
[0034] Assuming the space robot moves along the straight line between its initial and final positions, plan the space robot's position and velocity, i.e.: ; ; In the above formula, This represents the position of the space robot at time t; This represents the initial position of the space robot's motion, i.e., t is the position of the space robot at the initial moment; This represents the position planning function, through which the smooth movement of the space robot from its initial position to the target stage position is controlled; This indicates the target position of the space robot's movement at a certain stage, which is the target position that the space robot is about to reach. This represents the velocity of the space robot at time t; express The derivative with respect to time t.
[0035] When planning the posture of a space robot, select an appropriate trajectory. Satisfying the initial time Termination time Similar to planning the position and speed of a space robot, that is: ; In the above formula, This represents the attitude information of the space robot at time t; This represents the attitude planning function, whose variations control the smooth transition of the space robot's attitude from the initial state to the target state; This represents the angle by which the attitude rotates from the initial state to the target state. The rotation axis (unit vector) represents the attitude from the initial state to the target state. Then the angular velocity of the space robot It can be represented as: ; In the above formula, express The derivative with respect to time t.
[0036] Trapezoidal programming is used to plan the rate of change of position of a space robot in real time. The rate of change of attitude of space robots in real time : First, normalization is performed, letting , , ,but: ; In the above formula, Indicates normalized time; This represents the actual time variable, indicating the real-time progress of the space robot during its movement; This represents the total time taken for a space robot to move from its initial state to its target stage state. Represents relative to normalized time The attitude planning function; Represents the normalized time Location planning function; express normalized time The derivative of the attitude planning function in the normalized time domain is the rate of change of the attitude planning function. This represents the derivative of the position planning function with respect to actual time t in the actual time domain, i.e., the rate of change of the position of the space robot in the actual time. express normalized time The derivative; This represents the derivative of the attitude planning function with respect to actual time t in the actual time domain, i.e., the rate of change of the attitude of the space robot in actual time.
[0037] Considering the initial and final conditions and trajectory smoothness of the space robot, the following conditions must be met: ; The velocity of the space robot is planned using the trapezoidal programming method as follows: .
[0038] make , This represents the position vector of the space robot. This represents the attitude vector of the space robot. Based on the planned end-effector velocity and angular velocity, the joint angular velocities of the space robot can be planned using only the velocity-level inverse kinematics equations. : ; In the above formula, This represents the rate of change of location after planning. This represents the relevant quantity of angular velocity after planning; This represents the terminal velocity correlation vector.
[0039] Step 220: Use the damped minimum variance method to solve for the joint angular velocities of the planned space robot.
[0040] Generalized Jacobian matrix When a singularity occurs, the joint angular velocity of the space robot Infinity causes the corresponding trajectory planning to fail. The damped minimum variance method is one of the commonly used methods for solving inverse motion under singular conditions, i.e., planning the joint angular velocities of a space robot as follows: ; In the above formula, for The minimum variance inverse: ; in, E Represents the identity matrix; The damping coefficient is adjusted adaptively using the following formula: ; In the above formula, for The estimate of the minimum singular value; It is a judgment Is it a singular threshold? This is the maximum damping value within the neighborhood of the singular point. Clearly, the minimum variance damping method is used to handle dynamic singularities, utilizing... replace The reversal caused the space robot to deviate from its planned trajectory.
[0041] Step 300: Randomize the initial pose and desired pose of the space robot.
[0042] As an example, this step includes: Step 310: Initial pose randomization setting of the space robot.
[0043] In the space robot simulation environment, the arbitrary arm shape and spacecraft body attitude within the workspace of the space robot are taken as the initial nominal state, and the joint angles are set to a randomized range centered on the initial nominal value. If the initial arm shape after randomization is in a non-singular state, it is used as the initial pose of the space robot; otherwise, it is re-randomized until a non-singular initial pose is obtained.
[0044] In this embodiment, the specific design is as follows: Initial nominal state setting: Spacecraft body attitude Joint angle The randomization range corresponding to each joint angle .
[0045] Step 320: Randomization setting of the desired pose of the space robot.
[0046] Using the initial joint angles of the space robot and the spacecraft's body attitude corresponding to the space robot's pose as the center, the randomization range of the space robot's desired pose is set, as follows: The desired randomization range of positions is: the distance from the initial position in the x, y, and z directions all satisfy the following conditions. Within the range, This represents the minimum distance from the initial position in the x, y, and z directions. This represents the maximum distance from the initial position in the x, y, and z directions; Desired pose randomization range: Euler angles from the initial pose to the desired pose. , , Sizes all meet Within the range, This represents the minimum absolute value of each Euler angle. This represents the maximum absolute value of each Euler angle.
[0047] In this embodiment, it is specifically set as follows: .
[0048] Step 400: Based on the deep reinforcement learning algorithm, design a reward function considering the spacecraft's attitude perturbation factors, construct and train the stage target pose decision model until the reward function converges, and obtain the trained stage target pose decision model.
[0049] As an example, step 400 specifically includes: Step 410: Design the state space, action space, and reward function for training the target pose decision model in the design phase.
[0050] The state space includes the current pose, desired pose, singularity analysis and other related parameters of the space robot. The action space is the target pose of the space robot in the future stage. The reward function includes spacecraft attitude disturbance reward, dynamic singularity reward, space robot process accuracy reward and space robot final accuracy reward.
[0051] In this embodiment, the specific design is as follows: (1) State space: The state space is 26-dimensional, including 6-dimensional joint angles, 4-dimensional body pose, 6-dimensional initial pose at a certain stage, 6-dimensional expected pose, and 4-dimensional Jacobian matrix. Singularity-related information; Determine the 4-dimensional Jacobian matrix Is it a singular threshold? Using a trapezoidal programming method based on the damped minimum variance approach, we can plan the minimum value corresponding to the trajectory from the current state to the final desired state. , The corresponding relative time and the current time correspond Minimum singular value estimate .
[0052] (2) Action space: The action output is a 4-dimensional discrete action, which represents the stage target position information and stage target attitude information in the x, y and z directions, respectively. Each dimension takes an integer between 0 and 40.
[0053] Taking the x-axis position as an example, the relationship between the stage target position and the action output is as follows: ; in, The action output corresponds to the x-axis; This indicates adjusting the gain parameter in the location planning; This represents the desired target position of the space robot along the x-axis. This indicates the current position of the space robot along the x-axis. This indicates the target stage position of the space robot in the x-axis direction.
[0054] The relationship between the target posture and the action output at each stage is as follows: ; In the above formula, This indicates adjusting the gain parameter in attitude planning; This represents the target posture information of the space robot at each stage. This indicates the angle of rotation required for the attitude to change from the current state to the desired attitude. The rotation axis representing the rotation from the current attitude state to the desired attitude is a unit vector; Output the action corresponding to the posture. From the current posture... Rotation Afterwards, the posture changed from Adjust to From the current posture Rotation Afterwards, the posture changed from Adjust to .
[0055] (3) Reward Design: The reward function includes spacecraft attitude disturbance reward, dynamic singularity reward, space robot process accuracy reward, and space robot final accuracy reward. (a) Spacecraft attitude perturbation reward The spacecraft rotates around its rotation axis from the initial moment. Design spacecraft body perturbation reward : ; in, The threshold is set to 10° in this embodiment. For the desired pose randomization range, it was found in reinforcement learning training that 10° can achieve better training results.
[0056] (b) Dynamically exotic reward When transitioning from the current state to the desired pose using the traditional trapezoidal programming method, a reward function is designed based on whether a dynamic singularity is encountered. : ; in, During the movement of the aforementioned space robot corresponding Minimum value.
[0057] Furthermore, if the space robot is in a dynamically singular situation at the termination time, it will not only affect the current trajectory planning task but also make it difficult to avoid dynamically singular situations in future planning tasks. Therefore, to avoid dynamically singularities, the following reward function is added. : ; In the above formula, For the termination time .
[0058] (c) Space robot process accuracy reward: from the initial pose of the stage Transfer to target pose At that time, a reward function is designed based on the accuracy of the end-effector process of the space robot. : ; In the above formula, The position error can be expressed as , This represents the actual pose of the space robot's end effector. This refers to the attitude error, which is the error from the actual attitude. Rotation about the axis of rotation It can be adjusted to the target posture. ; Indicates the required positional accuracy for the trajectory planning task; This represents the attitude accuracy required for the trajectory planning task. In this embodiment, , .
[0059] (d) Final accuracy reward for space robots : From the initial pose of the stage to desired position At that time, a reward function is designed based on the final accuracy of the space robot's end effector: ; In the above formula, This represents the difference between the actual position and the expected position of the end effector of the space robot at the termination time. This represents the angle of rotation required to change the actual posture to the desired posture at the termination time. Indicates the required positional accuracy for the trajectory planning task; This represents the attitude accuracy required for the trajectory planning task. In this embodiment, , .
[0060] Step 420: Set a limit on the number of steps m required to complete trajectory planning in the deep reinforcement learning algorithm, i.e. In this embodiment, , .
[0061] Step 430: Build an action decision network and a value network, and use reinforcement learning to train the decision network and value network until the reward curve converges, thus obtaining a trained stage target pose decision model.
[0062] In this embodiment, the specific design is as follows: (1) Design of the action decision network, such as Figure 3 As shown, it consists of a common feature extraction network and four decision output networks with the same structure; the value network adopts a double hidden layer fully connected structure, taking the common features obtained after preprocessing the data by the common feature extraction network as input. The double hidden layer fully connected network has 128 and 32 nodes respectively, and the output layer has 1 node, representing the output value of the corresponding action value function. (2) The reinforcement learning method PPO is used to learn and train the stage target pose decision model in the constructed training framework until the reward curve converges, and the trained stage target pose decision model is obtained.
[0063] Step 500: Load the trained stage target pose decision model into the space robot system and combine it with the trapezoidal programming algorithm to generate the trajectory from the initial pose to the desired pose.
[0064] Using a trained stage target pose decision model, the corresponding stage target pose is obtained based on the current state of the space robot; using the trapezoidal programming method, the trajectory from the current pose to the stage target pose is formed.
[0065] Repeat the above steps until the robot reaches the desired pose, repeating the steps a certain number of times. m satisfy Finally, the trajectory from the initial pose to the desired pose is generated.
[0066] Taking a 6-DOF space robot as an example, a space robot simulation environment is built to learn and train a phased target pose decision model. For example... Figure 4 As shown, the convergence curve of the reward function is presented, indicating that it basically converges and tends to stabilize at around -20. The trained stage target pose decision model is then applied to the space robot system for algorithm verification. The following example is given: The initial joint angles of the space robot and the attitude of the spacecraft body are as follows: ; The corresponding space robot pose is: ; At the termination time, the desired pose of the space robot is: ; Using the method proposed in this invention and the traditional trapezoidal trajectory planning method based on damped minimum variance, the spacecraft's attitude was ultimately adjusted relative to the initial moment, rotating around the rotation axis. and In other words, the method of this invention reduces the attitude disturbance of the spacecraft. Furthermore, the position error of the space robot along the y-axis corresponding to the two methods is compared, for example... Figure 5 As shown, traditional trajectory planning methods sacrifice space robot error by introducing the damped minimum variance method to handle dynamic singularities. However, the method proposed in this invention can still plan the space robot to the desired state even when encountering dynamic singularities.
[0067] Simulation analysis shows that the trajectory planning method based on deep reinforcement learning proposed in this invention can not only optimize the attitude disturbance of the spacecraft body, but also handle dynamic singularity problems while ensuring the accuracy of the space robot's end effector.
[0068] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A spatial robot trajectory planning method based on deep reinforcement learning, characterized in that, Includes the following steps: Build a space robot simulation environment; The initial pose and desired pose of the space robot are set randomly; Based on deep reinforcement learning algorithms, a reward function is designed considering the attitude disturbance factors of the spacecraft itself. A stage target pose decision model is constructed and trained until the reward function converges, and a well-trained stage target pose decision model is obtained. The trained stage target pose decision model is loaded into the space robot system. Using the trained stage target pose decision model, the corresponding stage target pose is obtained based on the current state of the space robot. Using the trapezoidal programming method, a trajectory is formed from the current pose to the target pose of the stage; Repeat the above steps until the space robot reaches the desired pose, and finally generate the trajectory from the initial pose to the desired pose.
2. The spatial robot trajectory planning method based on deep reinforcement learning according to claim 1, characterized in that, The construction of the space robot simulation environment includes: building a kinematic model of a free-floating space robot, wherein the mathematical model of the free-floating space robot includes a kinematic model and attitude relationships.
3. The spatial robot trajectory planning method based on deep reinforcement learning according to claim 2, characterized in that, The construction of the space robot simulation environment also includes: building a trapezoidal programming method module based on the damped minimum variance method, which is used to generate the trajectory of the space robot from its current pose to the target pose of the stage.
4. The spatial robot trajectory planning method based on deep reinforcement learning according to claim 1, characterized in that, The initial pose of the space robot is randomized, including: The arbitrary arm shape and spacecraft body posture within the workspace of the space robot are taken as the initial nominal state, and the randomized range of each joint angle is set with the initial nominal value as the center. If the initial arm shape after randomization is in a non-singular state, it is used as the initial pose of the space robot; otherwise, it is re-randomized until a non-singular initial pose is obtained.
5. The spatial robot trajectory planning method based on deep reinforcement learning according to claim 1, characterized in that, Based on deep reinforcement learning algorithms, a phase target pose decision model is constructed and trained, including: The design phase training of the target pose decision model includes the state space, action space, and reward function; the state space includes the current state of the space robot, the expected pose of the end effector, and singular analysis; the action space is the target pose information of the space robot in future stages.
6. The spatial robot trajectory planning method based on deep reinforcement learning according to claim 5, characterized in that, The reward function includes spacecraft attitude disturbance reward, dynamic singularity reward, space robot process accuracy reward, and space robot final accuracy reward.