A Method for Generating Intelligent Grasping Strategies for Tools in a Manipulator-Dexterous Hand System for the Space Microgravity Environment
By building a simulation model in a spatial microgravity environment and using neural network training, an intelligent grasping strategy of the robotic arm-dexterous hand system was generated, and the problems of crawling failure and difficulty in converging in the existing technology were solved, and stable grasping in the spatial environment was achieved.
Patent Information
- Application Number
- CN202510071761.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-01-16
AI Technical Summary
In the spatial microgravity environment, it is difficult for the prior art to effectively train intelligent models to realize tool grabbing of robotic arm-dexterous hand systems, resulting in operation failure and the model is difficult to converge quickly.
By constructing a simulation environment under spatial microgravity, neural network training is used to generate intelligent grasping strategies of the robotic arm-dexterity hand system, input is the tool's position information and the real-time state of the robotic arm and dexterity hand, output is its actions, and trained using LSTM-PPO strategy and reward function.
The intelligent grasping strategy trained in the simulation environment can effectively deal with the spatial microgravity environment, improving the success rate of the grab and the convergence speed of the model.
Smart Images

Figure CN119795171B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent robots, and particularly relates to a method for generating an intelligent grasping strategy for a manipulator-dexterous hand system tool for a space microgravity environment. Background Art
[0002] At present, with the gradual development of space technology and the gradual maturity of robot technology, the use of a manipulator-dexterous hand system to work in cooperation with astronauts in the daily operation and maintenance of large facilities such as space stations has broad application prospects. Due to the dynamic and complex characteristics of the scenarios and targets in space missions, intelligent methods need to be used to achieve flexible control and autonomously achieve collision avoidance planning and stable operation. For the task of the robot to grasp and transfer tools, due to the types of targets and the microgravity characteristics of the space environment, its physical behavior is quite different from that on the ground. It is difficult to directly apply the intelligent models trained in ground scenarios to space manipulation tasks. Similar operation actions on the ground are likely to cause floating objects to be knocked away, resulting in task failure. At the same time, the above problems also become the training difficulties of intelligent models. Due to the small fault tolerance of operations and the large motion space of the dexterous hand, it is difficult for the model to converge quickly. Therefore, it is of great significance to develop a method for generating an intelligent grasping strategy for a manipulator-dexterous hand system tool for a space microgravity environment and form an intelligent dexterous operation ability in a space environment. Summary of the Invention
[0003] An embodiment of the present invention provides a method for generating an intelligent grasping strategy for a manipulator-dexterous hand system tool for a space microgravity environment, which can generate an intelligent grasping strategy for a manipulator-dexterous hand system tool for a space microgravity environment.
[0004] An embodiment of the present invention provides a method for generating an intelligent grasping strategy for a manipulator-dexterous hand system tool for a space microgravity environment, including:
[0005] Setting a gravity coefficient based on the installation configuration of the manipulator-dexterous hand system, the tool to be grasped, and the environmental scenario, and constructing a simulation environment under space microgravity;
[0006] Generating a grasping configuration of the dexterous hand based on the tool to be grasped and the configuration of the dexterous hand, so as to perform a fixed grasping test in the simulation environment;
[0007] Based on the simulation environment, taking the pose information of the tool to be grasped, and the real-time states of the manipulator and the dexterous hand as the input of neural network training, and taking the actions of the manipulator and the dexterous hand as the output of neural network training, and obtaining a grasping strategy through neural network training.
[0008] Optionally, the reward function of the neural network training includes a proximity reward function and a grasping reward function.
[0009] Optionally, the proximity reward function is as follows:
[0010]
[0011] where α, β, γ, ε, ρ, μ are parameters greater than 0, p h is the vector of the tool relative to the palm center, v is the velocity of the tool relative to the palm, a x , a y , a z correspond to the unit vectors of the three axes in the three-dimensional model of the palm, a ox corresponds to the x-axis of the tool, a hx the x-axis of the thumb tip, p ho is the orientation unit vector of the tool relative to the palm center.
[0012] Optionally, the grasping reward function is as follows:
[0013]
[0014] where r a is the proximity reward function, μ is a parameter greater than 0, p H is the deviation vector between the current joint angles of the dexterous hand and the grasping configuration angles.
[0015] Optionally, the neural network training includes:
[0016] Proximity stage: Guide the robotic arm-dexterous hand system to approach the tool to be grasped. When the relative position and orientation between the tool to be grasped and the palm of the dexterous hand are within the set first threshold for a preset time, it is determined that the proximity task is completed, and the robotic arm-dexterous hand system is guided to grasp the tool to be grasped; when the relative position and orientation between the tool to be grasped and the palm exceed the set second threshold or the force at the end of the palm is greater than the preset force, it is determined that the task fails, and the penalty value of the proximity reward function is increased;
[0017] Grasping stage: When guiding the robotic arm-dexterous hand system to grasp the tool to be grasped, when the relative position and orientation between the tool to be grasped and the palm are within the third threshold and the relative distance between the current dexterous hand and the target configuration is within the set fourth threshold for a preset time, it is determined that the grasping task is completed, and the reward value of the grasping reward function is increased; when the relative position and orientation between the tool to be grasped and the palm exceed the set second threshold or the force at the end of the palm is greater than the preset force, it is determined that the task fails, and the penalty value of the grasping reward function is increased.
[0018] Optionally, the reset policy for the neural network training is as follows:
[0019] When the judgments of the proximity stage and the grasping stage are completed, reset the simulation environment;
[0020] In the reset simulation environment, control a 1 / 10 probability to initialize the tool to be grasped in the palm, with its position and orientation both close to the correct grasping state;
[0021] In the reset environment, control a 1 / 10 probability to initialize the tool to be grasped near the palm, with its orientation close to the correct grasping orientation;
[0022] In the reset environment, control a 4 / 5 probability to make the position and orientation of the tool to be grasped both random.
[0023] Optionally, the actions of the dexterous hand and the robotic arm include the six-axis actions of the robotic arm and the joint actions of the seven degrees of freedom of the dexterous hand.
[0024] Optionally, the neural network training adopts the LSTM-PPO strategy, and the network architecture includes:
[0025] One layer of LSTM, with 512 neurons;
[0026] The first layer of MLP, with 256 neurons;
[0027] The second layer of MLP, with 128 neurons;
[0028] The third layer of MLP, with 64 neurons;
[0029] The output layer result is a 13-dimensional action.
[0030] Optionally, it further includes:
[0031] Perform inference verification on the trained neural network.
[0032] The present invention has at least the following beneficial effects compared with the prior art:
[0033] In this embodiment, by simulating the space microgravity environment through simulation, a simulation environment is obtained. Under the simulation environment, neural network training is carried out to obtain the intelligent grasping strategy of the robotic arm-dexterous hand system. Among them, the input of the neural network training is the pose information of the tool to be grasped, and the real-time states of the robotic arm and the dexterous hand, and the output of the neural network training is the actions of the robotic arm and the dexterous hand. Description of the Drawings
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0035] Figure 1Flowchart of the intelligent grasping strategy generation method for the manipulator - dexterous hand system facing the space microgravity environment of the present invention;
[0036] Figure 2 Image of the manipulator and dexterous hand system simulation environment in this embodiment of the present invention;
[0037] Figure 3 Image of the reasonable configuration of the human hand grasping the tool obtained in this embodiment of the present invention;
[0038] Figure 4 Image of the intelligent training network model adopted in this embodiment of the present invention;
[0039] Figure 5 Process image of the dexterous hand approaching the tool during the model inference test in this embodiment of the present invention;
[0040] Figure 6 Image of the dexterous hand completing the approach to the tool during the model inference test in this embodiment of the present invention;
[0041] Figure 7 Process image of the dexterous hand completing the grasping of the tool during the model inference test in this embodiment of the present invention. Detailed implementation manners
[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0043] In the description of the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance; unless otherwise specified or stated, the term "plural" means two or more; the terms "connection", "fixation", etc. shall be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, an integral connection, or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0044] In the description of this specification, it should be understood that the orientation terms such as "upper" and "lower" described in the embodiments of the present invention are described from the angles shown in the drawings and should not be construed as limitations on the embodiments of the present invention. In addition, in the context, it should also be understood that when it is mentioned that an element is connected "above" or "below" another element, it can not only be directly connected "above" or "below" another element, but also be indirectly connected "above" or "below" another element through an intermediate element.
[0045] As Figure 1 shown, the embodiments of the present invention provide a method for generating an intelligent grasping strategy for a manipulator-dexterous hand system tool for a space microgravity environment, including:
[0046] S1. Based on the installation configuration of the manipulator-dexterous hand system, the tool to be grasped, and the environmental scene, set the gravity coefficient to construct a simulation environment under space microgravity;
[0047] S2. Based on the tool to be grasped and the configuration of the dexterous hand, generate the grasping configuration of the dexterous hand to perform a fixed grasping test in the simulation environment;
[0048] S3. Based on the simulation environment, use the pose information of the tool to be grasped, and the real-time states of the manipulator and the dexterous hand as the input of neural network training, and use the actions of the manipulator and the dexterous hand as the output of neural network training to obtain the grasping strategy through neural network training.
[0049] The manipulator has more than 6 active degrees of freedom and can realize the translation of 3 degrees of freedom and the rotation of 3 degrees of freedom at the end of the manipulator; the dexterous hand has 5 fingers, and its joint distribution is similar to that of a human hand. Each finger has 4 degrees of freedom, and the number of active degrees of freedom is not less than 2; the manipulation target is a common tool such as a hammer, a tool with a certain initial velocity and linear velocity, and its pose and velocity can be obtained by calculation by a camera or directly obtained in the simulation environment; the space floating environment can be set from 0 to 1g.
[0050] In this embodiment, a simulation environment is obtained by simulating the space microgravity environment. In the simulation environment, neural network training is performed to obtain the intelligent grasping strategy of the manipulator-dexterous hand system. Among them, the input of neural network training is the pose information of the tool to be grasped, and the real-time states of the manipulator and the dexterous hand, and the output of neural network training is the actions of the manipulator and the dexterous hand.
[0051] For S1, the simulation environment can construct a simulation model according to the general description file of the manipulator and the dexterous hand mechanism, can parameterize the setting of gravity parameters, and supports parallel training of multiple simulation environments.
[0052] The specific steps for setting the simulation environment are as follows. The simulation environment is as Figure 2 shown:
[0053] (1) The simulation environment uses IsaacGym, the simulation physics engine uses PhysX, and the interface program for the reinforcement learning training of the intelligent grasping network uses IsaacGymEnvs;
[0054] (2) The UR5 robotic arm is adopted, which has a total of 6 active degrees of freedom, and its kinematic and dynamic models are configured based on the URDF general description file;
[0055] (3) The five-fingered dexterous hand is adopted, and each finger has 4 degrees of freedom, including active degrees of freedom and passive coupled degrees of freedom. Its simplified kinematic and dynamic models are configured based on the URDF general description file. For the passive coupled degrees of freedom, an online solution based on the given value of the active degrees of freedom and an active assignment in the simulation environment are used to avoid the situation where a complex hybrid mechanism cannot be solved in the simulation engine;
[0056] (4) Set the base coordinate system of the robotic arm to be fixedly connected to the world coordinate system;
[0057] (5) Set the tool to be grasped as a hammer. Its three-dimensional model uses RGB camera and real object image data, and is constructed and manually optimized and adjusted in cooperation with the NeRF three-dimensional reconstruction network. Its three-dimensional model is configured based on the URDF general description file, and its pose is initialized in the simulation environment. Set it not to be fixedly connected to any object to achieve the simulation of the floating state;
[0058] (6) Set simulation parameters such as the simulation step size and the pid parameters of position control.
[0059] For S2, the grasping configuration of the dexterous hand can obtain a reasonable configuration of human hand grasping according to the target, and generate and adjust the grasping configuration of the dexterous hand based on this.
[0060] Based on the configuration of the object to be manipulated and the dexterous hand, generate the grasping configuration of the dexterous hand, including the relative pose between the root of the dexterous hand and the object to be manipulated and the angles of each joint of the dexterous hand;
[0061] The method for generating the grasping configuration is as follows:
[0062] (1) Based on OakInk, obtain the typical configurations of human hand grasping, including relative poses and the values of each joint of the hand, as Figure 3 shown;
[0063] (2) In the simulation environment, perform fixed grasping tests according to the configurations obtained in step (1), and manually optimize and adjust the configurations to ensure adaptation to the five-fingered dexterous hand.
[0064] For S3, specifically, the observation input may include some information among the shape information of the target (the tool to be grasped), the position, orientation, linear velocity, and angular velocity of the target relative to the base coordinate system of the robotic arm, the angles, angular velocities of the joints of the robotic arm, the position, orientation, linear velocity, and angular velocity of the end TCP point, the angles, angular velocities of the joints of the dexterous hand, the positions of the end points of each finger of the dexterous hand, the six-axis force sensor at the wrist of the robotic arm, the force sensors of the fingers of the dexterous hand, and other status information.
[0065] The reward function calculates the reward function according to the difference between the input status information and the ideal close state, sets penalties for factors such as collisions causing the tool to move away, and determines the scope of action of the current reward function.
[0066] The determination of the observation input of the intelligent network is specifically as follows and can be adjusted according to the actual training effect:
[0067] (1) The first to the third dimensions are the position of the tool, and the fourth to the seventh dimensions are the quaternion of the orientation of the tool, and the values are all relative to the base coordinate system of the robotic arm;
[0068] (2) The eighth to the tenth dimensions are the linear velocity of the tool, and the eleventh to the thirteenth dimensions are the angular velocity of the tool, and the values are all relative to the base coordinate system of the robotic arm;
[0069] (3) The fourteenth to the nineteenth dimensions are the angles of the joints of the robotic arm, and the nineteenth to the twenty-fourth dimensions are the angular velocities of the joints of the robotic arm;
[0070] (4) The twenty-fifth to the forty-fourth dimensions are the angles of the joints of the dexterous hand, and the forty-fifth to the sixty-fourth dimensions are the angular velocities of the joints of the dexterous hand;
[0071] (5) The sixty-fifth to the sixty-seventh dimensions are the position of the root of the dexterous hand, and the sixty-eighth to the seventy-first dimensions are the quaternion of the orientation of the root of the dexterous hand, and the values are all relative to the base coordinate system of the robotic arm;
[0072] (6) The seventy-first to seventy-fifth dimensions are the contact forces at the ends of each finger of the dexterous hand;
[0073] (7) The seventy-sixth to seventy-eighth dimensions are the dimensions of the tool;
[0074] (8) The seventy-ninth to ninety-first dimensions are the output action values of the previous network;
[0075] (9) The ninety-second to one hundred and eleventh dimensions are the angle values of each joint of the dexterous hand configuration calculated by S2.
[0076] In some embodiments of the present invention, the reward function for neural network training includes a proximity reward function and a grasping reward function.
[0077] In some embodiments of the present invention, the proximity reward function is as follows:
[0078]
[0079] Among them, α, β, γ, ε, ρ, μ are parameters greater than 0, and p h is the vector of the tool relative to the palm center, v is the speed of the tool relative to the palm, and a x , a y , a x correspond to the unit vectors of the three axes in the three-dimensional model of the palm. a ox corresponds to the x-axis of the tool, and a hx is the x-axis of the thumb tip, and p ho is the orientation unit vector of the tool relative to the palm center.
[0080] Specifically, The term is the change in the relative distance between the palm and the target, encouraging approaching the target operating component.
[0081] The term v·(-a x ) is the speed reward for approaching in the normal direction of the palm center;
[0082] The term v·(-a z ) is the speed reward for approaching the normal direction of the end;
[0083] The term |a ox ·a y | is the rotational attitude reward, making the tool hammer handle direction correspond to the palm direction;
[0084] The term p ho ·a x is the orientation reward relative to the palm a x ;
[0085] The term p ho ·a hx is the orientation reward relative to the palm fingertip a hx .
[0086] In some embodiments of the present invention, the grasping reward function is as follows:
[0087]
[0088] Among them, r a is the proximity reward function, μ is a parameter greater than 0, and p H is the deviation vector between the current joint angles of the dexterous hand and the grasping configuration angles.
[0089] The term is the change in the relative distance between the current dexterous hand and the target configuration, encouraging the dexterous hand to achieve the desired configuration grasp.
[0090] In some embodiments of the present invention, the neural network training includes:
[0091] Approaching stage: Guide the robotic arm - dexterous hand system to approach the tool to be grasped. When the relative position and attitude between the tool to be grasped and the palm of the dexterous hand are within the set first threshold for a preset time, it is determined that the approaching task is completed, and the robotic arm - dexterous hand system is guided to grasp the tool to be grasped; when the relative position and attitude between the tool to be grasped and the palm exceed the set second threshold or the force at the end of the palm is greater than the preset force, it is determined that the task fails, and the penalty value of the approaching reward function is increased.
[0092] Grasping stage: When the robotic arm - dexterous hand system grasps the tool to be grasped, when the relative position and attitude between the tool to be grasped and the palm are within the third threshold, and the relative distance between the current dexterous hand and the target configuration is within the set fourth threshold for a preset time, it is determined that the grasping task is completed, and the reward value of the grasping reward function is increased; when the relative position and attitude between the tool to be grasped and the palm exceed the set second threshold or the force at the end of the palm is greater than the preset force, it is determined that the task fails, and the penalty value of the grasping reward function is increased.
[0093] In some embodiments of the present invention, the reset strategy for neural network training is as follows:
[0094] When the judgments of the approaching stage and the grasping stage are completed, reset the simulation environment;
[0095] In the reset simulation environment, control a probability of 1 / 10 to initialize the tool to be grasped in the palm of the hand, and its position and attitude are both close to the correct grasping state;
[0096] In the reset environment, control a probability of 1 / 10 to initialize the tool to be grasped near the palm of the hand, and its attitude is close to the correct grasping attitude;
[0097] In the reset environment, control a probability of 4 / 5 to make the position and attitude of the tool to be grasped both random.
[0098] In some embodiments of the present invention, the actions of the dexterous hand and the robotic arm include the six - axis actions of the robotic arm and the joint actions of the seven degrees of freedom of the dexterous hand.
[0099] Specifically, the output action includes a total of 6 degrees of freedom, corresponding to the six - axis actions of the robotic arm, and they are all relative increment values, that is, the expected joint position of the robotic arm is set as the sum of the current position and the output action value of the intelligent grasping network;
[0100] For the joint actions of the dexterous hand, to reduce the exploration space, each flexion - extension joint in each finger is set to be virtually coupled. The output action includes a total of 7 degrees of freedom, including the closing and extension of all flexion - extension joints, the lateral swing of the five fingers, and the rotational degree of freedom of the thumb along the root, and they are all relative increment values, that is, the expected joint position of the dexterous hand is set as the sum of the current position and the output action value of the intelligent grasping network.
[0101] In some embodiments of the present invention, neural network training adopts the LSTM-PPO strategy, and the network architecture ( Figure 4 ) includes:
[0102] One layer of LSTM with 512 neurons;
[0103] The first layer of MLP with 256 neurons;
[0104] The second layer of MLP with 128 neurons;
[0105] The third layer of MLP with 64 neurons;
[0106] The output layer result is a 13-dimensional action.
[0107] In some embodiments of the present invention, it further includes:
[0108] Performing inference verification on the trained neural network.
[0109] In this embodiment, the initial pose of the randomized tool is used, and the trained intelligent grasping model is used for inference verification, which can complete obstacle avoidance grasping of the tool, and tools with shapes similar to those used in the training process can be used to test the generalization ability of the intelligent method. The specific process is as Figure 5 、 Figure 6 、 Figure 7 shown.
[0110] The present invention can be widely applied to microgravity environments such as space stations, where space robots can operate autonomously or in cooperation with astronauts, such as tasks like the repair and replacement of key components. Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating an intelligent grasping strategy for a manipulator-dexterous hand system tool in a space microgravity environment, characterized in that Including: Based on the installation configuration of the robotic arm - dexterous hand system, the tool to be grasped, and the environmental scene, set the gravity coefficient to construct a simulation environment under space microgravity; Based on the configuration of the tool to be grasped and the dexterous hand, generate the grasping configuration of the dexterous hand to perform a fixed grasping test in the simulation environment; Based on the simulation environment, use the pose information of the tool to be grasped, and the real-time states of the robotic arm and the dexterous hand as the input for neural network training, and use the actions of the robotic arm and the dexterous hand as the output of neural network training to obtain a grasping strategy through neural network training; among them, the reward function of the neural network training includes an approaching reward function and a grasping reward function; The neural network training includes: Approaching stage: Guide the robotic arm - dexterous hand system to approach the tool to be grasped. When the relative position and attitude between the tool to be grasped and the palm of the dexterous hand are within the set first threshold for a preset time, it is judged that the approaching task is completed, and the robotic arm - dexterous hand system is guided to grasp the tool to be grasped; when the relative position and attitude between the tool to be grasped and the palm exceed the set second threshold or the force at the end of the palm is greater than the preset force, it is judged that the task fails, and the penalty value of the approaching reward function is increased; Grasping stage: When guiding the robotic arm - dexterous hand system to grasp the tool to be grasped, when the relative position and attitude between the tool to be grasped and the palm are within the third threshold, and the relative distance between the current dexterous hand and the target configuration is within the set fourth threshold for a preset time, it is judged that the grasping task is completed, and the reward value of the grasping reward function is increased; when the relative position and attitude between the tool to be grasped and the palm exceed the set second threshold or the force at the end of the palm is greater than the preset force, it is judged that the task fails, and the penalty value of the grasping reward function is increased; The reset strategy of the neural network training is as follows: When the judgments of the approaching stage and the grasping stage are completed, reset the simulation environment; In the reset simulation environment, control a probability of 1 / 10 to initialize the tool to be grasped in the palm of the hand, and its position and attitude are both close to the correct grasping state; In the reset environment, control a probability of 1 / 10 to initialize the tool to be grasped near the palm of the hand, and its attitude is close to the correct grasping attitude; In the reset environment, control a probability of 4 / 5 to make the position and attitude of the tool to be grasped both random.
2. The method according to claim 1, wherein The approaching reward function is as follows: where α, β, γ, ε, ρ, μ are parameters greater than 0, p h is the vector of the tool relative to the palm center, v is the velocity of the tool relative to the palm, a x , a y , a z correspond to the three-axis unit vectors in the three-dimensional model of the palm, a ox corresponds to the x-axis of the tool, a hx x-axis of the thumb tip, p ho is the unit vector of the tool's orientation relative to the palm center, The term is the change in the relative distance between the palm and the target.
3. The method according to claim 1, wherein The grasping reward function is as follows: where r a is the proximity reward function, μ is a parameter greater than 0, and p H is the deviation vector between the current joint angles of the dexterous hand and the grasping configuration angles, and the 4. The method according to claim 1, characterized in that, The actions of the dexterous hand and the robotic arm include the six-axis actions of the robotic arm and the joint actions of the seven degrees of freedom of the dexterous hand.
5. The method according to claim 1, wherein The neural network training adopts the LSTM-PPO strategy, and the network architecture includes: One layer of LSTM, 512 neurons; The first layer of MLP, 256 neurons; The second layer of MLP, 128 neurons; The third layer of MLP, 64 neurons; The output layer result is a 13-dimensional action.
6. The method according to claim 1, characterized in that It also includes: Perform inference verification on the trained neural network.
Citation Information
Patent Citations
Space manipulator reinforcement learning motion planning method for unfixed obstacles
CN116619380A