A parallel fixture shape generation method based on DDPG reinforcement learning algorithm

Through the parallel fixture shape generation method based on DDPG reinforcement learning algorithm, the problem of fixture design relies on manual experience in the prior art is solved, automated design and lightweight shape generation are realized, and the grab success rate and stability are improved, and manufacturing costs are reduced.

CN115890726BActive Publication Date: 2025-05-20DALIAN UNIV OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211404026.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-10
Publication Date
2025-05-20
Estimated Expiration
2042-11-10

AI Technical Summary

Technical Problem

The existing mechanical fixture design methods rely on manual experience, making it difficult to achieve efficient and stable grasping in a variety of scenes of objects to be grasped, and the resulting fixtures are bulky in shape, which increases the manufacturing cost of industrial applications.

Method used

The parallel fixture shape generation method based on DDPG reinforcement learning algorithm is adopted to explore the fixture design space in the simulation environment through reinforcement learning, reduce the dependence of manual experience, realize automated design, and generate lightweight fixture shapes.

Benefits of technology

降低了人工成本,提高了夹具在多种场景下的抓取成功率和稳定性,避免了笨重形状的生成,减少了制造成本。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115890726B_ABST
    Figure CN115890726B_ABST
Patent Text Reader

Abstract

The present invention proposes a parallel fixture design shape generation method based on the DDPG reinforcement learning algorithm, which belongs to the field of parallel fixture shape generation for robot grasping tasks. The present invention uses a reinforcement learning method to automatically generate lightweight parallel fixture shapes, specifically including the design of parallel fixture shape space, state representation and action space design, the setting of grasping reward mechanism, the construction of fixture shape generation network, the update and training of generation network, etc., to achieve action selection based on a given object and a trained strategy network, thereby automatically generating a practical fixture shape. The present invention uses a parallel fixture shape automatic generation method using a reinforcement learning algorithm, which can achieve the goals of automated, generalizable, and lightweight parallel fixture shape generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of parallel fixture shape generation for robot grasping tasks, and relates to a method for generating parallel fixture shapes based on the DDPG reinforcement learning algorithm. Background Art

[0002] The automatic grasping of mechanical fixtures has been deeply applied in many fields, such as warehousing and sorting in the logistics field, automated production lines in the manufacturing field, equipment self-inspection and automatic maintenance in the logistics field, and even automatic loading in the military field, etc.

[0003] The design of mechanical fixtures has a long history. Traditional methods mainly rely on manual design, and the shape and parameters of the fixtures are adjusted and updated based on geometric analysis and long-term production practice experience. With the application of grasping methods to more scenarios, the geometric shapes of many different objects to be grasped pose challenges to the versatility of the fixtures. The emergence of machine learning methods makes it possible to automatically generate practically usable fixtures based on object features. Common methods include geometric analysis methods, heuristic-based imprinting methods, data-driven deep learning methods, etc.

[0004] However, most of the above methods have the following problems: (1) The real grasping success rate is not high or the grasping process is unstable. Specifically, the fixture shapes designed in specific situations rely on manually given professional experience or data experience, and it is sometimes difficult to obtain high performance in real grasping situations; (2) The method versatility is poor. Specifically, different objects to be grasped on the production line often require different fixture shapes. However, for some unseen objects lacking professional experience or data experience, the above methods all need to go through complex processing procedures to obtain the final result, because the generalization ability is weak; (3) The method is complex or the practical cost is high. Specifically, the fixture shape design based on geometric methods relies too much on manual experience, and the fixture design based on deep learning also requires a large amount of training data, and the method implementation process is relatively complex. In addition, the fixture shapes generated by the above methods are generally relatively bulky, increasing their manufacturing costs in the industrial application stage. Therefore, the present invention intends to use the method of reinforcement learning to automatically generate lightweight parallel fixture shapes. Summary of the Invention

[0005] Aiming at the problems existing in the prior art, the present invention proposes a method for generating the shape of a parallel fixture based on the DDPG reinforcement learning algorithm. The Deep Deterministic Policy Gradient (DDPG) algorithm not only utilizes the experience replay in the Deep Q-Network algorithm to solve the problem of correlation between parameters before and after, but also has the advantages of double neural networks and the policy gradient algorithm, and can handle continuous action control problems. Combining the goals of automation, generalization, and lightweight of the parallel fixture shape generation, a method for automatically generating the shape of a parallel fixture based on the DDPG reinforcement learning algorithm is proposed, which is of great significance for the design of lightweight fixtures in intelligent grasping.

[0006] The generation of the shape of a parallel fixture based on the DDPG reinforcement learning algorithm, this method is designed and improved based on the reinforcement learning method and the simulation grasping of the parallel fixture. The specific design features are as follows:

[0007] Different from the traditional fixture shape design method, the present invention uses the reinforcement learning method to explore the fixture design space, minimizing the manual experience and data experience required by the method and reducing the method complexity; simulating the clamping situations of different objects in the simulation environment to ensure the grasping accuracy and fixture generalization; combining with the fixture shape space design to avoid the generation of bulky fixture shapes and realizing the automated design of the parallel fixture.

[0008] The present invention realizes the above technical purpose through the following technical means:

[0009] Step S1: Design of the parallel fixture shape space;

[0010] Further, the specific content of step S1 is as follows:

[0011] Step S11: Based on the WSG50 fixture, construct the basic shape template of the fixture shape in the embodiment of the present invention;

[0012] Step S12: Design an L-shaped finger composed of a fixed plate and a shape-variable plate, and obtain the fixture shape in the embodiment of the present invention;

[0013] Step S13: Represent the shape of the L-shaped finger parametrically. The length of the clamping teeth is controlled by the length parameter, thereby controlling the shape of the shape-variable plate, and the installation position is controlled by the height parameter;

[0014] Step S14: Realize the mapping from the 11-dimensional vector to the fixture shape space, that is, according to the 11-dimensional parameter representation, obtain the corresponding left and right finger shapes of the fixture respectively.

[0015] Step S2: State representation and action space design;

[0016] Further, the specific content of step S2 is as follows:

[0017] Step S21: Abstractly represent the geometric shape of the object and the shape of the fixture fingers. Specifically, use the Truncated Signed Distance Function (TSDF) and a 3D convolutional shape encoder to obtain a 128-dimensional object shape encoding; use an 11-dimensional vector for controlling the left jaw and an 11-dimensional vector for the right jaw as the parametric representation of the fixture finger shape;

[0018] Step S22: Define the state space that includes the object encoding and the parametric representations of the left and right fixture finger shapes;

[0019] Step S23: Use the parameter change amounts of the left and right fixture fingers as actions, and then define the action space;

[0020] Step S24: Define the action constraints in the action space.

[0021] Step S3: Set up the grasping reward mechanism and use PyBullet physical simulation for grasping evaluation;

[0022] Furthermore, the specific steps of Step S3 are as follows:

[0023] Step S31: Based on Pybullet, set the parameters of the simulated grasping environment. Considering the grasping success, grasping stability, and grasping robustness, 9 different grasping scenarios are set;

[0024] Step S32: Set up the reward score calculation mechanism. After each action is executed, calculate the reward of the current environment, specifically as follows:

[0025]

[0026] Among them, T f represents the number of successful grasps in 4 scenarios of applying external forces, T 10 represents the number of successful grasps in the scenario of rotating 10 degrees clockwise or counterclockwise, T 20 represents the grasping success in the scenario of rotating 20 degrees clockwise or counterclockwise.

[0027] Step S4: Build the fixture shape generation network, construct the policy network μ actor and the value network Q, where the policy network μ actor is responsible for inputting the state s at the current moment and outputting the predicted action a, that is:

[0028] a = μ actor (s),

[0029] The value network Q is responsible for inputting the current state s and the executed action a, and outputting the predicted value, that is;

[0030] q = Q(s, a).

[0031] Further, the specific steps of step S4 are as follows:

[0032] Step S41: Using a fully connected layer and an activation layer, design a feature extraction module in the policy network, and process the object shape representation and the left and right fixture shape representations as a whole into a 128×3-dimensional vector, which is used as the shape feature extracted from the state representation;

[0033] Step S42: Using a fully connected layer and an activation layer, design an action selection module in the policy network;

[0034] Step S43: Using a fully connected layer and an activation layer, construct a value network, and predict the current value based on the connection vector of the object shape, the left and right fixture shapes, and the action features.

[0035] Step S5: Based on the DDPG reinforcement learning algorithm, update and train the fixture shape generation network;

[0036] Step S6: Automatically generate the fixture shape based on the given object, use the trained policy network to implement action selection, and achieve a shape change with a maximum number of steps of 20, so as to automatically generate the final fixture shape.

[0037] Beneficial effects of the present invention:

[0038] The present invention generates the shape of a parallel fixture based on the DDPG reinforcement algorithm. Compared with the design method based on geometric analysis, it reduces the dependence on artificial professional experience for the fixture shape; compared with the data-driven deep learning method, it reduces the dependence on training data and avoids the bulkiness of the fixture shape. The present invention adopts the DDPG reinforcement learning algorithm, and through the process of selecting actions given the state, it conducts self-training and exploration, greatly reducing the labor cost. By maximizing the grasping reward, it finally automatically generates a lightweight fixture shape. Description of the drawings

[0039] Figure 1 It is an example diagram of the fixture shape design space provided by the embodiment of the present invention.

[0040] Figure 2 It is the overall method flow chart provided by the embodiment of the present invention.

[0041] Figure 3 It is the training flow chart of the fixture shape generation network provided by the embodiment of the present invention.

[0042] Figure 4 It is the schematic diagram of the automatic fixture shape generation method provided by the embodiment of the present invention. Specific implementation method

[0043] Embodiments of the present invention will be described in detail below. The embodiments described with reference to the attached drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.

[0044] In recent years, the application of reinforcement learning in the field related to robot control has achieved good results. According to the shape of the object and the characteristics of the parallel fixture, the present invention proposes a method for generating the shape of a parallel fixture based on the DDPG reinforcement learning algorithm to adapt to various problems of the parallel fixture for gripping an object. The specific steps include:

[0045] Step S1: Design of the shape space of the parallel fixture;

[0046] Furthermore, the specific content of step S1 is as follows:

[0047] Step S11: Based on the WSG50 fixture, retain the fixture base of the WSG50 fixture except for the fingers as the basic shape template of the fixture shape of the present invention;

[0048] Step S12: Design an L-shaped finger composed of a fixed plate and a shape-variable plate. The shape-variable plate is controllably embedded in the fixed plate in terms of height. Install the L-shaped finger on the fixture base of the WSG50 to obtain the initial shape of the fixture in the embodiment of the present invention;

[0049] Step S13: Define the degrees of freedom of the L-shaped finger and use the length value and height value to achieve shape change, as Figure 1 shown. Specifically, discretize the length of the shape-variable plate into 40 voxels, and take adjacent 4 positions as a group to form 10 groups of length-variable clamping teeth. Each group of length values is used to control the length of each group of clamping teeth, thereby controlling the shape of the shape-variable plate. Discretize the length of the fixed plate into 40 voxels, and the height value is used to control the relative installation position between the shape-variable plate and the fixed plate;

[0050] Step S14: Realize the mapping from an 11-dimensional vector to the fixture shape space. Given an 11-dimensional integer vector, where the first 10-dimensional vectors represent the length values used to control the shape of the shape-variable plate, and the last 1-dimensional represents the height value for the installation position of the shape-variable plate. Given the 11-dimensional vector, the corresponding fixture shape can be obtained in the fixture shape design space.

[0051] Step S2: State representation and action space design;

[0052] Furthermore, the specific content of step S2 is as follows:

[0053] Step S21: Abstractly represent the geometric shape of the object and the shape of the fixture fingers. Specifically, use the Truncated Signed Distance Function (TSDF) to represent the object to be grasped. The object is preprocessed in a voxel space of 40×40×40. The geometric shape of the object is composed of the truncated signed distances of each voxel, and finally a 40×40×40 tensor is obtained. After being processed by a pre-trained 3D convolutional shape encoder, a 128-dimensional vector is obtained as the final encoded representation of the object's geometric shape. The fixture shape is abstractly represented by an 11-dimensional vector for controlling the left gripper teeth and an 11-dimensional vector for controlling the right gripper teeth. The control process of the gripper teeth in the corresponding fixture shape is as described in Step S14;

[0054] Step S22: Define the state space. The input state s = [s o , s l , s r , where s o is the 128-dimensional encoded representation of the object after preprocessing, is the 11-dimensional encoded representation of the left fixture finger shape, where is the shape parameter for controlling the variable plate of the left finger shape, is the height parameter for controlling the installation position of the variable plate of the shape, and similarly, is the 11-dimensional encoded representation of the right fixture finger shape;

[0055] Step S23: Define the actions in the action space. The action executed in each step of the reinforcement learning is the parameter change amount of the left and right fixture fingers. Define the action a = [Δs l , Δs r as a 22-dimensional vector, where represents the change amount of the left fixture shape and position, represents the change amount of the right fixture shape and position. Define the maximum number of action execution steps for each object as 20 steps;

[0056] Step S24: Define the action constraints in the action space. The parameter change amount corresponding to each action is fixed within the range of [-2, +2], where 2 represents 2 voxel units; after executing the action, among the left and right fingers of the fixture, the installation height of the variable plate of the shape is within the range of [0, 40], and the length of each group of gripper teeth is within the range of [0, 20].

[0057] Step S3: Set up the grasping reward mechanism and use PyBullet physical simulation for grasping evaluation;

[0058] Furthermore, the specific content of Step S3 is as follows:

[0059] Step S31: Set the parameters for simulating grasping based on Pybullet. To evaluate whether the gripper can grasp successfully, in the simulation environment, the grasping force of the parallel gripper is set to 1.0 N. It is defined that if the gripper can lift the object's centroid upward by 30 cm and maintain the lifted state for 2 s without dropping, it is recorded as a successful grasp. To evaluate the grasping stability of the gripper, for the successfully grasped instances, a 1-N force is applied in 4 predefined directions respectively for interference. If the object still does not drop after the interference, it is recorded as a stable grasp. To evaluate the grasping robustness of the gripper, 4 scenarios are set where the object rotates 10 degrees and 20 degrees clockwise and 10 degrees and 20 degrees counterclockwise around the z-axis starting from the initial position. If the gripper can still grasp successfully in the corresponding scenarios, it is recorded as a robust grasp. In summary, a total of 9 different grasping scenarios are set in the simulation environment;

[0060] Step S32: Set the reward score calculation mechanism. After each action is executed, calculate the reward of the current environment. For an object to be grasped, if it cannot be grasped successfully in the first normal grasping scenario, a direct reward of -5 is given for this step; conversely, if the grasp is successful, rewards are given based on whether the grasp is successful in different grasping scenarios. Specifically, if the grasp is successful in the normal scenario, a reward of 1000 is given; if the grasp is successful in the 4 scenarios with external force interference, a reward of 100 is given respectively; if the grasp is successful in the scenarios of rotating 10 degrees clockwise or counterclockwise, a reward of 200 is given. Similarly, if the grasp is successful in the scenario of rotating 20 degrees, a reward of 100 is given. In summary, the final reward calculation can be expressed as follows:

[0061]

[0062] where T f represents the number of successful grasps in the 4 scenarios with external force applied, T 10 represents the number of successful grasps in the scenarios of rotating 10 degrees clockwise or counterclockwise, T 20 represents the successful grasp in the scenarios of rotating 20 degrees clockwise or counterclockwise.

[0063] Step S4: Build the gripper shape generation network. Based on the DDPG reinforcement learning algorithm framework, build the gripper shape generation network. The DDPG algorithm adds an Actor-Critic network structure on the basis of the DQN algorithm. Its construction mainly includes a policy network and a value network. Among them, the policy network μ actor is responsible for inputting the state s at the current moment and outputting the predicted action a, that is:

[0064] a = μ actor (s),

[0065] The value network Q is responsible for inputting the current state S and the executed action A and outputting the predicted value;

[0066] q = Q(s, a).

[0067] Furthermore, step S4 is specifically as follows:

[0068] Step S41: Design the feature extraction module in the policy network. For the object shape encoding after 3D convolution, build a fully connected layer with an input of 128 dimensions and an output of 128 dimensions and a Relu activation layer to obtain the object shape representation; for the 11-dimensional shape parameter representations of the left and right fingers of the fixture, build 2 fully connected layers with an input of 11 dimensions and an output of 128 dimensions and a Relu activation layer respectively to obtain the left and right shape representations; finally, connect the object shape representation with the left and right shape representations of the fixture into a 128×3-dimensional vector as the shape feature extracted from the state representation;

[0069] Step S42: Design the action selection module in the policy network. Using the 128×3-dimensional shape feature obtained in step S41 as the input, build an action selection network composed of 3 layers of fully connected layers and Relu activation layers, and finally output an action vector a = [Δs l , Δs r , where the first 11-dimensional vector represents the shape parameters of the left variable shape plate of the fixture, and the last 11-dimensional vector represents the shape parameters of the right variable shape plate of the fixture.

[0070] Step S43: Construct the value network. Specifically, first in the feature extraction part, build a fully connected layer with an input dimension of 22 dimensions and an output dimension of 128 dimensions and a Relu activation layer to process the 22-dimensional action vector obtained in step S42 into a 128-dimensional action representation feature. Secondly, the feature extraction of the object shape and the left and right shape representations of the fixture is consistent with the processing flow described in step S41. Finally, build a value prediction module with an input dimension of 128×4 dimensions and an output dimension of 1 dimension, and predict the current value based on the connection vector of the object shape, the left and right shape of the fixture, and the action feature.

[0071] Step S5: Network update and training;

[0072] Furthermore, step S5 is specifically as follows:

[0073] Step S51: Initialize the state space and initialize the simulation environment settings. For the initialization of the fixture shape, as described in step S13, the left and right fingers of the fixture are mainly controlled by the variable shape plates, which can be represented as an 11-dimensional vector in the form of . During the initialization process, set the first 10-dimensional shape parameters corresponding to the left and right fingers to 20 respectively, and set the last 1-dimensional height parameter to 0; for the initialization of the simulation environment, after each execution of the grasping scenario of the current object, replace the object and restore the fixture height to the state before grasping the object;

[0074] Step S52: Based on the DDPG algorithm framework, construct 2 policy networks and value networks with the same structure, where 1 group of policy networks and value networks is called the target network. Specifically, for the current time t, the policy network is represented as μ(s t |θ μ ); the value network is represented as Q(s t , a t |θ Q ); the weights of the target policy network and the target value network are copied from the policy network and the evaluation network, and are represented as μ′(s t |θ μ′ ) and Q′(s t+1 , a t+1 |θ Q′ ); that is, at regular intervals of training steps, θ μ → θ μ′ , θ Q → θ Q′ , where θ μ , θ Q represent the parameters of the current policy network and value network, and θ μ′ , θ Q′ represent the parameters of the current target policy network and target value network.

[0075] Step S53: Set the maximum number of training rounds E (E = 20), the size of the experience pool M (M = 10000). Based on the DDPG algorithm, explore the fixture shape space using a random policy before the experience pool is filled. Each time after executing the action a t , update the parameter representation of the left and right fingers of the fixture shape, so as to update the state representation s t to s t+1 . After the experience pool is filled and learning starts, the action selection policy will use the action predicted by the policy network as the action to be executed. The policy network and the value network randomly select a part of the samples from the experience pool for training. The Critic network calculates the current value, and the learning process is:

[0076] q t = r t + γQ′(s t+1 , μ′(s t+1 ∣θ μ′ )∣θ Q′ ),

[0077] where r t is the reward obtained at the current time, Q′ is the target value network, θ μ′ , θ Q′denote the parameters of the current target policy network and target value network, γ ∈ [0, 1] is the discount factor, which is used to calculate the cumulative return value of the network. The larger its value, the more the value network focuses on long-term benefits;

[0078] Step S54: Update the network parameters based on the loss function. The value network aims to make an accurate evaluation of the current state and action. The calculation formula of the loss function is as follows:

[0079]

[0080] The policy network aims to maximize the state value under action selection. The calculation formula of the loss function is as follows:

[0081] Loss Actor =-Q(s t ,a t ∣θ Q ).

[0082] According to the gradient descent method, update the network parameters:

[0083]

[0084] where denotes the policy gradient of the policy network under the parameter θ μ , and denote the evaluation network state-action value function gradient and the policy network policy function gradient respectively. After a certain training step, the target network weights are updated as follows:

[0085]

[0086] where τ is the soft update ratio coefficient. Specifically, the value network obtains the value loss function at the current iteration according to the actual value q t calculated by the current value network and the value of the next state-action predicted by the target value network. Use the above loss function to train the value network at the current iteration to obtain the value network at the next iteration. The policy network aims to maximize the value. The overall detailed training process follows the DDPG algorithm, as Figure 3 shown.

[0087] Step S6: Automatically generate the fixture shape based on the given object. Use the trained policy network to implement action selection. Specifically, given the object to be grasped, as described in step S21, the initial object is processed into a 128-dimensional shape encoding, and combined with the initial fixture shape to obtain the initial state s 0 , and input the state s 0 into the policy network to obtain the predicted action a 0, after performing the action, the shape of the fixture changes, and the status is updated to s 1 , and so on, to achieve a shape change with a maximum of 20 steps, thereby automatically generating the final fixture shape, as Figure 4 shown.

[0088] The present invention elaborates on the principle and implementation manner, but the above embodiments are not limitations of the present invention. They only help to understand the method and idea of the present invention. Any other changes, modifications, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are within the protection scope of the present invention.

Claims

1. A parallel fixture shape generation method based on DDPG reinforcement learning algorithm, characterized in that the steps include: Step S1: parallel fixture shape space design; Step S2: state representation and action space design; Step S3: Setting up the grabbing reward mechanism and using PyBullet physics simulation to evaluate grabbing; Step S4: Build the fixture shape generation network and construct the strategy network μ actor and value network Q, where the strategy network μ actor Responsible for inputting the current state s and outputting the predicted action a, that is: a=μ actor (s), The value network Q is responsible for inputting the current state s and the action a performed, and outputting the predicted value, i.e.; q=Q(s,a). Step S5: updating and training the fixture shape generation network based on the DDPG reinforcement learning algorithm; Step S6: automatically generate a fixture shape based on a given object, use the trained policy network to implement action selection, and achieve a shape change with a maximum number of steps of 20, thereby automatically generating the final fixture shape; The step S1 is specifically as follows: Step S11: constructing a basic shape template of the fixture shape based on the WSG50 fixture; Step S12: designing an L-shaped finger consisting of a fixed plate and a shape-variable plate, and obtaining a fixture shape; Step S13: the shape of the L-shaped finger is represented by parameters, the length of the clamping teeth is controlled by the length parameter, and then the shape of the shape-variable plate is controlled, and the installation position is controlled by the height parameter; Step S14: Implement the mapping from the 11-dimensional vector to the fixture shape space, that is, obtain the corresponding left and right finger shapes of the fixture respectively according to the 11-dimensional parameter representation.

2. According to claim 1, a parallel fixture shape generation method based on DDPG reinforcement learning algorithm is characterized in that: The step S2 is specifically as follows: Step S21: abstractly represent the geometric shape of the object and the shape of the gripper fingers; specifically: use the truncated signed distance function (TSDF) and the shape encoder based on 3D convolution to obtain a 128-dimensional object shape encoding; use the 11-dimensional vector of the left gripper teeth and the 11-dimensional vector of the right gripper teeth as the parameter representation of the gripper finger shape; Step S22: defining a state space including object coding and shape parameter representation of the left and right fingers of the gripper; Step S23: taking the parameter changes of the left and right gripper fingers as actions, and further defining the action space; Step S24: Define action constraints in the action space.

3. A parallel fixture shape generation method based on DDPG reinforcement learning algorithm according to claim 1 or 2, characterized in that: The step S3 is specifically as follows: Step S31: various parameters of the simulated grasping environment are set based on Pybullet, and nine different grasping scenarios are set from three aspects: grasping success, grasping stability, and grasping robustness; Step S32: Set up a reward score calculation mechanism to calculate the reward of the current environment after each action is performed, as follows: Among them, T f represents the number of successful grasping in the four situations where external force is applied, T 10 represents the number of successful grasps in the case of a 10-degree clockwise or counterclockwise rotation, T 20 Indicates a successful grasp in a 20 degree clockwise or counterclockwise rotation.

4. A parallel fixture shape generation method based on DDPG reinforcement learning algorithm according to claim 1 or 2, characterized in that: The step S4 is specifically as follows: Step S41: using the fully connected layer and the activation layer, designing a feature extraction module in the strategy network, processing the object shape representation and the left and right shape representations of the fixture as a 128×3-dimensional vector as the shape feature extracted from the state representation; Step S42: Designing an action selection module in the strategy network using the fully connected layer and the activation layer; Step S43: Use the fully connected layer and the activation layer to build a value network, and predict the current value based on the connection vector of the object shape, the left and right shapes of the fixture, and the action features.

5. The parallel fixture shape generation method based on DDPG reinforcement learning algorithm according to claim 3 is characterized in that: The step S4 is specifically as follows: Step S41: using the fully connected layer and the activation layer, designing a feature extraction module in the strategy network, processing the object shape representation and the left and right shape representations of the fixture as a 128×3-dimensional vector as the shape feature extracted from the state representation; Step S42: Designing an action selection module in the strategy network using the fully connected layer and the activation layer; Step S43: Use the fully connected layer and the activation layer to build a value network, and predict the current value based on the connection vector of the object shape, the left and right shapes of the fixture, and the action features.

Citation Information

Patent Citations

  • Robot adaptive grabbing method based on deep reinforcement learning

    CN106094516A

  • Mechanical arm end effector grabbing posture adjusting method and system based on reinforcement learning

    CN110909644A