Robot multi-finger grabbing method and device
By fusing visual and proprioceptive perception information and using a strategy generation model to optimize grasping actions, the stability problem of the robot's multi-fingered hand in complex scenarios is solved, enabling stable grasping of objects by the multi-fingered hand and improving the reliability of grasping.
Patent Information
- Application Number
- CN202510882635.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-11-25
AI Technical Summary
There is still room for improvement in the stability of existing multi-finger grasping operations in robots, especially when dealing with complex scenes and objects with low coefficient of friction, where there are problems such as uneven force application and object slippage.
By acquiring visual information and robot body perception information from the object grasping task scenario, cross-modal complementary information fusion is performed, and a policy generation model is used for training to generate grasping actions to control the multi-fingered hand to grasp objects. The policy generation model is generated based on minimizing the loss function and the target reward function, and the grasping actions are optimized to improve stability.
The robot's multi-fingered hand has improved its grasping stability in complex scenarios. Through cross-modal information fusion and optimization of the strategy generation model, it has achieved stable grasping of objects by the multi-fingered hand, thus improving the reliability and stability of grasping.
Smart Images

Figure CN121004598A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics technology, specifically to a method and apparatus for grasping with a multi-fingered hand of a robot. Background Technology
[0002] Robotic grasping operations have wide applications in industrial production, home services, and healthcare. However, due to the diverse shapes of objects, different surface materials, and complex scenarios, achieving intelligent and stable grasping remains a key technical challenge that urgently needs to be solved in the field of robotics.
[0003] Currently, research on robotic grasping technology can be broadly categorized into two types: two-finger gripper grasping technology and multi-finger dexterous hand grasping technology. Two-finger gripper grasping, where the end effector is a traditional two-finger parallel gripper, involves analyzing the target or grasping scenario and controlling the opening and closing of the gripper at the end of the robotic arm to grasp the object. However, due to hardware limitations, the two-finger gripper can only provide two points of contact force when grasping objects. When handling objects with complex geometries or low friction coefficients, it often faces problems such as uneven force application and object slippage, limiting its reliability in complex scenarios. Multi-finger dexterous hand grasping, where the end effector is a humanoid multi-finger dexterous hand, involves controlling the rotation of motors at multiple joints of the dexterous hand to grasp objects. This method possesses the ability for multi-contact point collaborative control. By adjusting the gripping direction and force applied by the multiple fingers, it can grasp objects of different shapes and physical properties, and the grasping posture is more flexible and expandable.
[0004] In industrial production, object sorting has always been done manually, which is time-consuming, labor-intensive, and rarely yields satisfactory results. While existing robotic multi-finger grasping operations can grasp objects and improve sorting efficiency to some extent, the stability of the grasp still has considerable room for improvement. Summary of the Invention
[0005] This application provides a method and apparatus for robotic multi-finger grasping, which addresses the technical problem that the stability of existing robotic multi-finger grasping still has considerable room for improvement.
[0006] In a first aspect, embodiments of this application provide a robotic multi-finger grasping method, including: Acquire visual information of the object grasping task scenario and the robot's proprioceptive information; The visual information and the ontological perception information are fused to obtain cross-modal complementary information; The cross-modal complementary information is input into the strategy generation model to obtain the grasping action output by the strategy generation model; Based on the grasping action, control the multi-finger hand to grasp objects; The policy generation model is trained based on historical cross-modal complementary information, with the goal of minimizing the loss function, on the basis of the policy network. The loss function is generated based on the target reward function, which is the reward function after fusing the force closure evaluation when the multi-finger hand grasps the object.
[0007] In one embodiment, the policy generation model is obtained based on the following steps: Historical cross-modal complementary information is input into the policy network to obtain the historical crawling actions output by the policy network; Based on the historical grasping actions, the multi-finger hand is controlled to grasp objects; The contact points between the multi-fingered hand and the object being grasped are clustered to obtain multiple clusters; Based on the center points of the multiple clusters and their corresponding average force vectors, the grasping stability score is calculated. Based on the grasping stability score, calculate the target reward value for the historical grasping action; Based on the target reward value, calculate the loss value of the policy network; If the loss value converges, then the policy network at this point is determined as the policy generation model; If the loss value does not converge, the parameters of the policy network are adjusted, and the process returns to the step of inputting historical cross-modal complementary information into the policy network to obtain the historical grabbing actions output by the policy network, until the loss value converges. At this point, the policy network is determined as the policy generation model.
[0008] In one embodiment, the historical grasping action includes historical hand end pose and historical finger grasping features, and controlling the multi-finger hand to grasp objects based on the historical grasping action includes: Based on the main grasping feature vector, the historical finger grasping features are convexly combined to obtain the historical target finger joint position; The historical target finger joint positions are subjected to an exponential moving average to obtain the final historical finger joint positions. The historical hand end pose is input into the robotic arm controller to generate robotic arm control commands. The historical final finger joint position is input into the multi-finger hand controller to generate multi-finger hand control commands; The robot's robotic arm is controlled to move according to the robotic arm control command until the palm end pose of the multi-fingered hand matches the historical palm end pose. The robot's multi-finger hand movement is controlled based on the multi-finger hand control commands until the finger joint positions of the multi-finger hand match the historical final finger joint positions.
[0009] In one embodiment, the primary feature vector is obtained based on the following steps: Select target objects that have typical set characteristics; Based on the target object, generate multiple grasping postures of the multi-fingered hand; From the various grasping postures, select multiple target postures that conform to human grasping habits; Principal component dimensionality reduction is performed on the various target poses to obtain the main grasping feature vectors corresponding to the main target poses.
[0010] In one embodiment, calculating the loss value of the policy network based on the target reward value includes: Acquire historical tactile information of the multi-finger hand and historical target information of the grasped object, corresponding to the historical cross-modal complementary information; the historical target information includes historical state information and historical expected target. The historical cross-modal complementary information, the historical tactile information, and the historical target information are input into the evaluation network to obtain the evaluation value of the evaluation network for the historical grasping action; The evaluation value and the target reward value are input into the advantage function to obtain the advantage value of the historical grabbing action output by the advantage function; The advantage value is input into the loss function to obtain the loss value.
[0011] In one embodiment, the visual information of the object grasping task scene is obtained based on the following steps: Obtain a depth image of the object grasping task scene; Multiple image points are obtained from the depth image; Randomly select one image point from the plurality of image points as the starting point to construct a sampling point set; Calculate the distances between the plurality of image points and the starting point to obtain a plurality of first distances; Select the maximum distance among the plurality of first distances, determine the image point corresponding to the maximum distance as the new starting point, and add it to the sampling point set; Calculate the distances between the plurality of image points and the new starting point to obtain a plurality of second distances; For each image point, retain the smaller of the first and second distances to obtain multiple third distances; Select the maximum distance among the plurality of third distances, determine the image point corresponding to the maximum distance as the new starting point, and add it to the sampling point set; After taking the multiple third distances as new multiple first distances, return to the step of calculating the distance between the multiple image points and the new starting point to obtain multiple second distances, until the number of image points in the sampling point set reaches a preset number, and the target image point is obtained; The target image points are input into a point cloud network to obtain the visual information.
[0012] In one embodiment, the robot's body perception information includes the robot's robotic arm information and multi-fingered hand information.
[0013] Secondly, embodiments of this application provide a robotic multi-finger grasping device, comprising: The information acquisition module is used to acquire visual information of the object grasping task scenario and the robot's body perception information. The information fusion module is used to: fuse the visual information and the ontology perception information to obtain cross-modal complementary information; The action generation module is used to: input the cross-modal complementary information into the strategy generation model to obtain the grasping action output by the strategy generation model; The object grasping module is used to control a multi-finger hand to grasp objects based on the grasping action. The policy generation model is trained based on historical cross-modal complementary information, with the goal of minimizing the loss function, on the basis of the policy network. The loss function is generated based on the target reward function, which is the reward function after fusing the force closure evaluation when the multi-finger hand grasps the object.
[0014] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory storing a computer program, wherein the processor executes the program to implement the steps of the robotic multi-finger grasping method described in the first aspect.
[0015] Fourthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the robotic multi-fingered hand grasping method described in the first aspect.
[0016] The robot multi-finger grasping method and apparatus provided in this application acquires visual information and robot proprioceptive information of the object grasping task scenario, fuses the visual information and proprioceptive information to obtain cross-modal complementary information, inputs the cross-modal complementary information into a policy generation model, obtains the grasping action output by the policy generation model, and controls the multi-finger hand to grasp objects based on the grasping action. The policy generation model is trained on the basis of a policy network with the goal of minimizing the loss function and based on historical cross-modal complementary information. The loss function is generated based on the target reward function, which is the reward function after fusing the force closure evaluation when the multi-finger hand grasps the object. In this application, visual information outside the robot body and perceptual information inside the robot body are first fused to obtain rich cross-modal complementary information. Then, this cross-modal complementary information is input into the policy generation model to output grasping actions. This enables the model to fully capture information inside and outside the robot body, guiding the model to output grasping actions that conform to the grasping scenario and the robot's perception. At the same time, since the policy generation model aims to minimize the loss function generated based on the target reward function, and the target reward function is the reward function after fusing the force closure evaluation when the multi-finger grasps the object, this means that the model will reward grasping actions with good force closure evaluation during the training process, encouraging the model to use more of these grasping actions to minimize the loss function. Therefore, the grasping actions output by the final policy generation model have high force closure evaluation performance, that is, the position distribution when the multi-finger grasps the object is effectively allocated. Controlling the multi-finger to grasp objects based on this grasping action can effectively improve grasping stability. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is one of the flowcharts illustrating the robot multi-finger grasping method provided in the embodiments of this application; Figure 2 This is a second flowchart illustrating the robot multi-finger grasping method provided in this application embodiment; Figure 3 This is a schematic diagram of the strategy generation model training in the robot multi-finger grasping method provided in the embodiments of this application; Figure 4 This is a schematic diagram of the cluster center points and their average force vectors in the robot multi-finger grasping method provided in this application embodiment; Figure 5This is the third flowchart illustrating the robot multi-finger grasping method provided in this application embodiment; Figure 6 This is a schematic diagram illustrating the generation of key grasping feature vectors in the robot multi-finger grasping method provided in this application embodiment; Figure 7 This is the fourth flowchart of the robot multi-finger grasping method provided in the embodiments of this application; Figure 8 This is the fifth flowchart illustrating the robot multi-finger grasping method provided in the embodiments of this application; Figure 9 This is a schematic diagram of the structure of the robotic multi-finger grasping device provided in the embodiments of this application; Figure 10 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] It should be noted that in the description of the embodiments of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. The terms "upper," "lower," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. Unless otherwise expressly specified and limited, the terms "installed," "connected," and "linked" should be interpreted broadly, for example, they can be fixed connections, detachable connections, or integral connections; they can be mechanical connections or electrical connections; they can be direct connections or indirect connections through an intermediate medium; and they can be internal connections between two elements. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.
[0021] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects; for example, a first object can be one or more. Furthermore, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects have an "or" relationship.
[0022] Figure 1 This is one of the flowcharts illustrating the robot multi-finger grasping method provided in this application embodiment; see reference. Figure 1 This application provides a robotic multi-finger grasping method, which may include: 101. Acquire visual information of the object grasping task scene and the robot's proprioceptive information; 102. By fusing visual information and proprioceptive information, cross-modal complementary information is obtained; 103. Input cross-modal complementary information into the strategy generation model to obtain the grasping action output by the strategy generation model; 104. Controlling multi-finger hand to grasp objects based on grasping motion.
[0023] The policy generation model is trained based on historical cross-modal complementary information, with the goal of minimizing the loss function, on the basis of the policy network. The loss function is generated based on the target reward function, which is the reward function after fusing the force closure evaluation when a multi-finger hand grasps an object.
[0024] In step 101, the visual information can be the visual point cloud information of the object grasping scene, and the dimension can be... This includes 512 3D points; the robot's body perception information can include the robot's robotic arm information and multi-finger hand information. The robotic arm information can include the 3D robotic arm end position and the 4D robotic arm end pose, and the multi-finger hand information can include 16-dimensional multi-finger hand joint angles, 3D multi-finger hand palm position, and 4D multi-finger hand palm pose, etc.
[0025] Visual information from the object grasping task scenario and the robot's proprioceptive information help the multi-fingered hand to robustly grasp objects of different types and sizes.
[0026] In step 104, the grasping action can be converted into a robot control command and transmitted to the robot, thereby controlling the robot's multi-fingered hand to grasp objects using the grasping action.
[0027] The multi-finger grasping method for robots provided in this embodiment acquires visual information and the robot's proprioceptive information of the object grasping task scenario, fuses the visual information and proprioceptive information to obtain cross-modal complementary information, inputs the cross-modal complementary information into a policy generation model, obtains the grasping action output by the policy generation model, and controls the multi-finger to grasp objects based on the grasping action. The policy generation model is trained on the basis of a policy network with the goal of minimizing the loss function, based on historical cross-modal complementary information. The loss function is generated based on the target reward function, which is the reward function after fusing the force closure evaluation when the multi-finger grasps the object. In this embodiment, visual information from outside the robot body and perceptual information from inside the robot body are first fused to obtain rich cross-modal complementary information from inside and outside the robot body. Then, this cross-modal complementary information is input into the strategy generation model to output grasping actions. This enables the model to fully capture information from inside and outside the robot body, guiding the model to output grasping actions that conform to the grasping scenario and the robot's perception. At the same time, since the strategy generation model aims to minimize the loss function generated based on the target reward function, and the target reward function is the reward function after fusing the force closure evaluation when the multi-finger grasps the object, this means that the model will reward grasping actions with good force closure evaluation during the model training process, encouraging the model to use more of these grasping actions to minimize the loss function. Therefore, the grasping actions output by the final strategy generation model have high force closure evaluation performance, that is, the position distribution when the multi-finger grasps the object is effectively allocated. Controlling the multi-finger to grasp objects based on this grasping action can effectively improve grasping stability.
[0028] Figure 2 This is a second schematic flowchart of the robot multi-finger grasping method provided in the embodiments of this application; see also Figure 2 and Figure 3 In one embodiment, the policy generation model can be obtained based on the following steps: 201. Input historical cross-modal complementary information into the policy network to obtain the historical grabbing actions output by the policy network; 202. Controlling multi-finger hand to grasp objects based on historical grasping actions; 203. Cluster the contact points between the multi-fingered hand and the object being grasped to obtain multiple clusters; 204. Calculate the grasping stability score based on the center points of multiple clusters and their corresponding average force vectors; 205. Based on the grasping stability score, calculate the target reward value for historical grasping actions; 206. Calculate the loss value of the policy network based on the target reward value; 207. If the loss value converges, then the policy network at this point is determined as the policy generation model; 208. If the loss value does not converge, adjust the parameters of the policy network and return to step 201 until the loss value converges. Then, determine the policy network at this point as the policy generation model.
[0029] In step 201, historical visual information and historical ontology perception information can be bilinearly fused based on the following formula to obtain historical cross-modal complementary information. : ; in, For historical ontology to perceive information, It is historical visual information. Historical weighting, This is a historical bias term.
[0030] Will Input Policy Network This allows us to obtain the historical capture actions of its output. .
[0031] In step 203, any clustering algorithm can be used to cluster the contact points between the multi-finger hand and the grasped object. There is no limitation here. In this embodiment, the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm is used to cluster the contact points on the same finger of the multi-finger hand into a cluster, thereby obtaining multiple clusters corresponding to multiple fingers.
[0032] In step 204, it is assumed that The center point of each cluster is Its corresponding average force vector is The capture stability score can then be calculated using the following formula. : ; in: ; ; ; in, for The identity matrix, for The cross product matrix, as shown above, is the vector... The first three dimensions represent the force balance, and the last three dimensions represent the torque balance. By calculating vectors The model is used to reflect the overall force closure situation; Figure 4 for hour and A schematic diagram.
[0033] In step 205, the target reward function can be calculated according to the following formula. Value: ; in: ; For multi-finger hand strength closure reward, Normalization to reduce the impact of the number of contact points on force closure assessment; ; The reward is the distance between the multi-fingered hand and the object being grasped. The position of a single finger of a multi-fingered hand at a historical moment. The location of the object captured at a historical moment. The center of the palm of the hand is often used to describe historical moments. The influencing factor for fingers, The palm-related factor; ; This is the distance reward between the position of the grasped object at a historical moment and its historical target position; in other words, it is the reward for the multi-finger hand to lift the grasped object towards its historical target position. The historical target location of the object being captured. The influence factor for lifting objects; ; For speed and posture rewards, This refers to the robot's historical movement speed, including the historical movement speed of the fingers and the historical movement speed of the hand. The initial historical posture of the object being grasped. The pose of the object being grasped at a specific historical moment, that is, the pose at a certain historical moment after the initial historical moment. It is an influencing factor on the speed of historical movement. It is an influencing factor on the amount of historical attitude change. and This is to prevent collisions between fingers caused by the robot's multi-fingered hand moving too fast, and to prevent grasping failure due to excessive changes in the posture of the object being grasped.
[0034] , , , This refers to the weight of the corresponding reward.
[0035] In steps 206 to 208, the higher the target reward value, the more the parameters of the policy network will be adjusted through backpropagation to encourage the policy network to select more corresponding historical crawling actions and update in the direction of reducing the loss value. Therefore, the loss value of the policy network can be calculated based on the target reward value, and the parameters of the policy network can be adjusted when the loss value has not converged to update in the direction of reducing the loss value. Finally, when the loss value reaches the minimum, the policy network at this time is determined to be the policy generation model.
[0036] This embodiment clusters the contact points between the multi-finger hand and the grasped object according to the fingers involved, resulting in multiple clusters. The center point of each cluster is determined as the average force point of the corresponding finger. Then, based on the average force point and its corresponding average force vector, a grasping stability score is calculated to evaluate the grasping closure force. This score, along with other rewards, generates the target reward value of the historical grasping actions output by the policy network to guide the parameter update of the policy network. Ultimately, the loss value of the policy network is minimized, resulting in a policy generation model. This policy generation model can generate grasping actions with better closure force evaluation based on a combination of factors, including the closure force between the multi-finger hand and the grasped object, thus achieving stable grasping of the object.
[0037] Figure 5 This is the third flowchart illustrating the robot multi-finger grasping method provided in this application embodiment; see also... Figure 3 and 5 In one embodiment, the historical grasping action includes the historical hand end pose and historical finger grasping features. Controlling a multi-finger hand to grasp objects based on the historical grasping action may include: 501. Based on the main grasping feature vector, perform a convex combination of historical finger grasping features to obtain the historical target finger joint position; 502. Perform an exponential moving average on the historical target finger joint positions to obtain the final historical finger joint positions; 503. Input the historical hand end-effector pose into the robotic arm controller to generate robotic arm control commands; 504. Input the final finger joint position from the past into the multi-finger hand controller to generate multi-finger hand control instructions; 505. Control the movement of the robot's robotic arm based on the robotic arm control instructions until the palm end pose of the multi-fingered hand matches the historical palm end pose; 506. Control the movement of the robot's multi-finger hand based on multi-finger hand control instructions until the finger joint positions of the multi-finger hand match the historical final finger joint positions.
[0038] In step 501, the main feature vector can be obtained based on the following steps: 501a. Select target objects with typical set characteristics; 501b. Generate multiple grasping postures for a multi-fingered hand based on the target object; 501c: Select multiple target postures that conform to human grasping habits from a variety of grasping postures; 501d. Principal component dimensionality reduction is performed on multiple target poses to obtain the main grasping feature vectors corresponding to the main target poses.
[0039] Reference Figure 6 : In step 501a, target objects with typical set features can be selected from the ShapeNet dataset. ShapeNet is one of the most widely used 3D object datasets, containing over 5,000 object categories and over 300,000 3D models. This dataset covers a wide range of categories, from furniture and vehicles to natural objects, providing rich benchmark data for tasks such as 3D object recognition, generation, and segmentation.
[0040] In step 501b, Dexgraspnet can be used to generate models that, based on the selected target objects, generate diverse grasping poses corresponding to these target objects.
[0041] In step 501c, grasping postures that do not conform to human grasping habits are removed from these diverse grasping postures, thereby forming a more realistic grasping prior database that resembles a human hand.
[0042] In step 501d, the action space is reduced for the target pose data in the prior grasping database to filter out the main target poses and obtain their corresponding main grasping feature vectors, as follows: 1. Based on the following formula, the target attitude data is analyzed. Standardization processing is performed to obtain standardized attitude data. : ; in, The average of multiple target pose data. The standard deviation is the set of multiple target attitude data.
[0043] 2. Calculate the covariance matrix : ; in, The number of target pose data.
[0044] 3. Eigenvalue decomposition: ; in, For the first 1 eigenvector For the first Each feature value.
[0045] 4. Select the top eigenvalues from the descending order based on the desired percentage of cumulative variance to be retained. The eigenvectors corresponding to each eigenvalue , , These feature vectors are the main feature vectors to be captured.
[0046] The historical target finger joint position can then be obtained based on the following formula. : ; in, To and The corresponding historical finger grasping features can be used to obtain 16-dimensional historical target finger joint positions, representing the 16 joint positions of the fingers in a multi-fingered hand.
[0047] In step 502, the historical final finger joint position can be obtained based on the following formula. : ; in, It refers to the finger joint position at a historical moment before reaching the target finger joint position. This is the smoothing coefficient.
[0048] Exponential moving averages can reduce abrupt changes in joint torque, ensuring a smooth motion trajectory.
[0049] In steps 503 to 506, the historical palm end pose and the historical final finger joint position are input into the robotic arm controller and the multi-finger hand controller, respectively. The controller generates corresponding control commands. During the action execution phase, the robotic arm controller and the multi-finger hand controller issue control commands to control the movement of the robotic arm and the movement of the multi-finger hand until the palm end pose of the multi-finger hand matches the historical palm end pose and the finger joint position of the multi-finger hand matches the historical final finger joint position, so that the multi-finger hand completes the object grasping action using the historical grasping action.
[0050] It should be noted that this embodiment adopts a parallel training method for robots, setting up multiple robots to train simultaneously and storing the training data in an experience replay pool to accelerate the learning efficiency of the strategy.
[0051] This embodiment performs principal component screening and combination on the historical finger grasping features output by the policy network to obtain the historical target finger joint positions. This reduces the dimension of the action space and improves the learning efficiency of the policy network. Then, an exponential moving average is applied to ensure the smoothness of the motion trajectory to obtain the historical final finger joint positions. Finally, the historical palm end pose and the historical final finger joint positions are input into the corresponding controller to generate control commands to control the movement of the robotic arm and the multi-finger hand respectively. This allows the palm end pose of the multi-finger hand to accurately match the historical palm end pose, and the finger joint positions of the multi-finger hand to accurately match the historical final finger joint positions, thereby enabling the multi-finger hand to accurately follow the historical grasping actions to perform stable object grasping.
[0052] Figure 7 This is the fourth flowchart illustrating the robot multi-finger grasping method provided in this application embodiment; see also... Figure 3 and 7 In one embodiment, calculating the loss value of the policy network based on the target reward value may include: 701. Acquire historical tactile information of the multi-finger hand and historical target information of the grasped object, which correspond to historical cross-modal complementary information; Historical target information includes historical status information and historical expected targets; 702. Input historical cross-modal complementary information, historical tactile information, and historical target information into the evaluation network to obtain the evaluation value of the evaluation network for historical grasping actions; 703. Input the evaluation value and target reward value into the dominance function to obtain the dominance value of the historical grabbing actions output by the dominance function; 704. Input the advantage value into the loss function to obtain the loss value.
[0053] In step 701, the historical tactile information refers to the contact between the finger joints and the grasped object, which can be represented using binary encoding; the historical state information can include the position and posture of the grasped object, and the historical desired target can be a pre-set object target position; then the evaluation network... and policy network The observation information and its feature dimensions used in the simulation environment are shown in the table below: Table 1 Comparison of Evaluation Network and Policy Network in, To evaluate the network parameters, These are the parameters of the policy network.
[0054] In step 702, from Table 1 and Figure 3It can be seen that the input of the evaluation network, based on the policy network, further integrates historical state information, historical tactile information, and historical expected goals, so as to evaluate the historical grasping actions generated by the policy network from multiple perspectives.
[0055] In step 703, the advantage value of the historical grabbing actions output by the advantage function can be calculated based on the following formula: ; in, Representing history The advantage value output by the advantage function at any given time. Representing history The target reward value output by the target reward function at time step [time]. Representing history Continuously evaluate the evaluation value output by the network. Representing history The evaluation value output by the network is evaluated at all times.
[0056] In step 704, the loss value of the policy network can be calculated based on the following formula: ; in, Representing history The loss value output by the loss function at time step. Representing history The time-based strategy network outputs historical capture actions. The probability of.
[0057] This embodiment integrates historical state information, historical tactile information, and historical expected target input into the evaluation network based on historical cross-modal complementary information. Then, based on the evaluation value output by the evaluation network and the target reward value, the advantage value output by the advantage function is obtained. Based on the advantage value and the probability of historical grasping actions output by the policy network, the loss value is obtained, forming a backpropagation path of the target reward value, ultimately minimizing the loss value.
[0058] Figure 8 This is the fifth flowchart illustrating the robot multi-finger grasping method provided in this application embodiment; see also... Figure 3 and 8 In one embodiment, the visual information of the object grasping task scene can be obtained based on the following steps: 801. Obtain depth images of the object grasping task scene; 802. Obtain multiple image points from a depth image; 803. Randomly select one image point from multiple image points as the starting point to construct a set of sampling points; 804. Calculate the distances between multiple image points and the starting point to obtain multiple first distances; 805. Select the maximum distance among multiple first distances, determine the image point corresponding to the maximum distance as the new starting point, and add it to the sampling point set; 806. Calculate the distances between multiple image points and the new starting point to obtain multiple second distances; 807. Retain the smaller of the first and second distances corresponding to each image point to obtain multiple third distances; 808. Select the maximum distance among multiple third distances, determine the image point corresponding to the maximum distance as the new starting point, and add it to the sampling point set; 809. After treating multiple third distances as multiple new first distances, return to step 806; 810. When the number of image points in the sampling point set reaches a preset number, the target image points are obtained; 811. Input the target image points into the point cloud network to obtain visual information.
[0059] In the object grasping task scenario, the robot's robotic arm is a six-axis robotic arm with a multi-finger dexterous hand at its end. The maximum load weight of the robotic arm is 5KG, and its working plane is a square area of 20*20 square centimeters. The multi-finger dexterous hand is controlled by 16 joint motors. Different types and sizes of objects are randomly placed in the area. A depth camera is placed in front of the area to obtain the global view of the scene.
[0060] In step 801, a depth image of the object grasping task scene can be obtained using a depth camera.
[0061] In steps 802 to 805, since there is only one starting point in the initial sampling point set, there is only a unique first distance between each image point and the starting point. This first distance can be considered as the minimum distance between the corresponding image point and the starting point. Then, the image point corresponding to the maximum distance is selected from these first distances, and it is used as the new starting point and added to the sampling point set.
[0062] In steps 806 to 808, since there are two starting points in the sampling point set, after obtaining the second distance between multiple image points and the new starting point, for each image point, there is a first distance between it and the old starting point and a second distance between it and the new starting point. At this time, the larger distance between the first distance and the second distance is removed, and the smaller distance is retained to obtain the minimum distance corresponding to each image point, that is, the third distance. Among these third distances, the image point corresponding to the largest distance is selected, and it is used as the new starting point and added to the sampling point set.
[0063] In steps 809 to 810, after using the third distance as the new first distance, the process returns to step 806. That is, the distances between multiple image points and the new starting point are calculated and compared with the new first distance. The larger distances are discarded, and the smaller distances are retained. Then, the image point corresponding to the largest distance is selected from the smaller distances and used as the new starting point. This process is repeated until the number of image points in the sampling point set reaches a preset number, such as 512. These 512 image points are then determined as the target image points.
[0064] In step 811, these target image points can be input into any point cloud network, which is not limited here. In this embodiment, these target image points can be input into the PointNet network to obtain visual point cloud information.
[0065] This embodiment calculates the distance between multiple image points and the starting point. For each image point, the minimum distance between it and the starting point is retained. Then, the image point with the maximum distance from the minimum distance is selected as the new starting point and added to the sampling point set. The distance is then recalculated, and so on, until the number of image points in the sampling set reaches a preset number. This achieves the sampling of image points in the depth image. The sampling points obtained by this sampling method are evenly distributed, while retaining the key geometric features of the point cloud.
[0066] The following describes the robotic multi-finger grasping device provided in the embodiments of this application. The robotic multi-finger grasping device described below and the robotic multi-finger grasping method described above can be referred to and correspond to each other.
[0067] Figure 9 This is a schematic diagram of the structure of the robotic multi-finger grasping device provided in an embodiment of this application. (Refer to...) Figure 9 This application provides a robotic multi-finger grasping device, which may include: The information acquisition module 901 is used to: acquire visual information of the object grasping task scene and the robot's body perception information; The information fusion module 902 is used to: fuse the visual information and the ontology perception information to obtain cross-modal complementary information; The action generation module 903 is used to: input the cross-modal complementary information into the strategy generation model to obtain the grasping action output by the strategy generation model; The object grasping module 904 is used to: control a multi-finger hand to grasp objects based on the grasping action; The policy generation model is trained based on historical cross-modal complementary information, with the goal of minimizing the loss function, on the basis of the policy network. The loss function is generated based on the target reward function, which is the reward function after fusing the force closure evaluation when the multi-finger hand grasps the object.
[0068] The multi-finger grasping device provided in this embodiment acquires visual information and the robot's proprioceptive information of the object grasping task scenario. It fuses the visual information and proprioceptive information to obtain cross-modal complementary information. The cross-modal complementary information is input into a policy generation model to obtain the grasping action output by the policy generation model. Based on the grasping action, the multi-finger grasping device is controlled to grasp objects. The policy generation model is trained on the basis of a policy network with the goal of minimizing the loss function. The loss function is generated based on the target reward function, which is the reward function after fusing the force closure evaluation when the multi-finger grasps the object. In this embodiment, visual information from outside the robot body and perceptual information from inside the robot body are first fused to obtain rich cross-modal complementary information from inside and outside the robot body. Then, this cross-modal complementary information is input into the strategy generation model to output grasping actions. This enables the model to fully capture information from inside and outside the robot body, guiding the model to output grasping actions that conform to the grasping scenario and the robot's perception. At the same time, since the strategy generation model aims to minimize the loss function generated based on the target reward function, and the target reward function is the reward function after fusing the force closure evaluation when the multi-finger grasps the object, this means that the model will reward grasping actions with good force closure evaluation during the model training process, encouraging the model to use more of these grasping actions to minimize the loss function. Therefore, the grasping actions output by the final strategy generation model have high force closure evaluation performance, that is, the position distribution when the multi-finger grasps the object is effectively allocated. Controlling the multi-finger to grasp objects based on this grasping action can effectively improve grasping stability.
[0069] In one embodiment, a policy generation model building module (not shown in the figure) is also included for: Historical cross-modal complementary information is input into the policy network to obtain the historical crawling actions output by the policy network; Based on the historical grasping actions, the multi-finger hand is controlled to grasp objects; The contact points between the multi-fingered hand and the object being grasped are clustered to obtain multiple clusters; Based on the center points of the multiple clusters and their corresponding average force vectors, the grasping stability score is calculated. Based on the grasping stability score, calculate the target reward value for the historical grasping action; Based on the target reward value, calculate the loss value of the policy network; If the loss value converges, then the policy network at this point is determined as the policy generation model; If the loss value does not converge, the parameters of the policy network are adjusted, and the process returns to the step of inputting historical cross-modal complementary information into the policy network to obtain the historical grabbing actions output by the policy network, until the loss value converges. At this point, the policy network is determined as the policy generation model.
[0070] In one embodiment, the policy generation model building module is specifically used for: Based on the main grasping feature vector, the historical finger grasping features are convexly combined to obtain the historical target finger joint position; The historical target finger joint positions are subjected to an exponential moving average to obtain the final historical finger joint positions. The historical hand end pose is input into the robotic arm controller to generate robotic arm control commands. The historical final finger joint position is input into the multi-finger hand controller to generate multi-finger hand control commands; The robot's robotic arm is controlled to move according to the robotic arm control command until the palm end pose of the multi-fingered hand matches the historical palm end pose. The robot's multi-finger hand movement is controlled based on the multi-finger hand control commands until the finger joint positions of the multi-finger hand match the historical final finger joint positions.
[0071] In one embodiment, the policy generation model building module is specifically used for: Select target objects that have typical set characteristics; Based on the target object, generate multiple grasping postures of the multi-fingered hand; From the various grasping postures, select multiple target postures that conform to human grasping habits; Principal component dimensionality reduction is performed on the various target poses to obtain the main grasping feature vectors corresponding to the main target poses.
[0072] In one embodiment, the policy generation model building module is specifically used for: Acquire historical tactile information of the multi-finger hand and historical target information of the grasped object, corresponding to the historical cross-modal complementary information; the historical target information includes historical state information and historical expected target. The historical cross-modal complementary information, the historical tactile information, and the historical target information are input into the evaluation network to obtain the evaluation value of the evaluation network for the historical grasping action; The evaluation value and the target reward value are input into the advantage function to obtain the advantage value of the historical grabbing action output by the advantage function; The advantage value is input into the loss function to obtain the loss value.
[0073] In one embodiment, the information acquisition module 901 is specifically used for: Obtain a depth image of the object grasping task scene; Multiple image points are obtained from the depth image; Randomly select one image point from the plurality of image points as the starting point to construct a sampling point set; Calculate the distances between the plurality of image points and the starting point to obtain a plurality of first distances; Select the maximum distance among the plurality of first distances, determine the image point corresponding to the maximum distance as the new starting point, and add it to the sampling point set; Calculate the distances between the plurality of image points and the new starting point to obtain a plurality of second distances; For each image point, retain the smaller of the first and second distances to obtain multiple third distances; Select the maximum distance among the plurality of third distances, determine the image point corresponding to the maximum distance as the new starting point, and add it to the sampling point set; After taking the multiple third distances as new multiple first distances, return to the step of calculating the distance between the multiple image points and the new starting point to obtain multiple second distances, until the number of image points in the sampling point set reaches a preset number, and the target image point is obtained; The target image points are input into a point cloud network to obtain the visual information.
[0074] In one embodiment, the robot's body perception information includes the robot's robotic arm information and multi-fingered hand information.
[0075] Figure 10 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application, such as... Figure 10 As shown, the electronic device may include: a processor 1010, a communication interface 1020, a memory 1030, and a communication bus 1040, wherein the processor 1010, the communication interface 1020, and the memory 1030 communicate with each other via the communication bus 1040. The processor 1010 can call a computer program in the memory 1030 to execute the steps of the robot's multi-finger grasping method, such as: Acquire visual information of the object grasping task scenario and the robot's proprioceptive information; The visual information and the ontological perception information are fused to obtain cross-modal complementary information; The cross-modal complementary information is input into the strategy generation model to obtain the grasping action output by the strategy generation model; Based on the grasping action, control the multi-finger hand to grasp objects; The policy generation model is trained based on historical cross-modal complementary information, with the goal of minimizing the loss function, on the basis of the policy network. The loss function is generated based on the target reward function, which is the reward function after fusing the force closure evaluation when the multi-finger hand grasps the object.
[0076] Furthermore, the logical instructions in the aforementioned memory 1030 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0077] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the steps of the robot multi-finger grasping method provided in the above embodiments, such as including: Acquire visual information of the object grasping task scenario and the robot's proprioceptive information; The visual information and the ontological perception information are fused to obtain cross-modal complementary information; The cross-modal complementary information is input into the strategy generation model to obtain the grasping action output by the strategy generation model; Based on the grasping action, control the multi-finger hand to grasp objects; The policy generation model is trained based on historical cross-modal complementary information, with the goal of minimizing the loss function, on the basis of the policy network. The loss function is generated based on the target reward function, which is the reward function after fusing the force closure evaluation when the multi-finger hand grasps the object.
[0078] On the other hand, embodiments of this application also provide a non-transitory computer-readable storage medium storing a computer program thereon, the computer program being used to cause a processor to execute the steps of the robotic multi-finger grasping method provided in the above embodiments, for example including: Acquire visual information of the object grasping task scenario and the robot's proprioceptive information; The visual information and the ontological perception information are fused to obtain cross-modal complementary information; The cross-modal complementary information is input into the strategy generation model to obtain the grasping action output by the strategy generation model; Based on the grasping action, control the multi-finger hand to grasp objects; The policy generation model is trained based on historical cross-modal complementary information, with the goal of minimizing the loss function, on the basis of the policy network. The loss function is generated based on the target reward function, which is the reward function after fusing the force closure evaluation when the multi-finger hand grasps the object.
[0079] The non-transitory computer-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memory (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical memory (e.g., CD, DVD, BD, HVD), and semiconductor memory (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).
[0080] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0081] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for grasping with a multi-fingered hand in a robot, characterized in that, include: Acquire visual information of the object grasping task scenario and the robot's proprioceptive information; The visual information and the ontological perception information are fused to obtain cross-modal complementary information; The cross-modal complementary information is input into the strategy generation model to obtain the grasping action output by the strategy generation model; Based on the grasping action, control the multi-finger hand to grasp objects; The policy generation model is trained based on historical cross-modal complementary information, with the goal of minimizing the loss function, on the basis of the policy network. The loss function is generated based on the target reward function, which is the reward function after fusing the force closure evaluation when the multi-finger hand grasps the object.
2. The robot multi-finger grasping method according to claim 1, characterized in that, The strategy generation model is obtained based on the following steps: Historical cross-modal complementary information is input into the policy network to obtain the historical crawling actions output by the policy network; Based on the historical grasping actions, the multi-finger hand is controlled to grasp objects; The contact points between the multi-fingered hand and the object being grasped are clustered to obtain multiple clusters; Based on the center points of the multiple clusters and their corresponding average force vectors, the grasping stability score is calculated. Based on the grasping stability score, calculate the target reward value for the historical grasping action; Based on the target reward value, calculate the loss value of the policy network; If the loss value converges, then the policy network at this point is determined as the policy generation model; If the loss value does not converge, the parameters of the policy network are adjusted, and the process returns to the step of inputting historical cross-modal complementary information into the policy network to obtain the historical grabbing actions output by the policy network, until the loss value converges. At this point, the policy network is determined as the policy generation model.
3. The robot multi-finger grasping method according to claim 2, characterized in that, The historical grasping actions include historical hand end pose and historical finger grasping features. Controlling the multi-finger hand to grasp objects based on the historical grasping actions includes: Based on the main grasping feature vector, the historical finger grasping features are convexly combined to obtain the historical target finger joint position; The historical target finger joint positions are subjected to an exponential moving average to obtain the final historical finger joint positions. The historical hand end pose is input into the robotic arm controller to generate robotic arm control commands. The historical final finger joint position is input into the multi-finger hand controller to generate multi-finger hand control commands; The robot's robotic arm is controlled to move according to the robotic arm control command until the palm end pose of the multi-fingered hand matches the historical palm end pose. The robot's multi-finger hand movement is controlled based on the multi-finger hand control commands until the finger joint positions of the multi-finger hand match the historical final finger joint positions.
4. The robot multi-finger grasping method according to claim 3, characterized in that, The main feature vectors are obtained based on the following steps: Select target objects that have typical set characteristics; Based on the target object, generate multiple grasping postures of the multi-fingered hand; From the various grasping postures, select multiple target postures that conform to human grasping habits; Principal component dimensionality reduction is performed on the various target poses to obtain the main grasping feature vectors corresponding to the main target poses.
5. The robot multi-finger grasping method according to claim 2, characterized in that, The step of calculating the loss value of the policy network based on the target reward value includes: Acquire historical tactile information of the multi-finger hand and historical target information of the grasped object, corresponding to the historical cross-modal complementary information; the historical target information includes historical state information and historical expected target. The historical cross-modal complementary information, the historical tactile information, and the historical target information are input into the evaluation network to obtain the evaluation value of the evaluation network for the historical grasping action; The evaluation value and the target reward value are input into the advantage function to obtain the advantage value of the historical grabbing action output by the advantage function; The advantage value is input into the loss function to obtain the loss value.
6. The robotic multi-finger grasping method according to claim 1, characterized in that, The visual information of the object grasping task scene is obtained based on the following steps: Obtain a depth image of the object grasping task scene; Multiple image points are obtained from the depth image; Randomly select one image point from the plurality of image points as the starting point to construct a sampling point set; Calculate the distances between the plurality of image points and the starting point to obtain a plurality of first distances; Select the maximum distance among the plurality of first distances, determine the image point corresponding to the maximum distance as the new starting point, and add it to the sampling point set; Calculate the distances between the plurality of image points and the new starting point to obtain a plurality of second distances; For each image point, retain the smaller of the first and second distances to obtain multiple third distances; Select the maximum distance among the plurality of third distances, determine the image point corresponding to the maximum distance as the new starting point, and add it to the sampling point set; After taking the multiple third distances as new multiple first distances, return to the step of calculating the distance between the multiple image points and the new starting point to obtain multiple second distances, until the number of image points in the sampling point set reaches a preset number, and the target image point is obtained; The target image points are input into a point cloud network to obtain the visual information.
7. The robot multi-finger grasping method according to claim 1, characterized in that, The robot's body perception information includes the robot's robotic arm information and multi-fingered hand information.
8. A robotic multi-finger grasping device, characterized in that, include: The information acquisition module is used to acquire visual information of the object grasping task scenario and the robot's body perception information. The information fusion module is used to: fuse the visual information and the ontology perception information to obtain cross-modal complementary information; The action generation module is used to: input the cross-modal complementary information into the strategy generation model to obtain the grasping action output by the strategy generation model; The object grasping module is used to control a multi-finger hand to grasp objects based on the grasping action. The policy generation model is trained based on historical cross-modal complementary information, with the goal of minimizing the loss function, on the basis of the policy network. The loss function is generated based on the target reward function, which is the reward function after fusing the force closure evaluation when the multi-finger hand grasps the object.
9. An electronic device comprising a processor and a memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the robot multi-finger grasping method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the robotic multi-finger grasping method according to any one of claims 1 to 7.