A Two-Stage Object Grasping Method and System Based on Visual-Tactile Feature Fusion

Through the two-stage grasping method of visual and tactile features fusion, the visual and touch fusion network and attention mechanism residual convolution network are used, combined with the depth camera and agile hand robotic arm, the problem of insufficient fusion of visual and tactile features is solved, and high-precision object position estimation and stable grasping are achieved.

CN120134325BActive Publication Date: 2025-07-25FOSHAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510611273.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-07-25
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

In the prior art, in the process of robot clever hand grabbing, visual and tactile features are not fully integrated, resulting in low pose estimation accuracy, poor slip detection effect, and insufficient grasp detection.

Method used

The residual convolutional network model based on the visual and touch fusion network model and attention mechanism are adopted, combining visual and tactile features to perform dual-stage grasping, including visual and touch feature fusion processing, pose estimation and slip detection, object information is obtained through depth cameras and dexterous hand robotic arms, feature extraction and pose estimation are used for feature extraction and pose estimation, and tactile sensors are used for stable grasping.

Benefits of technology

The accuracy and stability of object grasping is improved. Through the full fusion of visual and tactile characteristics, detailed prediction and stable grasp of object position are achieved, and the success rate and stability of robot grasping is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120134325B_ABST
    Figure CN120134325B_ABST
Patent Text Reader

Abstract

The present invention discloses a two-stage object grasping method and system based on visual-tactile feature fusion. The method includes: performing visual-tactile feature fusion processing on the object to be grasped based on a visual-tactile fusion network model to obtain a visual-tactile fusion feature vector of the object; performing pose estimation on the object to be grasped based on an attention mechanism residual convolution network model to obtain a grasping rotation angle and a grasping quality score; performing pre-grasping on the object to be grasped based on the grasping rotation angle and the grasping quality score to obtain physical parameter information of the object to be grasped; and combining the visual-tactile fusion feature vector of the object and the physical parameter information of the object to perform slip detection on the object to be grasped. The present invention can fully fuse the visual features and tactile features of the object to be grasped and effectively predict the pose information of the grasped object, improving the accuracy of object grasping. As a two-stage object grasping method and system based on visual-tactile feature fusion, the present invention can be widely applied to the field of machine vision technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine vision, and particularly to a two-stage object grasping method and system based on visual-tactile feature fusion. Background Art

[0002] With the rapid progress of artificial intelligence technology, robots have begun to play an increasingly important role in many fields, and the demand for the ability of robots to perform complex operations has also increased accordingly. Among these operations, the grasping function of a dexterous hand has become a research hotspot, and a certain degree of refined operation can be achieved through the dexterous hand. In order to improve the grasping efficiency of the dexterous hand, it is crucial to accurately identify objects and adjust the finger positions in real time.

[0003] For objects with various shapes, it is often difficult to achieve the required accuracy by relying solely on the dexterous hand to complete the grasping at one time. This requires the dexterous hand to first make micro-contact with the object to obtain its approximate shape and determine a suitable initial grasping posture. In this process, not only the slippage of the object during the grasping process needs to be considered, but also the contact state of each finger with the object needs to be considered. In recent years, researchers have applied multi-modal fusion technology and cooperative control strategies to the grasping detection method of robots. They use a visual-tactile feature fusion network to detect the slippage during the robot grasping process, and adopt a multi-finger cooperative control method based on torque perception to fine-tune the position of each finger to ensure stable grasping, and good results have been achieved. However, most of the related technologies are aimed at performing operations such as slippage detection after direct grasping, without considering the preliminary planning path of the manipulator and the initial state suitable for grasping the object. Secondly, the extracted visual and tactile features cannot be fully utilized, and there is a problem of insufficient fusion, resulting in poor slippage detection effect and low pose estimation accuracy, leading to inaccurate grasping detection. Summary of the Invention

[0004] In order to solve the above technical problems, the purpose of the present invention is to provide a two-stage object grasping method and system based on visual-tactile feature fusion, which can fully fuse the visual features and tactile features of the object to be grasped and effectively predict the detailed pose information of the grasped object to perform two-stage grasping on the object, improving the accuracy of object grasping.

[0005] The first technical solution adopted by the present invention is: a two-stage object grasping method based on visual-tactile feature fusion, comprising the following steps:

[0006] Performing visual-tactile feature fusion processing on the object to be grasped based on a visual-tactile fusion network model to obtain a visual-tactile fusion feature vector of the object to be grasped;

[0007] Performing pose estimation on the object to be grasped based on an attention mechanism residual convolution network model to obtain a grasping rotation angle and a grasping quality score;

[0008] Pre - grasp the object to be grasped based on the grasping rotation angle and the grasping quality score to obtain the physical parameter information of the object to be grasped;

[0009] Combine the visual - tactile fusion feature vector of the object to be grasped and the physical parameter information of the object to be grasped to perform slip detection on the object to be grasped, and achieve stable grasping of the object to be grasped.

[0010] Furthermore, the step of performing visual - tactile feature fusion processing on the object to be grasped based on the visual - tactile fusion network model to obtain the visual - tactile fusion feature vector of the object to be grasped specifically includes:

[0011] Obtain the depth image data of the object to be grasped through a depth camera and perform data pre - processing to obtain the pre - processed depth image data;

[0012] Obtain the tactile array information of the object to be grasped through a dexterous robotic arm with a tactile sensor;

[0013] Construct a visual - tactile fusion network model;

[0014] Based on the visual - tactile fusion network model, perform visual - tactile feature fusion processing on the pre - processed depth image data and the tactile array information of the object to be grasped to obtain the visual - tactile fusion feature vector of the object to be grasped.

[0015] Furthermore, the visual - tactile fusion network model specifically includes a convolutional neural network module, an attention mechanism module, a long - short - term memory network module, and a fully - connected layer, and the convolutional neural network module, the attention mechanism module, the long - short - term memory network module, and the fully - connected layer are connected in sequence.

[0016] Furthermore, the step of performing visual - tactile feature fusion processing on the pre - processed depth image data and the tactile array information of the object to be grasped based on the visual - tactile fusion network model to obtain the visual - tactile fusion feature vector of the object to be grasped specifically includes:

[0017] Input the pre - processed depth image data and the tactile array information of the object to be grasped into the visual - tactile fusion network model;

[0018] Based on the convolutional neural network module of the visual - tactile fusion network model, perform feature extraction processing on the pre - processed depth image data and the tactile array information of the object to be grasped respectively to obtain the local features of the depth image data and the local features of the tactile array information;

[0019] Based on the attention mechanism module of the visual - tactile fusion network model, perform global average pooling, compression, convolution, and adaptive calculation processing on the local features of the depth image data and the local features of the tactile array information in sequence to obtain a visual - tactile feature map with channel attention;

[0020] The long - short - term memory network module based on the visual - tactile fusion network model performs temporal transformation processing on the visual - tactile feature map with channel attention to obtain the transformed visual - tactile feature map;

[0021] Based on the fully - connected layer of the visual - tactile fusion network model, the transformed visual - tactile feature map is classified and output to obtain the visual - tactile fusion feature vector of the object to be grasped.

[0022] Furthermore, the step of performing pose estimation on the object to be grasped by the residual convolutional network model based on the attention mechanism to obtain the grasping rotation angle and the grasping quality score specifically includes:

[0023] Obtain the depth image data of the object to be grasped through a depth camera and perform data pre - processing to obtain the pre - processed depth image data;

[0024] Construct a residual convolutional network model based on the attention mechanism;

[0025] Based on the residual convolutional network model based on the attention mechanism, perform pose estimation on the pre - processed depth image data to obtain the grasping rotation angle and the grasping quality score.

[0026] Furthermore, the residual convolutional network model based on the attention mechanism specifically includes a convolutional module, a residual module, an ESIP attention mechanism module, and a transposed convolutional module. The convolutional module, the residual module, the ESIP attention mechanism module, and the transposed convolutional module are connected in sequence. The ESIP attention mechanism module is formed by connecting the EMA attention mechanism module and the SE attention mechanism module in parallel.

[0027] Furthermore, the step of performing pose estimation on the pre - processed depth image data by the residual convolutional network model based on the attention mechanism to obtain the grasping rotation angle and the grasping quality score specifically includes:

[0028] Input the pre - processed depth image data into the residual convolutional network model based on the attention mechanism;

[0029] Based on the convolutional module of the residual convolutional network model based on the attention mechanism, perform a convolutional calculation operation on the pre - processed depth image data to obtain the local spatial features of the object to be grasped;

[0030] Based on the residual module of the residual convolutional network model based on the attention mechanism, perform a residual operation on the local spatial features of the object to be grasped to obtain the enhanced local spatial features of the object to be grasped;

[0031] The ESIP attention mechanism module of the attention mechanism residual convolution network model performs multi-scale feature extraction and fusion processing on the enhanced local spatial features of the object to be grasped, and obtains the feature map of the object to be grasped after weighted combination;

[0032] The transposed convolution module of the attention mechanism residual convolution network model performs upsampling operation on the feature map of the object to be grasped after weighted combination, and obtains the grasping rotation angle and the grasping quality score.

[0033] Further, the step of performing slip detection on the object to be grasped by combining the visual-tactile fusion feature vector of the object to be grasped and the physical parameter information of the object to be grasped, and realizing stable grasping of the object to be grasped specifically includes:

[0034] Performing secondary grasping on the object to be grasped based on the visual-tactile fusion feature vector of the object to be grasped and the physical parameter information of the object to be grasped;

[0035] If the object to be grasped does not show a slipping phenomenon, mark the force value of the dexterous robotic arm at this time as the first force value;

[0036] If the object to be grasped shows a suspected slipping phenomenon, mark the force value of the dexterous robotic arm at this time as the second force value;

[0037] If the object to be grasped shows a slipping phenomenon, mark the force value of the dexterous robotic arm at this time as the third force value;

[0038] Perform slip detection on the object to be grasped according to the force value of the dexterous robotic arm;

[0039] If the dexterous robotic arm does not touch the object to be grasped, control the dexterous hand to slowly approach the object to be grasped until the force value of the dexterous robotic arm is the first force value;

[0040] Grasp the object to be grasped;

[0041] If the object to be grasped shows a suspected slipping phenomenon, control the force of the dexterous hand so that the force value of the dexterous robotic arm is the second force value;

[0042] If the object to be grasped shows a slipping phenomenon, control the force of the dexterous hand so that the force value of the dexterous robotic arm is the third force value, and realize stable grasping of the object to be grasped.

[0043] The second technical solution adopted by the present invention is: a dual-stage object grasping system based on visual-tactile feature fusion, including:

[0044] The first module is used to perform visual-tactile feature fusion processing on the object to be grasped based on the visual-tactile fusion network model, and obtain the visual-tactile fusion feature vector of the object to be grasped;

[0045] The second module is used to estimate the pose of the object to be grasped based on the attention mechanism residual convolutional network model, and obtain the grasping rotation angle and the grasping quality score;

[0046] The third module is used to perform pre-grasping on the object to be grasped based on the grasping rotation angle and the grasping quality score, and obtain the physical parameter information of the object to be grasped;

[0047] The fourth module is used to combine the visual-tactile fusion feature vector of the object to be grasped and the physical parameter information of the object to be grasped, perform slip detection on the object to be grasped, and achieve stable grasping of the object to be grasped.

[0048] The beneficial effects of the method and system of the present invention are as follows: By performing visual-tactile feature fusion processing on the object to be grasped based on the visual-tactile fusion network model, the visual-tactile fusion feature vector of the object to be grasped is obtained, which can fully fuse the visual features and tactile features of the object to be grasped. Then, based on the attention mechanism residual convolutional network model, the pose of the object to be grasped is estimated, and the grasping rotation angle and the grasping quality score are obtained, which can effectively predict the detailed pose information of the grasped object. Furthermore, based on the grasping rotation angle and the grasping quality score, pre-grasping is performed on the object to be grasped, and then combined with the visual-tactile fusion feature vector of the object to be grasped and the physical parameter information of the object to be grasped, slip detection is performed on the object to be grasped, combined with the physical characteristics of the object obtained in the first-stage pre-grasping, and finally stable grasping in the second stage is performed, thereby improving the accuracy of object grasping. Description of the Drawings

[0049] Figure 1 is the flowchart of the steps of a method for two-stage object grasping based on visual-tactile feature fusion according to the present invention;

[0050] Figure 2 is the structural block diagram of a system for two-stage object grasping based on visual-tactile feature fusion according to the present invention;

[0051] Figure 3 is the schematic framework diagram of two-stage object grasping provided by a specific embodiment of the present invention;

[0052] Figure 4 is the schematic structural diagram of the visual-tactile fusion network model provided by a specific embodiment of the present invention;

[0053] Figure 5 is the schematic structural diagram of the attention mechanism residual convolutional network model provided by a specific embodiment of the present invention. Detailed Embodiments

[0054] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of description and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0055] Referring to Figure 1 and Figure 3 , the present invention provides a two-stage object grasping method based on visual-tactile feature fusion, and the method includes the following steps:

[0056] S100. Perform visual-tactile feature fusion processing on the object to be grasped based on the visual-tactile fusion network model to obtain the visual-tactile fusion feature vector of the object to be grasped;

[0057] S110. Obtain the depth image data of the object to be grasped through a depth camera and perform data preprocessing to obtain the preprocessed depth image data;

[0058] In this embodiment, the present invention uses a D435i depth camera to collect RGB-D image information of the target object, and then uses a Convolutional Neural Networks (CNN) to extract image features. The processed image data is trained through the CNN convolutional neural network so that it can extract the features of the object position in the image. In order to improve the image quality and enhance the robustness of the image processing system, before the image is input into the CNN, it is necessary to preprocess the image data, including adjusting the image size, normalizing the pixel values, data augmentation, etc.

[0059] S120. Obtain the tactile array information of the object to be grasped through a dexterous robotic arm with tactile sensors; construct a visual-tactile fusion network model;

[0060] First of all, it should be noted that compared with visual information, tactile information, as another key sensor information utilized in the robot system, is not only a supplement to vision, but also reflects the interaction between the robot and the environment, and also depicts and describes the physical properties and spatial information of the contacted object. Tactile perception can be used to perceive the subtle changes in the contact force during the operation process, plan repeated grasping, and adjust the current grasping configuration to generate a more stable grasping posture. As a supplement to visual perception, tactile perception has been widely applied in the field of robot grasping and has shown excellent effects in improving the perception of the manipulator for the target object and evaluating the grasping stability.

[0061] In this embodiment, the Freedom Dexterous Hand is used in the embodiment of the present invention, and the model of the tactile sensor equipped is TS161010. This sensor uses CMOS technology to implement a composite sensing structure, thus constituting a single-piece tactile sensing unit (i.e., the contact point), and adopts flexible technology to realize the connection between sensing units and signal reading. When the sensor is subjected to pressure, the corresponding resistance value will decrease with the increase of pressure. Therefore, its piezoresistive characteristic shows that the resistance and pressure are in a power function relationship, while the reciprocal of the resistance and pressure will show an approximate linear relationship. Each sensing unit can be similar to a pressure variable resistor. The upper and lower electrode designs of this fingertip tactile sensor are in the form of "five horizontal and five vertical", automatically dividing the whole piece of piezoresistive material into a 5×5 matrix form, that is, having 25 independent array points, which correspondingly have 25 piezoresistive units. When the dexterous hand performs a grasping operation, the tactile sensors installed on each finger will provide real-time pressure value feedback. Once the tactile sensor senses the pressure information, the pressure values of each contact point will be displayed in real time in the 25 array display frames.

[0062] In addition, touch can also provide the physical properties and more accurate geometric information of the target object. By acquiring and analyzing this information, such as hardness properties, geometric shapes, etc., and combining the pose information obtained previously with it, the pre-grasping in the first stage can be realized, and then a stable grasping operation can be carried out according to the corresponding control strategy, and finally stable grasping is achieved.

[0063] S130. Perform visual-tactile feature fusion processing on the preprocessed depth image data and the tactile array information of the object to be grasped based on the visual-tactile fusion network model to obtain the visual-tactile fusion feature vector of the object to be grasped.

[0064] Among them, the visual-tactile fusion network model specifically includes a convolutional neural network module, an attention mechanism module, a long short-term memory network module and a fully connected layer, and the convolutional neural network module, the attention mechanism module, the long short-term memory network module and the fully connected layer are connected in sequence.

[0065] In this embodiment, the CNN-LSTM network is a deep learning model that combines a convolutional neural network (CNN) and a long short-term memory network (LSTM). CNN is suitable for extracting the spatial features of data and reducing the data dimension, while LSTM is good at obtaining the time features of data, has long-term memory function, and is suitable for processing time series. In the application of visual-tactile fusion, LSTM can process time series data from tactile sensors, such as information about force, pressure, texture, etc. In this way, LSTM can understand and predict the dynamic characteristics of objects. For example Figure 4As shown in the figure, an embodiment of the present invention proposes a CNN-ECA-LSTM visual-tactile fusion network. This network mainly extracts features from visual images and tactile arrays through a convolutional neural network (CNN), then further accurately extracts important features through an attention mechanism, and finally passes through an LSTM network to achieve slip detection.

[0066] Furthermore, the preprocessed depth image data and the tactile array information of the object to be grasped are input into the visual-tactile fusion network model; based on the convolutional neural network module of the visual-tactile fusion network model, feature extraction processing is respectively performed on the preprocessed depth image data and the tactile array information of the object to be grasped to obtain the local features of the depth image data and the local features of the tactile array information; based on the attention mechanism module of the visual-tactile fusion network model, global average pooling, compression, convolution, and adaptive calculation processing are sequentially performed on the local features of the depth image data and the local features of the tactile array information to obtain a visual-tactile feature map with channel attention; based on the long short-term memory network module of the visual-tactile fusion network model. Temporal transformation processing is performed on the visual-tactile feature map with channel attention to obtain the transformed visual-tactile feature map; based on the fully connected layer of the visual-tactile fusion network model, classification output is performed on the transformed visual-tactile feature map to obtain the visual-tactile fusion feature vector of the object to be grasped.

[0067] It should be noted that during the robot grasping process, simply relying on the visual system to guide the dexterous hand to grasp an object often results in less than ideal effects and is prone to dropping the object. Therefore, it is necessary to introduce the tactile perception system on the dexterous hand to fuse the corresponding visual and tactile features of the object, achieve slip detection, ensure that the object will not slip during grasping, and thus achieve stable grasping.

[0068] In this embodiment, the convolutional neural network module is described as follows:

[0069] First, the visual feature and the tactile feature are respectively used as the inputs of the CNN module, and convolutional calculation processing is performed on the image feature and the tactile array to obtain the corresponding visual and tactile local feature maps;

[0070] Then, the obtained feature maps are transmitted to the SE squeeze-and-excitation network attention module (Squeeze-and-Excitation Networks, SE), and the feature maps are squeezed, that is, the output features of the convolutional layer are compressed into a feature vector through global average pooling operation, and a channel weight vector is generated through excitation operation, and then reweighted to obtain important features, capturing the object features of the image and the dynamic pressure changes on the tactile sensor;

[0071] Finally, through the average pooling operation, the feature map is dimensionally reduced by calculating the average value of the elements within each pooling window, while important information in the feature map is retained.

[0072] Further elaboration on the attention mechanism module:

[0073] The features extracted above are transmitted to the efficient frequency attention mechanism module ECA (Efficient Channel Attention for Deep Convolutional Neural Networks, ECA). First, the feature information is compressed again through the global average pooling (GAP) operation by the ECA attention module, retaining the channel dimension;

[0074] Then, the feature is subjected to a 1×1 convolution operation through an adaptively determined one-dimensional convolution kernel, where the kernel size k is adaptively calculated according to the number of channels C of the input feature, and its calculation formula is as follows:

[0075] ;

[0076] After the 1×1 convolution, the Sigmoid function is used to obtain the attention weight for each channel;

[0077] Finally, the learned channel attention weights are multiplied with the features obtained from the CNN module channel by channel to obtain a feature map with channel attention, and this feature map is a one-dimensional sequence.

[0078] Even further elaboration on the long short-term memory network module and the fully connected layer:

[0079] The sequence obtained above is input into the LSTM network module, and the resulting one-dimensional sequence is processed through time steps. Each time step corresponds to a set of features in the sequence, and its internal state is updated simultaneously, finally extracting and transforming the features;

[0080] Then, the fully connected layer is used to splice the data processed from the image features and the data processed from the tactile array to obtain a one-dimensional feature vector that fuses visual and tactile information. Finally, the obtained visual-tactile fusion feature vector is input into the fully connected layer for classification output, and its output will be used for the next slip detection.

[0081] S200, Pose estimation of the object to be grasped is performed based on the attention mechanism residual convolutional network model to obtain the grasping rotation angle and the grasping quality score;

[0082] First of all, it should be noted that pose estimation involves determining the position and orientation of an object in space, which plays an important role in applications such as robot navigation, autonomous driving, and medical image analysis. As Figure 5 shown, an embodiment of the present invention proposes a network called ESIPNet (Generative Residual Convolutional Neural Network with EMA and SE in parallel, ESIPNet) for pose estimation of target objects, which is used to generate the rotation angle and grasping quality of the robotic arm.

[0083] Specifically, depth image data of the object to be grasped is obtained through a depth camera and preprocessed to obtain preprocessed depth image data; an attention mechanism residual convolutional network model is constructed; based on the attention mechanism residual convolutional network model, pose estimation is performed on the preprocessed depth image data to obtain the grasping rotation angle and the grasping quality score.

[0084] Among them, the attention mechanism residual convolutional network model specifically includes a convolutional module, a residual module, an ESIP attention mechanism module, and a transposed convolutional module. The convolutional module, the residual module, the ESIP attention mechanism module, and the transposed convolutional module are connected in sequence. The ESIP attention mechanism module is formed by the parallel connection of an EMA attention mechanism module and an SE attention mechanism module.

[0085] Furthermore, the preprocessed depth image data is input into the attention mechanism residual convolutional network model; based on the convolutional module of the attention mechanism residual convolutional network model, convolutional calculation operations are performed on the preprocessed depth image data to obtain the local spatial features of the object to be grasped; based on the residual module of the attention mechanism residual convolutional network model, residual operations are performed on the local spatial features of the object to be grasped to obtain enhanced local spatial features of the object to be grasped; based on the ESIP attention mechanism module of the attention mechanism residual convolutional network model, multi-scale feature extraction and fusion processing are performed on the enhanced local spatial features of the object to be grasped to obtain a weighted combined feature map of the object to be grasped; based on the transposed convolutional module of the attention mechanism residual convolutional network model, upsampling operations are performed on the weighted combined feature map of the object to be grasped to obtain the grasping rotation angle and the grasping quality score.

[0086] In this embodiment, the convolutional module is first described:

[0087] Taking an n-channel image as the input, after passing through 3 convolutional layers, which mainly perform convolutional calculation operations on the input image to extract local spatial features, and then using normalization and introducing the ReLU function to obtain a series of feature maps.

[0088] Further elaboration on the residual module:

[0089] The feature map obtained from the previous layer is transmitted into the residual layer. After five-layer residual layer operations, the effect of feature enhancement is obtained.

[0090] Further elaboration on the ESIP attention mechanism module:

[0091] For the ESIP attention mechanism module (EMA and SE in parallel, ESIP), that is, the EMA attention mechanism module and the SE attention mechanism module are connected in parallel, which can capture the multi-scale features of image information and the relationships between channels. Among them, for the EMA module, short-term and long-term dependencies are established for the upper-layer features through parallel sub-networks, while the SE module learns the relationships between each feature channel through GAP. Finally, their outputs are fused through multiplication operations to enhance the feature representation and obtain the weighted combined feature map. In this network, this module is mainly interspersed after the convolutional module and the residual module, which can extract corresponding important features, thereby improving the model performance of pose estimation.

[0092] Even further elaboration on the transposed convolution module:

[0093] Then, the features extracted from the previous layer go through three transposed convolution layers, mainly performing upsampling operations to increase the spatial dimension of the feature map. Finally, two images are generated, namely the grasping quality score Quantity and the rotation angle θ during grasping.

[0094] S300. Pre-grasp the object to be grasped based on the grasping rotation angle and the grasping quality score to obtain the physical parameter information of the object to be grasped;

[0095] It should be noted that after obtaining the pose information through the pose estimation network, the dexterous hand can accurately judge the position of the object, and then the dexterous hand can start to grasp the object. The embodiment of the present invention divides the process of grasping an object into two stages. The first stage is pre-grasping, mainly enabling the dexterous hand to have a general grasp of the object first, so as to judge the corresponding grasping state and the general shape of the object. The second stage is stable grasping. After determining that the dexterous hand has slightly contacted the object, according to the information feedback by slip detection, a multi-finger cooperative control strategy is adopted to adjust the joints of each finger, so as to stably grasp the object.

[0096] In the pre-grasping stage, the dexterous hand needs to generate and select a suitable grasping posture. The main tasks of this stage are to determine the approaching path of the dexterous hand, the grasping point, and the initial closing position of the fingers. The pose information generated by the above ESIPNet network can determine the object grasping point, and the tactile sensors on the fingertips of the dexterous fingers can identify the physical information of the hardness property and geometric shape of the unknown object. Through the combination of pose information and physical characteristics for grasping planning, the pre-grasping in the first stage is thus achieved.

[0097] S400. Combine the visual-tactile fusion feature vector of the object to be grasped and the physical parameter information of the object to be grasped, perform slip detection on the object to be grasped, and achieve stable grasping of the object to be grasped.

[0098] Specifically, based on the visual-tactile fusion feature vector of the object to be grasped and the physical parameter information of the object to be grasped, perform secondary grasping on the object to be grasped; if there is no slipping phenomenon of the object to be grasped, mark the force value of the dexterous hand robot arm at this time as the first force value; if there is a suspected slipping phenomenon of the object to be grasped, mark the force value of the dexterous hand robot arm at this time as the second force value; if there is a slipping phenomenon of the object to be grasped, mark the force value of the dexterous hand robot arm at this time as the third force value; perform slip detection on the object to be grasped according to the force value of the dexterous hand robot arm; if the dexterous hand robot arm does not touch the object to be grasped, control the dexterous hand to slowly approach the object to be grasped until the force value of the dexterous hand robot arm is the first force value; grasp the object to be grasped; if there is a suspected slipping phenomenon of the object to be grasped, control the force of the dexterous hand so that the force value of the dexterous hand robot arm is the second force value; if there is a slipping phenomenon of the object to be grasped, control the force of the dexterous hand so that the force value of the dexterous hand robot arm is the third force value, and achieve stable grasping of the object to be grasped.

[0099] In this embodiment, in the stable grasping stage, the dexterous hand will experience a change from micro-contact to dynamic contact until stable grasping. After completing the pre-grasping, it is necessary to monitor the state of the object in real time, especially when the object slides. Robot grasping slip detection is of great significance for the grasping task. During the grasping process, vision and touch, as the key modal information for judging the grasping state, it is still challenging to achieve their efficient fusion. Since the dexterous hand only obtains the approximate grasping posture and force of the object in the pre-grasping stage, for heavier or stiffer objects, the object is extremely likely to slip. Therefore, it is necessary to utilize the signal feedback from slip detection and adjust the joint postures and forces through a multi-finger coordination control strategy.

[0100] Tactile sensors can effectively detect contact slippage between a dexterous hand and the grasped object, estimate the contact force, and locate the object. After the object is pre-grasped, the sensing units on the tactile sensor array will have corresponding pressure information. During the object grasping process, the information of the tactile sensor changes in real time. We can judge whether the object will slip according to the change of the pressure information on the tactile sensing unit and make corresponding grasping adjustments, which can be divided into three cases:

[0101] 1) If the picture information displayed by the sensing unit on the tactile sensing array is constant, it means that the grasped object has no slipping phenomenon and is marked as 0. At this time, the force-bearing situation of the dexterous hand's grasp is recorded as ;

[0102] 2) If the picture information displayed by the sensing unit on the tactile sensing array gradually becomes smaller but does not completely disappear, it means that the grasped object has a slight slip and is marked as 1. Therefore, it is necessary to finely adjust the joint rotation of the dexterous hand to be able to grasp the object tightly. At this time, the force-bearing situation of the dexterous hand's grasp is recorded as ;

[0103] 3) If the picture information displayed by the sensing unit on the tactile sensing array gradually becomes smaller until it disappears, it means that the grasped object has completely slipped and is marked as 2. This situation indicates that the force at this time is too small, resulting in the object falling off. Therefore, it is necessary to re-grasp the object. Increase the force to grasp. If there is no slip, the force-bearing situation of the dexterous hand's grasp at this time is recorded as . The effect of slippage can be represented by the values on each contact point of the sensing unit on the tactile sensing array.

[0104] After performing the slip detection operation, multi-finger cooperative control is required. First, the expected contact force between the dexterous hand finger and the object has been defined as . When the sensing unit of the tactile sensor on a single finger reaches this value, it means that the finger has touched the object. Therefore, use as the critical value, represents the reading of the sensing unit of the tactile sensor on a single finger. According to the force-bearing situation of the dexterous hand's finger and the slipping situation, it can be divided into three states:

[0105] 1) If the object is not touched, it is necessary to control the finger to slowly approach the object until .

[0106] 2) If the object is touched and the object has a slight slip, it is necessary to control the finger to adjust the force until .

[0107] 3) If the object is touched and the object has completely slipped, it is necessary to re-grasp until And there is no more slipping phenomenon.

[0108] After the multi-finger cooperative control of the dexterous hand to grasp an object, the final stable grasping can be achieved.

[0109] In summary, vision cannot fully perceive the contact surface characteristics of the grasped object, the changes in the object pose, and the changes in the contact force information; the tactile-based method often requires multiple attempts and is difficult to achieve an efficient grasping task. Therefore, when analyzing the grasping stability, the accuracy and stability of grasping can be improved by the mutual complementation of tactile modality information and visual modality information.

[0110] In summary, the embodiment of the present invention proposes a two-stage grasping strategy based on a dexterous hand, and uses a multi-modal fusion method to integrate visual information and tactile information, thereby improving the success rate and stability of robot grasping. First, the RBG image and depth image of the target object are collected by a depth camera. Secondly, the visual feature information extracted by the neural network is input into the Generative Residual Convolutional Neural Network with EMA and SE in parallel (ESIPNet) to obtain the pose information of the object. This network is a parallel connection of an efficient multi-scale attention module EMA and a squeeze-and-excitation network SE attention module added on the basis of the generative residual convolutional network. Then, it is combined with the physical characteristics of the object obtained in the first-stage pre-grasping, and finally the second-stage stable grasping is performed. In addition, the tactile image collected by the TS161010 tactile sensor on the dexterous fingertip will also extract tactile feature information through the neural network. Subsequently, the visual feature information and tactile feature information are input into the CNN-ECA-LSTM visual-tactile fusion network for slip detection of the object, and the second-stage stable grasping is guided according to the feedback of the visual-tactile fusion information.

[0111] Therefore, the embodiment of the present invention can provide an effective way to predict the detailed pose information of the grasped object. Using the feature fusion method can solve the problem of insufficient or mismatched existing visual-tactile fusion feature extraction. For the multi-finger grasping of the dexterous hand, considering the grasping state of each finger according to the slip detection situation of the grasp, and judging the grasping state of a single finger through the corresponding information of the tactile sensing unit.

[0112] Referring to Figure 2 , a two-stage object grasping system based on visual-tactile feature fusion, comprising:

[0113] The first module 201 is used to perform visual-tactile feature fusion processing on the object to be grasped based on the visual-tactile fusion network model to obtain the visual-tactile fusion feature vector of the object to be grasped;

[0114] The second module 202 is used to perform pose estimation on the object to be grasped based on the attention mechanism residual convolution network model, and obtain the grasping rotation angle and the grasping quality score;

[0115] The third module 203 is used to perform pre-grasping on the object to be grasped based on the grasping rotation angle and the grasping quality score, and obtain the physical parameter information of the object to be grasped;

[0116] The fourth module 204 is used to combine the visual-tactile fusion feature vector of the object to be grasped and the physical parameter information of the object to be grasped, perform slip detection on the object to be grasped, and realize stable grasping of the object to be grasped.

[0117] The content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented by the system embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0118] The above has made a specific description of the preferred embodiment of the present invention, but the present invention is not limited to the described embodiment. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A two-stage object grasping method based on the fusion of visual and tactile features, characterized in that Including the following steps: Performing visual-tactile feature fusion processing on the object to be grasped based on a visual-tactile fusion network model to obtain a visual-tactile fusion feature vector of the object to be grasped; Obtaining depth image data of the object to be grasped through a depth camera and performing data preprocessing to obtain preprocessed depth image data; Constructing an attention mechanism residual convolutional network model; The attention mechanism residual convolutional network model specifically includes a convolutional module, a residual module, an ESIP attention mechanism module, and a transposed convolutional module. The convolutional module, the residual module, the ESIP attention mechanism module, and the transposed convolutional module are connected in sequence. The ESIP attention mechanism module is formed by the parallel connection of an EMA attention mechanism module and an SE attention mechanism module; Inputting the preprocessed depth image data into the attention mechanism residual convolutional network model; Based on the convolutional module of the attention mechanism residual convolutional network model, performing a convolutional calculation operation on the preprocessed depth image data to obtain the local spatial features of the object to be grasped; Based on the residual module of the attention mechanism residual convolutional network model, performing a residual operation on the local spatial features of the object to be grasped to obtain enhanced local spatial features of the object to be grasped; Based on the ESIP attention mechanism module of the attention mechanism residual convolutional network model, performing multi-scale feature extraction and fusion processing on the enhanced local spatial features of the object to be grasped to obtain a weighted combined feature map of the object to be grasped; Based on the transposed convolutional module of the attention mechanism residual convolutional network model, performing an upsampling operation on the weighted combined feature map of the object to be grasped to obtain the grasping rotation angle and the grasping quality score; Performing pre-grasping on the object to be grasped based on the grasping rotation angle and the grasping quality score to obtain the physical parameter information of the object to be grasped; Combining the visual-tactile fusion feature vector of the object to be grasped and the physical parameter information of the object to be grasped, performing slip detection on the object to be grasped, and realizing stable grasping of the object to be grasped.

2. The dual-stage object grasping method based on visual-tactile feature fusion according to claim 1, wherein, The step of performing visual-tactile feature fusion processing on the object to be grasped based on a visual-tactile fusion network model to obtain a visual-tactile fusion feature vector of the object to be grasped specifically includes: Obtaining depth image data of the object to be grasped through a depth camera and performing data preprocessing to obtain preprocessed depth image data; Obtaining the tactile array information of the object to be grasped through a dexterous robotic arm with tactile sensors; Constructing a visual-tactile fusion network model; Based on the visual-tactile fusion network model, performing visual-tactile feature fusion processing on the preprocessed depth image data and the tactile array information of the object to be grasped to obtain a visual-tactile fusion feature vector of the object to be grasped.

3. The dual-stage object grasping method based on visual-tactile feature fusion according to claim 2, wherein The visual-tactile fusion network model specifically includes a convolutional neural network module, an attention mechanism module, a long short-term memory network module, and a fully connected layer. The convolutional neural network module, the attention mechanism module, the long short-term memory network module, and the fully connected layer are connected in sequence.

4. The dual-stage object grasping method based on visual-tactile feature fusion according to claim 3, wherein The step of performing visual-tactile feature fusion processing on the preprocessed depth image data and the tactile array information of the object to be grasped based on the visual-tactile fusion network model to obtain a visual-tactile fusion feature vector of the object to be grasped specifically includes: Input the preprocessed depth image data and the tactile array information of the object to be grasped into the vision-tactile fusion network model; Based on the convolutional neural network module of the vision-tactile fusion network model, perform feature extraction processing on the preprocessed depth image data and the tactile array information of the object to be grasped respectively, and obtain the local features of the depth image data and the local features of the tactile array information; Based on the attention mechanism module of the vision-tactile fusion network model, perform global average pooling, compression, convolution and adaptive calculation processing on the local features of the depth image data and the local features of the tactile array information in sequence, and obtain a vision-tactile feature map with channel attention; Based on the long short-term memory network module of the vision-tactile fusion network model, perform temporal conversion processing on the vision-tactile feature map with channel attention, and obtain the converted vision-tactile feature map; Based on the fully connected layer of the vision-tactile fusion network model, perform classification output on the converted vision-tactile feature map, and obtain the vision-tactile fusion feature vector of the object to be grasped.

5. The dual-stage object grasping method based on visual-tactile feature fusion according to claim 4, wherein The step of performing slip detection on the object to be grasped by combining the vision-tactile fusion feature vector of the object to be grasped and the physical parameter information of the object to be grasped, and realizing stable grasping of the object to be grasped specifically includes: Perform secondary grasping on the object to be grasped based on the vision-tactile fusion feature vector of the object to be grasped and the physical parameter information of the object to be grasped; If the object to be grasped does not show a slipping phenomenon, mark the force value of the robotic hand at this time as the first force value; If the object to be grasped shows a suspected slipping phenomenon, mark the force value of the robotic hand at this time as the second force value; If the object to be grasped shows a slipping phenomenon, mark the force value of the robotic hand at this time as the third force value; Perform slip detection on the object to be grasped according to the force value of the robotic hand; If the robotic hand does not touch the object to be grasped, control the robotic hand to slowly approach the object to be grasped until the force value of the robotic hand is the first force value; Grasp the object to be grasped; If the object to be grasped shows a suspected slipping phenomenon, control the force of the robotic hand so that the force value of the robotic hand is the second force value; If the object to be grasped shows a slipping phenomenon, control the force of the robotic hand so that the force value of the robotic hand is the third force value, and realize stable grasping of the object to be grasped.

6. A two-stage object grasping system based on the fusion of visual and tactile features, characterized in that A method for dual-stage object grasping based on vision-tactile feature fusion as described in any one of claims 1-5, comprising the following modules: The first module is used to perform vision-tactile feature fusion processing on the object to be grasped based on the vision-tactile fusion network model, and obtain the vision-tactile fusion feature vector of the object to be grasped; The second module is used to perform pose estimation on the object to be grasped based on the attention mechanism residual convolutional network model, and obtain the grasping rotation angle and the grasping quality score; The third module is used to perform pre-grasping on the object to be grasped based on the grasping rotation angle and the grasping quality score, and obtain the physical parameter information of the object to be grasped; The fourth module is used to perform slip detection on the object to be grasped by combining the vision-tactile fusion feature vector of the object to be grasped and the physical parameter information of the object to be grasped, and realize stable grasping of the object to be grasped.

Citation Information

Patent Citations

  • Electric appliance identification method based on channel attention mechanism residual convolutional network

    CN115795308A

  • Grabbing stability evaluation method based on visual touch fusion perception and multi-modal space-time convolution

    CN116945170A