Transform-CNN (Convolutional Neural Network)-based plane grabbing pose processing method and system
By combining the characteristics of Transformer and CNN, using technical means such as ITC-Grasp network architecture and frequency slope structure, the problems of low grab accuracy and inaccurate pose estimation in complex environments are solved, and higher accuracy and generalization capabilities of grabbing pose prediction are achieved.
Patent Information
- Application Number
- CN202510104796.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-23
AI Technical Summary
The problem of low grasping accuracy and inaccurate pose estimation in complex and changing environments.
Using the plane grab pose processing method based on Transformer-CNN, through the ITC-Grasp network architecture, the local feature extraction capability of CNN is combined with the global information processing capability of Transformer to obtain multi-scale grab feature representations, and a frequency slope structure and ShuffleAttention attention mechanism are introduced to weigh high-level and low-level features.
It significantly improves the accuracy and generalization ability of the robot to predict the grab position on complex scenes and variable objects, enhances feature extraction and representation capabilities, and improves the grab accuracy and efficiency.
Smart Images

Figure CN120031965A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual capture and prediction technology, and more specifically, to a Transformer-CNN-based planar capture posture processing method and system. Background Art
[0002] With the development of science and technology, robots have been widely used in the field of industrial manufacturing and daily life. Human dependence and demand for robots are increasing day by day. Therefore, robots are expected to demonstrate efficient, stable and reliable operation capabilities in more complex and changeable environments to assist in completing various tasks more conveniently. As an important skill of robots, visual grasping has been widely used in operation scenarios such as item sorting and product packaging in fixed places or structured environments, greatly improving production efficiency and automation level.
[0003] However, when the robot is in an unstructured environment, it is still a very challenging and complex technical problem to accurately estimate the pose of the object and achieve precise grasping. Inappropriate grasping position selection can easily lead to operational errors, which in turn affects the complete realization and efficiency of the overall operation process. Traditional grasping prediction methods mostly rely on the shape characteristics and physical properties of the object itself, and their effective implementation is often based on the assumptions of a stable external environment, friction, stiffness, and an accurate three-dimensional object model. However, in actual application scenarios, the diversity of grasping objects and working environments makes it difficult to fully meet these assumptions, resulting in poor performance of traditional methods in actual applications and difficulty in meeting the grasping requirements in complex and changing environments. At present, some scholars have proposed a robot grasping position prediction method based on convolutional neural network (CNN), which extracts local features of the image through CNN, but it is also difficult to capture global dependencies in real scenarios. Therefore, when performing grasping tasks, robots face the problems of low grasping accuracy and inaccurate pose estimation.
[0004] In the related technology, for example, Chinese patent CN117381780A provides a dynamic grasping method and system of a composite robot based on visual servoing, which acquires the depth image of the composite robot in the direction of travel in real time, obtains the coordinate information of the object to be grasped and the current posture of the composite robot based on the depth information of the depth image and the position of the object to be grasped in the image, determines the position of the composite robot in the world coordinate system, estimates the coordinates of the object to be grasped based on the coordinate information of the object to be grasped and the world coordinates of the composite robot, and grasps the object to be grasped. Its disadvantage is that each grasping requires a complex calculation process, which is prone to errors, and the accuracy of the calculation results is difficult to guarantee, and the process is relatively complicated, resulting in low work efficiency.
[0005] From the above, we can see that the relevant technology does not provide any technical inspiration on how to solve the problems of low grasping accuracy and inaccurate posture estimation of robots in complex and changing environments. Summary of the invention
[0006] 1. Technical problems to be solved
[0007] In view of the problem of how to improve the grasping accuracy and efficiency of robots in complex environments in the prior art, the present invention provides a method and system for processing planar grasping posture based on Transformer-CNN, which can be implemented based on the ITC-Grasp network architecture, using the Inception mixer with a channel splitting mechanism to graft the efficient local feature extraction capability of CNN onto the Transformer, obtain better multi-scale grasping feature representation and global information, and introduce a frequency ramp structure to flexibly weigh high-level features and low-level features.
[0008] 2. Technical solution
[0009] The purpose of the present invention is achieved through the following technical solutions.
[0010] The content of this application is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this application is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.
[0011] Some embodiments of the present application propose a method and system for processing planar grasping posture based on Transformer-CNN to solve the technical problems mentioned in the above background technology section.
[0012] As a first aspect of the present application, some embodiments of the present application provide a method for processing a plane grasping posture based on Transformer-CNN, comprising the following steps: preprocessing a prediction data set to obtain an input image; constructing a plane grasping posture prediction model based on Transformer-CNN; the plane grasping posture prediction model includes an encoder module, a decoder module and a grasping prediction module;
[0013] The acquired input image is input into the constructed plane grasping pose prediction model for model training and testing: the trained plane grasping pose prediction model is used to obtain the grasping configuration and generate the prediction result.
[0014] Furthermore, the encoder module has an Inception mixer at each stage; the Inception mixer includes a high-frequency mixer and a low-frequency mixer; the high-frequency mixer is composed of deep convolution and maximum pooling, and the low-frequency mixer is implemented by flat pooling and self-attention.
[0015] Furthermore, assuming that the input of the planar grasping pose prediction model is X, X divides the channel dimension C into C h and C l Channels, the calculation process in the Inception mixer is as follows:
[0016] Y h1 =FC(MaxPool(X h1 ));
[0017] Y h2 =DWConv(FC(X h2 ));
[0018] Y l =Upsample(MSA(Avepooling(X l )));
[0019] Y c =Concat(Y l ,Y h1 ,Y h2 );
[0020] Y h1 represents the output of the high-frequency mixer after maximum pooling, Y h2 represents the output of the high-frequency mixer after depthwise separable convolution, Y l is the output of the low frequency mixer, Y c is the output of the Inception mixer, X h1 To feed Y h1 The number of channels is C h / 2 feature map, X h2 To feed Y h2 The number of channels is C h / 2 feature map, X l The number of channels in the low frequency mixer is C l feature map, MSA represents multi-head self-attention; FC() represents full connection; MaxPool() represents maximum pooling; DWConv() represents depthwise separable convolution; Avepooling() represents average pooling; MSA() represents multi-head self-attention mechanism; Upsample() represents upsampling operation; Concat() represents concatenation.
[0021] Furthermore, the acquired input image is input into the trained planar grasping pose prediction model, and the grasping configuration is obtained through the subtask network of the grasping prediction module, including the quality heat map, width heat map and angle heat map.
[0022] Furthermore, the prediction data set includes a Cornell grasping data set and a Jacquard grasping data set; the process of preprocessing the Cornell grasping data set is specifically as follows: performing data enhancement on all images in turn; cropping the image size to a fixed pixel; and normalizing the cropped image.
[0023] Furthermore, the Cornell grasping dataset is divided into a training set and a test set in a ratio of 9:1 using image segmentation and object segmentation;
[0024] During the model training process, the training parameters are adjusted until the convergence of the plane grasping pose prediction model slows down. The training parameters include the training batch size, the number of training rounds, and the learning rate.
[0025] Furthermore, the ITC-Grasp model is trained using the Cornell grasp prediction dataset, using a five-dimensional grasp pose representation method. The five-dimensional grasp representation g is defined as follows:
[0026] g = {x, y, w, h, θ};
[0027] (x, y) represents the two-dimensional coordinates of the center point of the grasping rectangle, θ represents the rotation angle of the grasping rectangle, and (w, h) represents the size of the grasping rectangle, where w represents the width and h represents the height.
[0028] Furthermore, the indicators for evaluating whether the predicted grasping rectangle is correct include: the offset between the grasping box of the predicted grasping rectangle and the ground truth rectangle is less than 30°; the Jaccard index between the predicted grasping rectangle and the ground truth rectangle is greater than 25%.
[0029] Furthermore, Smooth L1 is used as the loss function in the model training process, which is defined as follows:
[0030]
[0031] represents the total loss, n is the number of samples, Z i is the loss of the i-th sample, where i represents the sample number.
[0032] As the second aspect of the present application, some embodiments of the present application provide a plane grasping pose processing system based on Transformer-CNN, including a data acquisition module: preprocessing the prediction data set to obtain an input image; a model construction module: constructing a plane grasping pose prediction model based on Transformer-CNN; the plane grasping pose prediction model includes an encoder module, a decoder module and a grasping prediction module; a model training module: inputting the acquired input image into the constructed plane grasping pose prediction model for model training and testing: a result prediction module: using the trained plane grasping pose prediction model to obtain the grasping configuration and generate a prediction result.
[0033] 3. Beneficial effects
[0034] Compared with the prior art, the advantages of the present invention are:
[0035] (1) Enhanced core feature extraction and representation capabilities: The ITC-Grasp model of the present invention combines the ability of CNN to extract high-frequency features with the global modeling capability of Transformer. Through the Inception mixer with a channel splitting mechanism in the encoding stage, the processing method for plane grasping pose prediction can more effectively capture the detailed features in the image, and combined with the global vision of Transformer, a more robust and comprehensive grasping feature representation is obtained. This core improvement significantly improves the model's grasping pose prediction capability for complex scenes and variable objects.
[0036] (2) Multi-level feature interaction and information retention: In order to reduce the loss of underlying feature information and effectively model global context information, the technical solution of the present invention introduces a frequency ramp structure, which enables the encoder to effectively weigh high- and low-frequency components between different layers, thereby retaining more key information; in addition, in the decoding upsampling stage, the ShuffleAttention mechanism is used to realize information exchange between features at each level, further enhancing the model's attention to the grasping area and the separability of objects and backgrounds, generating more accurate grasping posture outputs.
[0037] (3) Improved grasping pose prediction accuracy and generalization capability: The technical solution of the present invention not only improves the model’s grasping pose prediction accuracy for specific objects, but also significantly enhances its generalization capability. By combining the global information extraction capability of Transformer and the local information modeling capability of CNN, the model can better adapt to different scenarios and object forms, thereby demonstrating higher stability and reliability in practical applications. In summary, the technical solution of the present invention achieves the enhancement of grasping feature representation, the optimization of multi-level feature interaction and information retention, and the improvement of grasping pose prediction accuracy and generalization capability. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a flowchart of a method for processing plane grasping posture based on Transformer-CNN in one embodiment of the present invention;
[0039] Figure 2 It is a network structure diagram of the ITC-Grasp model of planar grasping posture based on Transformer-CNN in one embodiment of the present invention;
[0040] Figure 3 A network structure diagram of an Inception mixer in a network structure in one embodiment of the present invention;
[0041] Figure 4 A network structure diagram of the Shuffle Attention mechanism in the network structure in one embodiment of the present invention;
[0042] Figure 5 It is a partial prediction graph of the model on the test set of the preprocessed Cornell crawling data set in one embodiment of the present invention;
[0043] Figure 6 This is a comparative prediction diagram of the model in one embodiment of the present invention and GGCNN and GR-ConvNet on the Cornell captured dataset. DETAILED DESCRIPTION
[0044] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0045] Combination Figures 1 to 6 A method for processing plane grasping posture based on Transformer-CNN of the present invention comprises the following steps:
[0046] S1. Preprocess the prediction data set to obtain the input image:
[0047] Specifically, the grasped prediction dataset is preprocessed to generate input images suitable for ITC-Grasp model training.
[0048] The prediction datasets include the Cornell grasping dataset and the Jacquard grasping dataset. The Cornell grasping dataset provides 885 sets of RGB images, depth images, and related point cloud data. Each set of images contains only one common object in life. The images in the dataset are taken in a real environment and annotated with multiple grasping rectangles. The Jacquard grasping dataset contains RGB and depth images of 11,619 different objects in 54,485 different scenes.
[0049] Among them, the Cornell grasping dataset is a standard dataset widely used in grasping prediction research. It contains 885 groups of images, each of which includes an RGB image, a depth image, and related point cloud data. Each group of images only shows one daily object. The images in the Cornell grasping dataset are taken in a real environment and annotated with multiple grasping rectangles (i.e., labeled grasping areas) to indicate possible grasping positions and directions. The Jacquard grasping dataset is another larger and more diverse grasping prediction dataset, which contains RGB and depth images of 11,619 different objects in 54,485 different scenes. These images provide rich grasping scenes and object diversity, which helps to train a more robust grasping ITC-Grasp model.
[0050] Since the size of the Cornell grab dataset is small, in order to prevent overfitting of the ITC-Grasp model, data augmentation is used to increase the size of the dataset, and the image size is cropped as the input of the network model.
[0051] Specifically, the process of preprocessing the Cornell crawl dataset includes: performing data augmentation on all images in the prediction dataset in turn; cropping the image size to a fixed pixel size, and normalizing the cropped image as the input of the network model. The Jacquard crawl dataset itself is large enough, so there is no need to perform additional processing on the Jacquard crawl dataset. By preprocessing the prediction dataset, the generalization ability of the deep learning model is improved.
[0052] In a specific embodiment, data enhancement includes randomly translating and flipping the image horizontally and vertically, randomly rotating the image angle, randomly reducing or enlarging the image size, randomly cropping or splicing the image, and randomly changing the image color and brightness; the image is cropped to a size of 224×224 pixels. By normalizing the cropped image, that is, scaling the pixel value from the range of [0, 255] to the range of [0, 1], gradient stabilization and accelerated convergence can be achieved during the deep learning model training process.
[0053] S2. Constructing the plane grasping posture prediction model—ITC-Grasp model:
[0054] Based on the encoder-decoder network architecture, a network model of ITC-Grasp planar grasping posture based on Transformer-CNN, namely the ITC-Grasp model, is constructed.
[0055] like Figure 2As shown in the figure, it is the specific structure of the ITC-Grasp model. The ITC-Grasp model includes an encoder module, a decoder module and a grasp prediction module, wherein the encoder module is used to extract the feature information of the input image, the decoder module is used to restore the feature map resolution and retain important feature information, and the grasp prediction module is used to generate the grasp configuration (grasp quality, grasp angle and grasp width). Norm is standardization, FFN is a feedforward network, ConvT2D is a transposed convolution, Batch Norm is batch normalization, and Relu is an activation function. The ITC-Grasp model is based on the encoder-decoder network architecture, including an encoder module, a decoder module and the final grasp prediction module. The encoder module is responsible for feature extraction of the input image, the decoder module is responsible for upsampling to restore the feature map resolution to the same size as the input image, and the grasp prediction module is used to generate the grasp configuration required by the robot (including grasp quality, grasp angle and grasp width). Each stage of the encoder module has an Inception mixer, which aims to enhance the perception of visual input by capturing high-frequency and low-frequency features in the data and effectively capture specific frequency information on the corresponding channels, thereby extracting richer features over a wider frequency range. Secondly, the encoding stage can effectively weigh the high-frequency and low-frequency components of each layer through a flexible frequency ramp structure.
[0056] Specifically, the Shuffle Attention mechanism is introduced in each stage of the decoder to realize information exchange between features at each level. While restoring the feature map to the original image size, the ITC-Grasp model can pay more attention to the graspable position, and finally input it into the grasp prediction module for grasping posture prediction.
[0057] S3. Input the input image obtained after preprocessing into the constructed plane grasping pose prediction model for training and testing:
[0058] The Cornell grasp dataset in the input image obtained after preprocessing is divided into a training set and a test set, and a five-fold cross-validation method is used to train and test the performance of the ITC-Grasp model on the Jacquard grasp dataset.
[0059] Specifically, the Cornell crawling dataset was randomly divided into a training set and a test set in a ratio of 9:1 using image segmentation and object segmentation. Image segmentation is used to test the model's prediction performance for previously seen objects in different directions and postures, while object segmentation refers to the instance division of all images in the dataset to ensure that objects in the training set do not appear in the test set, which is used to test the network's generalization ability for unknown objects.
[0060] The ITC-Grasp model uses the training set of the Cornell crawling dataset during training and the test set of the crawling dataset during testing. The training parameters are continuously adjusted during the ITC-Grasp model training process to make the model converge quickly.
[0061] Specifically, adjusting the training parameters includes changing the training batch size, the number of training rounds and the learning rate. More specifically, when the training parameters are increased or decreased, if the convergence of the ITC-Grasp model slows down, then the critical point is reached and the adjustment ends.
[0062] like Figure 3 The figure shows the specific structure of the Inception mixer, which consists of a high-frequency mixer and a low-frequency mixer. The high-frequency mixer consists of deep convolution and maximum pooling, and the low-frequency mixer is implemented by average pooling and self-attention. Linear is a linear layer, MaxPool is maximum pooling, AvePool is average pooling, and DWConv is a deep convolution.
[0063] Assume that the input of the ITC-Grasp model is X, and X is split into C along the channel dimension C of the Inception mixer h and C l channels, and put X h and X l The high frequency mixer uses a parallel structure to learn the high frequency components. First, the input X h Divided into X h1 and X h2 (All C h / 2 channels), X h1 Embedded MaxPool layer and linear layer, X 2 Feed to the linear layer and the deep convolution layer. The low-frequency mixer is the traditional multi-head self-attention. Because the high-frequency mixer brings additional calculation, the average pooling operation is performed here first, and then the upsampling layer restores the spatial dimension to the original state. The calculation process in the Inception mixer is as follows:
[0064] Y h1 =FC(MaxPool(X h1 )) (1)
[0065] Y h2 =DWConv(FC(Y h2 )) (2)
[0066] Y l =Upsample(MSA(Avepooling(Xl ))) (3)
[0067] Y c =Concat(Y l ,Y h1 ,Y h2 ) (4)
[0068] Among them, Y h1 represents the output of the high-frequency mixer after maximum pooling, Y h2 represents the output of the high-frequency mixer after depthwise separable convolution, Y l is the output of the low frequency mixer, Y c is the output of the Inception mixer, X h1 To feed Y h1 The number of channels is C h / 2 feature map, X h2 To feed Y h2 The number of channels is C h / 2 feature map, X l The number of channels in the low frequency mixer is C l feature map, MSA represents multi-head self-attention; FC() represents full connection; MaxPool() represents maximum pooling; DWConv() represents depthwise separable convolution; Avepooling() represents average pooling; MSA() represents multi-head self-attention mechanism; Upsample() represents upsampling operation; Concat() represents concatenation.
[0069] Specifically, the calculation process of an encoding stage is as follows:
[0070] Y=X+ITM(LN(X)) (5)
[0071] H=Y+FFN(LN(Y)) (6)
[0072] Among them, ITM represents the calculation process of an Inception mixer, X is the input, Y is the residual output, H represents the output of an encoding stage, FFN represents the feedforward network, and LN represents layer normalization.
[0073] like Figure 4 As shown in the figure, the specific structure of the Shuffle Attention mechanism is that the input features are grouped into multiple sub-features along the channel dimension. For each group of sub-features, the structure uses the Shuffle Unit to integrate the channel attention and spatial attention into a block of each group, and then aggregates all sub-features to realize information exchange between different sub-features through channel shuffling operations.
[0074] S4. Use the trained ITC-Grasp model to obtain the grasping configuration:
[0075] The image obtained in step S3 is input into the trained ITC-Grasp model, and finally the ITC-Grasp model obtains the grasping configuration through the subtask network of the grasping prediction module.
[0076] Specifically, the grasping configuration includes grasping quality heat map, grasping width heat map and grasping angle heat map. The grasping prediction module is used to generate the grasping configuration. The grasping prediction module includes four subtask networks responsible for grasping quality heat map Q, grasping width heat map W, grasping angle heat map A (sin 2θ) and grasping angle heat map A (cos 2θ). Among them, the value range of the grasping quality score is [0, 1], the value range of the grasping width is [0, 100], and the grasping angle θ corresponds to [-π / 2, π / 2] one by one. The final grasping angle is
[0077] The quality heat map is used to display the potential quality or success rate of each pixel in the image as a grasping point. In the quality heat map, the depth of color (such as red to blue) indicates the quality of grasping. The darker the color (such as red), the higher the quality of the point as a grasping point, that is, the greater the possibility of successful grasping; the lighter the color (such as blue), the lower the grasping quality. Therefore, the quality heat map helps the robot quickly identify which areas are more suitable for grasping operations.
[0078] The grasp width heat map shows the required gripper opening width for each potential grasping point in the image. Similar to the quality heat map, light and dark colors or different tones represent different grasping width ranges in the grasp width heat map. The robot can adjust the opening width of its grippers according to the grasp width heat map to adapt to objects of different shapes and sizes.
[0079] The grasping angle heat map shows the grasping angle required for each potential grasping point in the image. The grasping angle is represented by two components, sin 2θ and cos 2θ. This is because using this representation method can avoid the problems caused by the periodicity of the angle and facilitate subsequent calculations and processing. Through the grasping angle heat map, the robot can intuitively understand which angles are more suitable for grasping, and adjust the rotation angle of its gripper according to the grasping angle heat map, thereby optimizing its grasping posture to ensure stable grasping of objects.
[0080] In a specific embodiment, the ITC-Grasp model is trained and tested using the Cornell Grasp dataset.
[0081] Specifically, the Cornell grasping dataset is randomly divided into training and test sets in a ratio of 9:1 using image segmentation (Image-Wise) and object segmentation (Object-Wise). Image segmentation is used to test the model's prediction performance for previously seen objects in different directions and postures. Object segmentation refers to instance division of all images in the dataset to ensure that objects in the training set do not appear in the test set, which is used to test the network's generalization ability for unknown objects.
[0082] In a specific embodiment, the ITC-Grasp model is trained using the Cornell grasp prediction dataset. The rectangular box grasp pose representation method used in the training process can be a five-dimensional grasp pose representation method. In order to better handle the instability caused by abnormal gradients in training, Smooth L1 is used as the loss function in the training process.
[0083] In this embodiment, the five-dimensional grasping representation g used is defined as:
[0084] g={x,y,w,h,θ} (7)
[0085] Among them, (x, y) represents the two-dimensional coordinates of the center point of the grasping rectangle, θ represents the rotation angle of the grasping rectangle, and (w, h) represents the size of the grasping rectangle, including the width w and the height h.
[0086] In order to perform real-time 2D plane grasping pose prediction, g is optimized and represented as G, which is defined as follows:
[0087]
[0088] Q represents the grasping quality score. The higher the confidence of Q, the more suitable G is for grasping objects. W represents the grasping width, and A represents the grasping angle. represents a real number, T is the width of the feature map, and P is the height of the feature map.
[0089] In this embodiment, the Smooth L1 loss function is used, which is defined as follows:
[0090]
[0091] in, represents the total loss, n is the number of samples, Z i is the loss of the i-th sample, where i represents the sample number.
[0092] Specifically, Z i is defined as follows:
[0093]
[0094] Among them, G i represents the predicted grab rectangle, represents the actual grasping rectangle, and n is the number of samples.
[0095] Specifically, a five-dimensional grasping representation method is used to represent the grasping posture of the predicted grasping rectangle, where the indicators for evaluating whether the predicted grasping rectangle is correct are as follows:
[0096] (1) The offset between the predicted grasping rectangle and the ground truth rectangle is less than 30°;
[0097] (2) The Jaccard index between the predicted grasped rectangle and the ground truth rectangle is greater than 25%.
[0098] More specifically, the Jaccard index is defined as follows:
[0099]
[0100] Among them, J(G P ,G t ) represents the Jaccard index; G p , G t are the predicted grasp rectangle and the ground truth rectangle, respectively.
[0101] In this embodiment, the training and testing of the grasping pose prediction model are carried out under the Windows 11 system. The hardware facilities are Intel(R) Core(TM) i5-11400H processor, 16GB memory, NVIDIA RTX Geforce 8GB graphics card, the deep learning framework is Pytorch1.7 deep learning framework and CUDA 11.6 running version, and cuDNN 11.0 deep learning GPU acceleration library. In each training step, the initial learning rate is set to 0.001, and the optimizer used is Adam optimizer.
[0102] In a specific embodiment, the Jaccard index model is trained and tested using the Cornell dataset to verify the grasping posture prediction performance of the model proposed in this embodiment. The model achieved an accuracy rate of 98.2% on the Cornell dataset. Compared with other similar methods, the grasping accuracy is higher, and the prediction speed for each image is 32.7ms, which meets the real-time requirements.
[0103] As shown in Table 1, there is a comparison between the prediction results of this embodiment and those of the prior art.
[0104] Table 1 Comparison of grasp prediction results on the Cornell dataset
[0105]
[0106] It can be seen from Table 1 that the Transformer-CNN-based planar grasping posture processing method of this embodiment has higher accuracy in predicting the grasping posture.
[0107] like Figure 5 The figure shows the test results of the ITC-Grasp model provided in this embodiment on the Cornell grasping test set, where the first row is the grasping rectangle prediction map, and the second, third and fourth rows are the grasping quality heat map, grasping angle heat map and grasping width heat map respectively. Figure 5 It can be seen that the Transformer-CNN-based planar grasping posture processing method of this embodiment can accurately predict the optimal grasping position of the object, and the grasping quality score is high.
[0108] The technical solution of the present invention can combine the powerful ability of Transformer to extract global information with the local information modeling ability of CNN to obtain more robust and comprehensive feature extraction and representation capabilities, while strengthening the multi-level feature interaction capabilities, and improving the accuracy and generalization ability of grasping posture prediction. It not only improves the accuracy of the model's grasping posture prediction for specific objects, but also significantly enhances its generalization ability.
[0109] In a specific embodiment, a processing system for plane grasping posture based on Transformer-CNN includes a data acquisition module: preprocessing a prediction data set to acquire an input image;
[0110] Model building module: build a plane grasping pose prediction model based on Transformer-CNN;
[0111] Model training module: Input the acquired input image into the constructed plane grasping pose prediction model for model training and testing:
[0112] Result prediction module: Use the trained planar grasping pose prediction model to obtain the grasping configuration and generate prediction results.
[0113] like Figure 6As shown, the prediction results of the ITC-Grasp model of this embodiment on the Cornell grasping dataset are compared with the prediction results of GGCNN and GR-CovnNet on the Cormell grasping dataset. GGCNN (Generative Grasping Convolutional Neural Network) is a generative grasping convolutional neural network that can directly generate grasping postures through images. GR-ConvNet (Generative Residual Convolutional Neural Network) is a generative residual convolutional neural network that can generate robust contralateral grasps from multi-channel inputs. Among them, the Grasp grasping map represents an image related to grasping, including information about the location of the grasping point, the grasping direction, and the magnitude of the grasping force; the Quality grasping quality heat map represents the potential quality or success rate of each pixel in the image as a grasping point, and its color depth (such as red to blue) indicates the quality of the grasping, and the darker the color, the greater the possibility of successful grasping; the Angle grasping angle heat map shows the grasping angle required for each potential grasping point in the image; the Width grasping width heat map shows the jaw opening width required for each potential grasping point in the image. Through Figure 6 It can be seen that GGCNN cannot identify and locate the correct grasping area well, and the grasping quality heat map score is low, while GR-ConvNet fails to distinguish the outline and background of the object well, resulting in an inaccurate optimal grasping rectangle. The ITC-Grasp model proposed in this embodiment not only obtains an accurate predicted grasping rectangle with a high grasping quality score, but also improves the separability of the object and the background, and the prediction of the grasping angle and grasping width is also within the correct range.
[0114] The above schematically describes the invention and its implementation methods, which is not restrictive. Without departing from the spirit or basic features of the invention, the invention can be implemented in other specific forms. What is shown in the accompanying drawings is only one of the implementation methods of the invention. The actual structure is not limited to this, and any figure mark in the claims should not limit the claims involved. Therefore, if a person of ordinary skill in the art is inspired by it, without departing from the purpose of the invention, a structural method and an embodiment similar to the technical solution are designed without creativity, which should all fall within the scope of protection of this patent. In addition, the word "including" does not exclude other elements or steps, and the word "one" before the element does not exclude the inclusion of "multiple" elements. The multiple elements stated in the product claim can also be implemented by one element through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any specific order.
Claims
1. A method for processing plane grasping posture based on Transformer-CNN, comprising the following steps: Preprocess the prediction data set to obtain the input image; Construct a plane grasping pose prediction model based on Transformer-CNN; the plane grasping pose prediction model includes an encoder module, a decoder module and a grasping prediction module; Input the acquired input image into the constructed plane grasping pose prediction model for model training and testing: Use the trained planar grasping pose prediction model to obtain the grasping configuration and generate prediction results.
2. The method for processing plane grasping posture based on Transformer-CNN according to claim 1, characterized in that: The encoder module has an Inception mixer at each stage; The Inception mixer includes a high-frequency mixer and a low-frequency mixer; the high-frequency mixer is composed of deep convolution and maximum pooling, and the low-frequency mixer is implemented by average pooling and self-attention.
3. The method for processing plane grasping posture based on Transformer-CNN according to claim 2, characterized in that: Assume that the input of the planar grasping pose prediction model is X, X divides the channel dimension C into C h and C l Channels, the calculation process in the Inception mixer is as follows: Y h1 =FC(MaxPool(X h1 )); Y h2 =DWConv(FC(X h2 )); Y l =Upsample(MSA(Avepooling(X l ))); AND c =Concat(And l ,AND h1 ,AND h2 ); Y h1 represents the output of the high-frequency mixer after maximum pooling, Y h2 represents the output of the high-frequency mixer after depthwise separable convolution, Y l is the output of the low frequency mixer, Y c is the output of the Inception mixer, X h1 To feed Y h1 The number of channels is C h / 2 feature map, X h2 To feed Y h2 The number of channels is C h / 2 feature map, X l The number of channels in the low frequency mixer is C l feature map, MSA represents multi-head self-attention; FC() represents full connection; MaxPool() represents maximum pooling; DWConv() represents depthwise separable convolution; Avepooling() represents average pooling; MSA() represents multi-head self-attention mechanism; Upsample() represents upsampling operation; Concat() represents concatenation.
4. The method for processing plane grasping posture based on Transformer-CNN according to claim 2, characterized in that: The acquired input image is input into the trained planar grasping pose prediction model, and the grasping configuration is obtained through the subtask network of the grasping prediction module, including the quality heat map, width heat map and angle heat map.
5. The method for processing plane grasping posture based on Transformer-CNN according to claim 1, characterized in that: The prediction dataset includes the Cornell grasping dataset and the Jacquard grasping dataset. The process of preprocessing the Cornell grasping dataset is as follows: data enhancement is performed on all images in turn; the image size is cropped to a fixed pixel; and the cropped image is normalized.
6. The method for processing plane grasping posture based on Transformer-CNN according to claim 5, characterized in that: The Cornell grasping dataset is divided into a training set and a test set in a ratio of 9:1 using image segmentation and object segmentation. During the model training process, the training parameters are adjusted until the convergence of the plane grasping pose prediction model slows down. The training parameters include the training batch size, the number of training rounds, and the learning rate.
7. The method for processing plane grasping posture based on Transformer-CNN according to claim 5, characterized in that: The ITC-Grasp model is trained using the Cornell grasp prediction dataset, using a five-dimensional grasp pose representation method. The five-dimensional grasp representation g is defined as follows: g = {x, y, w, h, θ}; (x, y) represents the two-dimensional coordinates of the center point of the grasping rectangle, θ represents the rotation angle of the grasping rectangle, and (w, h) represents the size of the grasping rectangle, where w represents the width and h represents the height.
8. The method for processing plane grasping posture based on Transformer-CNN according to claim 7, characterized in that: The indicators for evaluating whether the predicted grasping rectangle is correct include: the offset between the grasping box of the predicted grasping rectangle and the ground truth rectangle is less than 30°; the Jaccard index between the predicted grasping rectangle and the ground truth rectangle is greater than 25%.
9. The method for processing plane grasping posture based on Transformer-CNN according to claim 5, characterized in that: Smooth L1 is used as the loss function in the model training process, which is defined as follows: represents the total loss, n is the number of samples, Z i is the loss of the i-th sample, where i represents the sample number.
10. The Transformer-CNN-based plane grasping posture processing system according to any one of claims 1 to 9, characterized in that: It includes a data acquisition module: preprocessing the prediction data set and acquiring the input image; Model building module: Construct a plane grasping pose prediction model based on Transformer-CNN; the plane grasping pose prediction model includes an encoder module, a decoder module and a grasping prediction module; Model training module: Input the acquired input image into the constructed plane grasping pose prediction model for model training and testing: Result prediction module: Use the trained planar grasping pose prediction model to obtain the grasping configuration and generate prediction results.
Citation Information
Patent Citations
Composite robot dynamic grabbing method and system based on visual servo
CN117381780A