An object grasping method based on multimodal fusion
Through multimodal fusion and knowledge distillation methods, a capture inference network is built, which solves the problem of robot capture detection forgetting in different environments, and achieves the improvement of capture accuracy and balance of learning ability.
Patent Information
- Application Number
- CN202510381684.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-28
AI Technical Summary
Existing robotic capture and detection technologies are prone to catastrophic forgetting when facing different working environments or objects, and it is difficult to balance the stability and plasticity of capture learning.
The object grasping method of multimodal fusion is adopted, and the capture inference network is constructed, and color images and depth feature maps are used, multimodal fusion modules, residual blocks and transposed convolution modules are combined, and feature extraction and fusion is extracted and fusion is performed with attention mechanisms. The knowledge distillation method is used for model training, and the teacher model is used to guide the learning of students' models.
It effectively improves the robot's grasping accuracy in new scenarios, reduces the forgetting of old tasks, balances the stability and plasticity of grab learning, and improves object-level and image-level grasping accuracy.
Smart Images

Figure CN119888209B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent grasping detection, and particularly relates to an object grasping method based on multimodal fusion. Background Art
[0002] Robot grasping technology involves multiple key fields such as grasping pose detection, robot control, and motion planning. Among them, during the robot grasping process, the detection of the grasping pose plays a crucial role, providing guidance for subsequent control and motion planning of the robot's execution of grasping, and is the primary step of robot grasping. Currently, robot grasping detection technologies are mainly divided into analysis-based methods and data-driven methods. Data-driven grasping methods are further divided into discriminative grasping algorithms and generative grasping algorithms. The generative grasping algorithm does not rely on sampling of grasping candidate objects, but directly generates grasping postures at the pixel level, so it has good real-time performance and is suitable for closed-loop grasping. The discriminative type is divided into two steps: generating candidate grasping frames and selecting the optimal grasping frame. The generative grasping algorithm currently generally uses color images and depth images as inputs, but the existing methods often fuse color image information and depth image information in the channel dimension, which will increase the computational complexity.
[0003] On the other hand, in the real world, robots often face different working environments or different grasping objects. When the robot switches to a completely new environment for grasping operations, due to problems such as data loss and waste of computing power in reusing all data to train the model, it is obviously unrealistic to retrain the robot using all data again. If only a very limited amount of data in the new task scenario is simply used to train the robot, although the robot may show a certain degree of excellent performance in the new scenario, this locally optimized strategy hides huge risks. When there are significant differences between the new scenario and the old scenario, the robot is extremely prone to catastrophic forgetting in the old scenario. Therefore, it is particularly important to explore a training method that can make the robot perform excellently in new tasks without excessive forgetting of old tasks, that is, to balance the stability and plasticity of the robot's grasping learning, so that the robot maintains its original grasping ability and has the ability to learn new grasping skills. Summary of the Invention
[0004] The present invention provides an object grasping method based on multimodal fusion, which can make the robot perform excellently in the grasping detection task of the new scenario without excessive forgetting of the grasping detection task of the old scenario, and effectively balance the stability and plasticity of the robot's grasping learning.
[0005] To achieve the above technical objectives, the present invention adopts the following technical solutions:
[0006] An object grasping method based on multi-modal fusion, comprising:
[0007] S1. Construct m identical grasping inference networks, with the input being a color image and a depth feature map, and the output being a feature map of the grasping pose;
[0008] The grasping inference network includes a multi-modal fusion module, a residual block, and a transposed convolution module; the multi-modal fusion module includes 3 feature extraction modules and 2 attention layers; the inputs of the first, second, and third feature extraction modules are respectively: a depth image, a color image, a dark image and a depth image; the first attention layer is used to fuse the output features of the first and second feature extraction modules, and the second attention layer is used to fuse the output features of the first attention layer and the third feature extraction module;
[0009] S2. Obtain training data sets for different scenarios, and mark the grasping poses of each training data to obtain a labeled data set for the corresponding scenario;
[0010] S3. Use the labeled data set of the first scenario to train each grasping inference network to obtain m grasping detection models;
[0011] Use each trained grasping detection model to perform grasping detection on the first scenario to generate a grasping pose;
[0012] S4. Let i = 1;
[0013] S5. Select the best one from all the grasping detection models of the i-th scenario as the initialization benchmark for the teacher model and the student model of the (i + 1)-th scenario;
[0014] Take partial training data from the labeled data sets of the previous i scenarios and the labeled data set of the (i + 1)-th scenario, and use the knowledge distillation method and the teacher model as a guide to conduct guiding training on the currently initialized student model;
[0015] Use the current grasping detection models to perform grasping detection on the previous (i + 1) scenarios to generate grasping poses;
[0016] Update i = i + 1, and repeat this step S5 until all the grasping detection models for all scenarios are obtained.
[0017] Furthermore, the running process of each attention layer is as follows: both of the input feature maps pass through an average pooling layer and a max pooling layer; then the four obtained output features are stacked in the channel dimension, and finally the convolution layer is used to output the weights of the two feature maps input to the attention layer; expressed as:
[0018] ;
[0019] ;
[0020] ;
[0021] ;
[0022] wherein, , , respectively represent the features extracted by the first, second, and third feature extraction modules from the depth image, color image, dark image, and depth image, represents the first attention layer, , are respectively the output weights of the first attention layer for its two input feature maps, represents the feature map after fusing the color image feature and the depth image feature, represents the second attention layer, , are respectively the output weights of the second attention layer for its two input feature maps, represents the output of the multi-modal fusion module.
[0023] Furthermore, the grasping pose feature map output by the grasping inference network includes: a grasping angle sine map, a grasping angle cosine map, a grasping width map, and a grasping quality map; all the output grasping pose feature maps are synthesized to obtain the following grasping pose : the position of the center of the gripper in the world coordinate system , the rotation angle of the gripper around the Z axis of the world coordinate system , the opening width of the gripper , the width of the gripper itself , the grasping quality score , expressed as .
[0024] Furthermore, based on the two measurement metrics of object-level grasping accuracy and image-level grasping accuracy, the best model is selected from all the grasping detection models; if the two measurement metrics of object-level grasping accuracy and image-level grasping accuracy do not reach the best simultaneously, the model with higher image-level grasping accuracy is preferentially considered as the best model; wherein, the calculation formulas of object-level grasping accuracy and image-level grasping accuracy are:
[0025] ;
[0026] ;
[0027] wherein, OL represents object-level, IL represents image-level, represents object-level grasping accuracy, represents image-level grasping accuracy, The number of objects in the i-th figure, represents the number of objects successfully grasped in the i-th figure, The total number of pictures in the validation set, represents the number of pictures in which all objects in the validation set pictures are correctly grasped.
[0028] Furthermore, in step S5, the knowledge distillation method is adopted and the teacher model is used as a guide to train the currently initialized student model. Its loss function includes the distillation loss of the teacher model guiding the student model training and the label loss of the student model training.
[0029] Furthermore, the distillation loss includes the differences between the student model and the teacher model in each dimension of the output grasping pose.
[0030] Furthermore, the distillation loss also includes the spatial feature loss between the output feature maps of the intermediate convolutional layers of the student model and the teacher model;
[0031] The distillation loss is expressed as:
[0032] ;
[0033] In the formula, represents the distillation loss, respectively represent the cosine distillation loss of the grasping angle, the sine distillation loss of the grasping angle, the distillation loss of the grasping width, and the distillation loss of the grasping quality, represents the spatial feature loss of the intermediate convolutional layer;
[0034] ;
[0035] In the formula, represents the two-dimensional vector after flattening the output feature map of the teacher model, represents the two-dimensional vector after flattening the output feature map of the student model, represents the cosine similarity calculation;
[0036] ;
[0037] ;
[0038] ;
[0039] In the formula, is used to measure the difference between the outputs of the intermediate convolutional layers of the teacher model and the student model in the width direction, is used to measure the difference between the outputs of the intermediate convolutional layers of the teacher model and the student model in the height direction, represents the feature map output by the intermediate convolutional layer of the student model, Represents the feature map output by the intermediate convolutional layer of the teacher model, respectively represent the number of channels, height, and width of the feature map, respectively represent the channel number of the feature map and the numbers of the pixel points in the height and width directions.
[0040] Furthermore, the label loss includes the sum of the losses in each dimension between the grasping pose output by the student model and the true label, expressed as:
[0041] ;
[0042] In the formula, represents the label loss, respectively represent the cosine label loss of the grasping angle, the sine label loss of the grasping angle, the label loss of the grasping width, and the label loss of the grasping quality. The loss in each dimension adopts the SmoothL1 loss:
[0043] ;
[0044] In the formula, represents the grasping pose output by the student model, represents the true grasping pose label.
[0045] The simple and efficient multi-modal fusion module proposed by the present invention is different from the previous grasping inference models that fuse color image information and depth image information in the channel dimension. It makes full use of color image information and depth image information. And the continuous learning strategy proposed by the present invention greatly reduces the forgetting of the model for old tasks, balancing plasticity and stability. Compared with the existing technologies, it has the following advantages:
[0046] Different from the previous fusion of color image information and depth image information in the channel dimension, the present invention makes full use of color image information and depth image information on the basis of only increasing a small amount of computational effort. On the KWG dataset, compared with the basic model, the object-level grasping accuracy is improved by 0.86%, and the image-level grasping accuracy is improved by 0.65%.
[0047] The model proposed by the present invention is more suitable for continuous learning than the benchmark model. When using the same continuous learning strategy, the model proposed by the present invention performs the best. At the same time, the present invention proposes to use the calculation of cosine similarity for the distillation loss in the continuous learning of the robot grasping inference model, making full use of the rich knowledge reserve and mature experience contained in the teacher model to accurately guide and effectively standardize the learning process of the student model, thereby greatly reducing the forgetting of the student model for past knowledge during the continuous learning process. The training strategy balances the stability and plasticity of the student model, making the average accuracy effectively improved during the continuous learning of five tasks. Description of the Drawings
[0048] Figure 1 This is a schematic structural diagram of the present invention for multi-modal fusion in the spatial dimension.
[0049] Figure 2 This is a schematic network flow diagram of the method for generating grasping poses based on modal fusion of the present invention.
[0050] Figure 3 This is the continuous learning training process proposed by the present invention.
[0051] Figure 4 This is a schematic diagram of the teacher-student architecture adopted for continuous learning of the present invention. Detailed Embodiment
[0052] The following provides a detailed description of the embodiments of the present invention. These embodiments are carried out based on the technical solutions of the present invention, presenting detailed implementation manners and specific operation processes, and further explaining the technical solutions of the present invention.
[0053] To make full use of color image information and depth image information, and endow the robot with the ability to learn new grasping skills while maintaining its original grasping ability, the object grasping method based on multi-modal fusion in this embodiment includes:
[0054] Step 1: Construct m identical grasping inference networks. The input is a color image and a depth feature map, and the output is the grasping pose of the object.
[0055] The grasping inference network includes a multi-modal fusion module, a residual block, and a transposed convolution module.
[0056] Currently, the existing grasping inference networks mainly integrate depth image information and color image information in the channel dimension. However, this method will greatly increase the computational complexity. The multi-modal fusion module of the present invention, as Figure 1 shown, includes 3 feature extraction modules with the same network structure but independent of each other, and 2 attention layers.
[0057] The inputs of the first, second, and third feature extraction modules are respectively: a depth image, a color image, and the fusion information of the color image and the depth image.
[0058] In addition, the 2 attention layers use the method of generating weights by spatial attention to fuse multi-modal information in the spatial dimension: the first attention layer is used to fuse the output features of the first and second feature extraction modules, and the second attention layer is used to fuse the output features of the first attention layer and the third feature extraction module.
[0059] The operation process of each attention layer is as follows: The two input feature maps are respectively passed through the average pooling layer and the max pooling layer; subsequently, the four obtained output features are stacked in the channel dimension, and finally, the convolution layer outputs the weights of the two feature maps input to the attention layer. Furthermore, the two weights generated by the attention mechanism are used to perform weighted summation processing on the corresponding two feature maps. In this way, the information carried by the relatively unimportant parts in the color image and the depth image can be effectively suppressed by the attention mechanism, so that the fused multi-modal information focuses more on the key features, improving the accuracy and effectiveness of the overall information processing. It is expressed by the formula as follows:
[0060] ;
[0061] ;
[0062] ;
[0063] ;
[0064] In the formula, 、 、 respectively represent the features extracted from the depth image, color image, dark image and depth image by the first, second and third feature extraction modules, represents the first attention layer, 、 are respectively the output weights of the first attention layer for its two input feature maps, represents the feature map after fusing the color image feature and the depth image feature, represents the second attention layer, 、 are respectively the output weights of the second attention layer for its two input feature maps, represents the output of the multi-modal fusion module.
[0065] The feature obtained by the multi-fusion module is input into the residual block. After being processed by multiple residual blocks, finally, upsampling is performed through transposed convolution to output the grasping quality map, grasping width map, sine map of the grasping angle, and cosine map of the grasping angle for synthesizing the grasping pose, as Figure 2 shown.
[0066] The grasping pose finally output by the grasping inference network includes: the position of the center of the gripper in the world coordinate system , the rotation angle of the gripper around the Z axis of the world coordinate system, the grasping width of the gripper, and the width of the gripper itself., grasping quality score , denoted as .
[0067] Step 2: Obtain the training data sets for different scenarios, and mark the grasping poses of each training data to obtain the marked data sets for the corresponding scenarios.
[0068] Step 3: Use the marked data sets of the first scenario to train each grasping inference network to obtain m grasping detection models.
[0069] In this embodiment, in step 3, an existing training method can be used to train each grasping inference network, that is, each trained grasping detection model can be used to perform grasping detection on the first scenario to generate grasping poses.
[0070] Step 4: Let i = 1.
[0071] Step 5: When performing the grasping detection task for other new scenarios currently, it is necessary to further train the existing trained grasping detection models using the annotation data sets of the new scenarios.
[0072] The embodiment of the present invention adopts a continual learning training strategy, and the training process is as Figure 3 , this strategy constructs a teacher-student architecture mode. In this architecture, the structures of the teacher model and the student model are the same. By fully leveraging the rich knowledge reserve and mature experience contained in the teacher model, the learning process of the student model is accurately guided and effectively regulated, thereby greatly reducing the forgetting of past knowledge by the student model during the process of continual learning. When starting the training of the next new task, the model with the best overall performance in the previous task will be selected as the initialization benchmark for the teacher model and the student model of the next task. For example, the best model selected after the completion of task 0 will be set as the teacher model in task 1, and this model will be used to initialize the student model of task 1. During the continual learning process of a certain new task, the teacher model will be frozen and its internal parameters will no longer be updated.
[0073] Step 5.1: Select the best one from all the grasping detection models of the i-th scenario as the initialization benchmark for the teacher model and the student model of the (i + 1)-th scenario.
[0074] To evaluate the ability of the network to predict the grasping pose, the correct grasping box should meet the following two conditions:
[0075] (1) The area of the intersection part of the predicted grasping box and the true grasping box divided by the area of the union of the two grasping boxes should be greater than 25%.
[0076] (2) The deviation angle between the predicted grasping box and the true grasping box should be less than 30°.
[0077] To evaluate the effect of the trained grasping detection model, the embodiments of the present invention select the best model from all grasping detection models based on two measurement metrics: object-level grasping accuracy and image-level grasping accuracy. If the two measurement metrics of object-level grasping accuracy and image-level grasping accuracy do not reach the best at the same time, the model with higher image-level grasping accuracy is preferentially considered as the best model.
[0078] The object-level grasping accuracy represents the proportion of objects that can be grasped among all objects, and the image-level grasping accuracy is the proportion of images considered correct. If there is only one effective grasp for each object in the image, the image is considered correct; otherwise, it is incorrect.
[0079] Among them, the calculation formulas for object-level grasping accuracy and image-level grasping accuracy are as follows:
[0080] ;
[0081] ;
[0082] In the formula, OL represents object-level, IL represents image-level, represents object-level grasping accuracy, represents image-level grasping accuracy, the number of objects in the i-th image, represents the number of objects successfully grasped in the i-th image, represents the total number of pictures in the validation set, represents the number of pictures in which all objects in the validation set pictures are correctly grasped.
[0083] Step 5.2: Take partial training data from the labeled data sets of the previous i scenarios and the labeled data set of the (i + 1)-th scenario, and use the knowledge distillation method and the teacher model as a guide to conduct guided training on the currently initialized student model.
[0084] In this embodiment, the knowledge distillation method is used and the teacher model is used as a guide to conduct guided training on the currently initialized student model. Its loss function includes the distillation loss of the teacher model guiding the student model training and the label loss of the student model training.
[0085] (1) Distillation loss.
[0086] In the embodiments of the present invention, the distillation loss includes the differences between the output grasping poses of the student model and the teacher model in each dimension, and further may include the spatial feature loss between the output feature maps of the intermediate convolutional layers of the student model and the teacher model. It is expressed as:
[0087] ;
[0088] In the formula, denotes the distillation loss, respectively denote the cosine distillation loss of the grasping angle, the sine distillation loss of the grasping angle, the distillation loss of the grasping width, and the distillation loss of the grasping quality. denotes the spatial feature loss of the intermediate convolutional layer.
[0089] As Figure 4 shown, the output form of the generative machine grasping point generation network is the sine map of the grasping angle, the cosine map of the grasping angle, and the width map of the grasping, which is significantly different from the traditional classification model. Therefore, if the output of the teacher model is subjected to conventional softening processing and the KL divergence or cross-entropy loss is used to calculate the distillation loss, it is not applicable to the robot grasping point generation network. Based on this, the present invention proposes to use the cosine similarity to calculate the distillation loss of the grasping quality, the grasping width, the sine value of the grasping angle, and the cosine value of the grasping angle. In the invention, the output feature map size of the grasping inference network is , and the calculation steps are to first transform the output shapes of the student model and the teacher model into ( ), and then calculate the cosine similarity, as shown in the following formula:
[0090] ;
[0091] ;
[0092] ;
[0093] ;
[0094] In the formula, denotes the two-dimensional vector after flattening the output of the teacher model, denotes the two-dimensional vector after flattening the output of the student model, denotes the cosine similarity calculation.
[0095] The embodiment of the present invention also uses the spatial feature loss as part of the distillation loss, which is applied to the intermediate layer of the network, and the formula is expressed as follows:
[0096] ;
[0097] ;
[0098] In the formula, is used to measure the difference in the width direction between the output of the intermediate convolutional layer of the teacher model and the student model, is used to measure the difference in the height direction between the output of the intermediate convolutional layer of the teacher model and the student model, denotes the feature map of the output of the intermediate convolutional layer of the student model, Represents the feature map output by the intermediate convolutional layer of the teacher model, respectively represent the number of channels, height, and width of the feature map, respectively represent the channel number of the feature map and the numbers of the pixel points in the height and width directions.
[0099] In the embodiment of the present invention, the differences in the width and height directions between the outputs of the intermediate convolutional layers of the teacher model and the student model , together constitute the spatial feature loss , which is expressed by the formula as follows:
[0100] ;
[0101] (2) Label loss.
[0102] The label loss includes the sum of the losses in each dimension between the grasping pose output by the student model and the true label, and is expressed as:
[0103] ;
[0104] In the formula, represents the label loss, respectively represent the cosine label loss of the grasping angle, the sine label loss of the grasping angle, the label loss of the grasping width, and the label loss of the grasping quality.
[0105] In this embodiment, the loss in each dimension adopts the SmoothL1 loss:
[0106] ;
[0107] In the formula, represents the grasping pose output by the student model, represents the true grasping pose label.
[0108] Step 5.3, use the current grasping detection models to perform grasping detection on the first i + 1 scenes, and generate grasping poses.
[0109] Step 5.4, update i = i + 1, and return to Step 5.1 until the grasping detection models for all scenes are obtained.
[0110] In the continuous learning of new scene tasks, the goal is to balance stability and plasticity. The backward transfer (BF) measures the degree of forgetting. If the degree of forgetting is high, it means that the model tends to adapt to new tasks and the stability is insufficient. The backward transfer after training on the i-th scene task is defined as the following formula. The lower the BF value, the stronger the ability of the model to overcome forgetting.
[0111] ;
[0112] To evaluate the overall performance of the model on all tasks trained during the continuous learning process, average precision is usually used as a performance metric. It is expressed by the following formula.
[0113] ;
[0114] where t represents the number of tasks that have been trained so far, represents the grasping precision of the model for the j-th task when training the t-th task, and s represents the number of all new tasks.
[0115] Example:
[0116] Step 1: Construct the robot grasping continuous learning dataset KWG-CL. The dataset is a conveyor belt grasping dataset captured by a Realsense D455 camera, with a total of 10 categories, including paper cups, paper boxes, plastic bottles, plastic boxes, glass bottles, metal bottles, wooden blocks, ceramic cups, foam, and others. According to different scenarios, the dataset is divided into five tasks. Task 1 is a simple background, containing 209 RGB-D images; Task 2 is a background with debris, containing 242 RGB-D images; Task 3 is a background with a plastic bag laid on the conveyor belt, containing 256 RGB-D images; Task 4 is a background with a plastic bag and non-oily impurities, containing 217 RGB-D images; Task 5 is a background with a plastic bag and oily and moist impurities, containing 240 RGB-D images. All color images are aligned with the depth images.
[0117] Step 2: Use GR-ConvNet as the baseline network and add a three-branch multi-modal fusion module. The multi-modal fusion module is as Figure 1, the model is trained on the KWG2024 dataset, and the dataset is divided into a training set and a test set in a ratio of 8:2. In this example, the training requires a system of Ubuntu 18.04.5 LTS or higher, and the system environment requires Python 3.8.5, Pytroch 1.10.1 or higher. The hardware platform needs to meet the requirements that the graphics card is NVIDIA GeForce RTX4090 (24G), the memory is more than 16G, and the hard disk capacity is not less than 256G. When performing model training, a total of 150 epochs are trained, the Adam optimizer is used for parameter update, and the learning rate adjustment strategy adopts the OneCycleLR learning rate scheduler. The main purpose of OneCycleLR is to dynamically change the learning rate during the training process to train the neural network more efficiently. Its characteristic is that within one training cycle, the learning rate first rises and then falls according to a specific pattern, and this change pattern is similar to the shape of a "triangle wave" or "periodic wave". This scheduling strategy helps the model converge faster and can improve the performance of the model to a certain extent. A random seed is used during training to obtain the same random split when running the training at different times, so that the experimental results can be reproduced. The comparison of the accuracy of the basic model and the improved model of the method of the present invention is shown in Table 1.
[0118] ;
[0119] Step 3: Regard the grasping scenarios included in the KWG2024 dataset as Task0, and continuously learn the model trained on the KWG2024 dataset on the KWG-CL dataset. Each task is trained for 20 rounds, and the continuous learning training process is as Figure 3 , when training each new task, re-initialize the learning rate scheduler, use the Adam optimizer for parameter update, the learning rate adjustment strategy adopts OneCycleLR, and the model is tested in the test set after each round of training, and the average accuracy is calculated for all the previously trained tasks.
[0120] In this example, continuous learning is divided into five tasks. The model with the highest average grasping accuracy in the validation set is selected as the basic prediction model for the subsequent steps. The continuous learning method proposed in the present invention is compared with the continuous learning method of memory replay. In the memory replay method, the total training dataset for each task is composed of 20 samples from the training sets of each old task and the training set of the current task. The memory replay method is widely regarded as one of the most effective means in the field of continuous learning. As shown in Table 2, the experimental results show that in the first three tasks, the continuous learning strategy proposed in the present invention is superior to memory replay. However, as the number of tasks increases, the number of samples in memory replay increases. Therefore, in Tasks 4 and 5, the continuous learning strategy proposed in the present invention is slightly inferior to memory replay. At the same time, the model of the present invention shows more excellent continuous learning adaptability compared with the baseline model. Under the same training strategy and after continuous learning of the same number of tasks, the average grasping accuracy of the model proposed in the present invention is higher than that of the baseline model in the first four tasks, and the grasping accuracy is basically the same in the fifth task, fully demonstrating its advantages in continuous learning tasks and being able to adapt to various different types of tasks. The forgetting rate in the fifth task is higher than that of the baseline model because the learning ability shown by the baseline model in the first four tasks is poor.
[0121] Meanwhile, the present invention proposes to use cosine similarity as the calculation method for the distillation loss. Compared with using SmoothL1 loss as the distillation loss, using cosine similarity to calculate the distillation loss results in an average accuracy of 1.16% after continuous learning of five tasks.
[0122] ;
[0123] The above embodiments are the preferred embodiments of the present application. Those of ordinary skill in the art can also make various transformations or improvements based on this. Without departing from the overall concept of the present application, these transformations or improvements should all fall within the scope of protection required by the present application.
Claims
1. An object grasping method based on multimodal fusion, characterized in that: include: S1, build m identical grasping inference networks, with color images and depth feature maps as input and feature maps of grasping postures as output; The crawling inference network includes a multimodal fusion module, a residual block and a transposed convolution module; the multimodal fusion module includes three feature extraction modules and two attention layers; the inputs of the first, second and third feature extraction modules are respectively: a depth image, a color image, a dark image and a depth image; the first attention layer is used to fuse the output features of the first and second feature extraction modules, and the second attention layer is used to fuse the output features of the first attention layer and the third feature extraction module; S2, obtaining training data sets of different scenes, and marking the grasping posture of each training data to obtain a marked data set of the corresponding scene; S3, using the labeled data set of the first scenario to train each grasping inference network to obtain m grasping detection models; Use the trained grasping detection models to perform grasping detection on the first scene and generate grasping postures; S4, let i=1; S5, select the best one from all grasp detection models of the i-th scene as the teacher model of the i+1-th scene and the initialization benchmark of the student model; Take part of the training data from the labeled datasets of the first i scenes, as well as the labeled dataset of the i+1th scene, adopt the knowledge distillation method and use the teacher model as a guide to guide the training of the currently initialized student model; Use the current grasping detection models to perform grasping detection on the previous i+1 scenes and generate grasping postures; Update i=i+1 and repeat step S5 until the grasping detection models for all scenes are obtained.
2. The object grasping method based on multimodal fusion according to claim 1, characterized in that: The operation process of each attention layer is as follows: both input feature maps are passed through the average pooling layer and the maximum pooling layer; The four output features are then stacked in the channel dimension, and finally the convolutional layer is used to output the weights of the two feature maps input to the attention layer; it is expressed as ; In the formula, , , Respectively represent the features extracted by the first, second and third feature extraction modules for the depth image, color image, dark image and depth image, respectively. represents the first attention layer, , are the output weights of the first attention layer for its two input feature maps, Represents the feature map after the fusion of color image features and depth image features, represents the second attention layer, , are the output weights of the second attention layer for its two input feature maps, Represents the output of the multimodal fusion module.
3. The object grasping method based on multimodal fusion according to claim 1, characterized in that: The grasping posture feature maps output by the grasping inference network include: grasping angle sine map, grasping angle cosine map, grasping width map and grasping quality map; all the output grasping posture feature maps are synthesized to obtain the following grasping posture : The position of the center of the gripper in the world coordinate system , the rotation angle of the gripper around the Z axis of the world coordinate system , the gripping width of the clamping jaws , the width of the gripper itself , crawl quality score , expressed as .
4. The object grasping method based on multimodal fusion according to claim 1 is characterized in that: Based on the two measurement indicators of object-level grasping accuracy and image-level grasping accuracy, the best model is selected from all grasping detection models; if the two measurement indicators of object-level grasping accuracy and image-level grasping accuracy do not reach the best at the same time, the model with higher image-level grasping accuracy is given priority as the best model; the calculation formulas of object-level grasping accuracy and image-level grasping accuracy are: ; ; In the formula, OL represents the object level, IL represents the image level, represents the object-level grasping accuracy, represents the image-level capture accuracy, The number of objects in the i-th image, Indicates the number of objects successfully captured in the i-th image, Represents the total number of images in the validation set, Indicates the number of images in the validation set where all objects are captured correctly.
5. The object grasping method based on multimodal fusion according to claim 1, characterized in that: In step S5, a knowledge distillation method is adopted and the teacher model is used as a guide to guide the training of the currently initialized student model. The loss function includes the distillation loss of the teacher model guiding the student model training and the label loss of the student model training.
6. The object grasping method based on multimodal fusion according to claim 5 is characterized in that: The distillation loss includes the difference between the student model and the teacher model in each dimension of the output grasp pose.
7. The object grasping method based on multimodal fusion according to claim 6, characterized in that: The distillation loss also includes the spatial feature loss between the output feature maps of the intermediate convolutional layers of the student model and the teacher model; The distillation loss is expressed as: ; In the formula, represents the distillation loss, They represent the cosine distillation loss of the grasping angle, the sine distillation loss of the grasping angle, the distillation loss of the grasping width, and the distillation loss of the grasping quality. Represents the spatial feature loss of the intermediate convolutional layer; ; In the formula, represents the two-dimensional vector after the teacher model output feature map is flattened, Represents the two-dimensional vector of the flattened output feature map of the student model, Indicates cosine similarity calculation; ; ; ; In the formula, It is used to measure the difference in width between the intermediate convolutional layer outputs of the teacher model and the student model. It is used to measure the difference in height between the output of the intermediate convolutional layer of the teacher model and the student model. Represents the feature map output by the intermediate convolutional layer of the student model, Represents the feature map output by the intermediate convolutional layer of the teacher model, Respectively represent the number of channels, height, and width of the feature map, They represent the channel number of the feature map and the pixel number in the height and width directions respectively.
8. The object grasping method based on multimodal fusion according to claim 5, characterized in that: The label loss includes the sum of the losses in each dimension between the grasping pose output by the student model and the true label, expressed as: ; In the formula, represents the label loss, They represent the cosine label loss of the grasping angle, the sine label loss of the grasping angle, the label loss of the grasping width, and the label loss of the grasping quality respectively. The loss of each dimension adopts the SmoothL1 loss: ; In the formula, represents the grasping pose output by the student model, Represents the actual grasp pose label.
Citation Information
Patent Citations
Water surface floating object monitoring method and system based on continuous learning
CN114022811A
Quasi-increment radiation source individual identification method based on knowledge distillation mechanism
CN114492745A