RGB-D multimodal semantic segmentation method
By improving the multi-head attention mechanism in Transformer and designing a self-attention multi-modal information interaction module, the challenges of multi-scale feature extraction and real-time performance in RGB-D semantic segmentation are solved, and RGB-D multi-modal semantic segmentation with high precision and high real-time performance are achieved.
Patent Information
- Application Number
- CN202310283961.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-03-22
AI Technical Summary
Existing RGB-D semantic segmentation methods have challenges in dealing with multi-scale features and real-time performance, especially under large-scale differences and real-time perception requirements.
By improving the multi-head attention mechanism in Transformer, a self-attention multi-modal information interaction module is designed to realize cross-modal information interaction and correction between color maps and depth maps, and combining multi-modal channel attention correction and global feature aggregation module to build an efficient RGB-D multi-modal semantic segmentation model.
It realizes high-precision RGB-D semantic segmentation, while maintaining high real-time performance, effectively solving the needs of multi-scale feature extraction and real-time performance.
Smart Images

Figure CN116597135B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and deep learning, and specifically relates to an RGB-D multimodal semantic segmentation method. Background Art
[0002] Image semantic segmentation is one of the important tasks in the field of computer vision. It is an effective scene understanding technology that aims to assign a category label to each pixel in the image and predict the position and outline of the object. Semantic segmentation has been widely used in industries such as autonomous driving, robot perception, and automatic navigation. In recent years, semantic segmentation methods based on color images have received increasing attention and have made significant progress in aspects such as segmentation accuracy. Current semantic segmentation methods cannot extract high-quality features in some cases. For example, when two objects have similar colors or textures, they cannot be accurately distinguished by color images alone. With the development of depth sensors, depth information is an important auxiliary information for semantic segmentation of color images. Compared with color images, depth images can provide richer geometric information. Therefore, studying the RGB-D semantic segmentation problem and exploring effective multimodal information fusion methods are of great significance to the application field of computer vision. At present, the RGB-D semantic segmentation method mainly faces the following problems:
[0003] (1) The scales of different objects in an image vary greatly. How to make full use of the multi-scale features in an image is a key issue.
[0004] (2) In practical applications, different devices need to perceive the surrounding environment in real time. How the RGB-D semantic segmentation method can meet the real-time performance under the premise of high accuracy is another key issue.
[0005] In summary, to address the above problems, an RGB-D multimodal semantic segmentation method is proposed. Transformer is introduced into RGB-D semantic segmentation. By utilizing the advantages of Transformer in the multimodal field, high-precision RGB-D semantic segmentation is achieved, which solves the multi-scale problem while maintaining high real-time performance. Summary of the invention
[0006] In view of the above problems, the purpose of the present invention is to provide an RGB-D multimodal semantic segmentation method. The method improves the multi-head attention mechanism in Transformer, and on this basis corrects and fuses the features of different modalities, thereby obtaining high-precision segmentation results while retaining high real-time performance.
[0007] The RGB-D multimodal semantic segmentation method includes the following steps:
[0008] S1. Design a self-attention multimodal information interaction module: The self-attention multimodal information interaction module is mainly used to realize cross-modal information interaction between color images and depth images in two dimensions: channel dimension and spatial dimension;
[0009] S2. Establish a multimodal semantic segmentation model based on RGB-D: the multimodal semantic segmentation model based on RGB-D includes a dual-stream feature extraction backbone network, a multimodal channel attention correction module, a multimodal global feature aggregation module and a feature pyramid decoder module; the multimodal feature extraction backbone network is used to extract features of the color image and the depth image respectively to generate feature maps of different sizes; the multimodal channel attention correction module is used to perform feature correction on the channel dimension of the feature maps of different sizes generated by the multimodal feature extraction backbone network to generate multimodal features after channel correction; the multimodal global feature aggregation module is used to perform feature aggregation on the spatial dimension of the corrected multimodal features generated by the multimodal channel attention correction module; the feature pyramid decoder module is used to decode the aggregated features generated by the multimodal global feature aggregation module to achieve prediction of two-dimensional semantic segmentation areas;
[0010] S3. Perform RGB-D based multimodal semantic segmentation model training: input the color image, the depth image and the semantic segmentation true label into the RGB-D based multimodal semantic segmentation model for training, and obtain the trained RGB-D based multimodal semantic segmentation model.
[0011] Furthermore, the self-attention multimodal information interaction module first calculates the query vector, key vector and value vector for the input color features and depth features respectively, wherein the acquisition of the three vectors is completed by the fully connected layer, and then the query vectors of the two modalities are exchanged, and the query vector of one modality and the transpose of the key vector of the other modality are used for matrix multiplication to calculate the self-attention matrix of each modality, and then the obtained self-attention matrix is matrix multiplied with the value vector to obtain the color features and depth features after information interaction, and finally the final output result is obtained through a fully connected layer to realize cross-modal information interaction. The above operation is expressed by the following formula:
[0012]
[0013]
[0014] RGB ii , Depth ii =FC(Attention RGB V RGB , Attention Depth V Depth )
[0015] In the formula, Q Depth , K Depth , V Depth Represent the query vector, key vector and value vector of the color feature respectively, Q Depth , K Depth , V Depth Represent the query vector, key vector and value vector of the deep feature respectively, d head Represents the dimension of the vector, Softmax represents the Softmax activation function, Attention RGB and Attention Depth Represents the self-attention matrix of color features and depth features respectively, FC represents the fully connected layer, RGB ii and Depth ii They respectively represent the color features and depth features after information interaction.
[0016] Furthermore, the RGB-D based multimodal semantic segmentation model is composed of a dual-stream pvt_v2 backbone network, four multimodal channel attention correction modules and four multimodal global feature aggregation modules. The dual-stream pvt_v2 backbone network extracts features of the color image and the depth image respectively, and passes the features into the four multimodal channel attention correction modules after each downsampling; the four multimodal channel attention correction modules perform feature correction on the color features and depth features extracted by the dual-stream pvt_v2 backbone network in the channel dimension, and pass the corrected features into the multimodal global feature aggregation module; the four multimodal global feature aggregation modules perform feature aggregation on the corrected color features and depth features output by the four multimodal channel attention correction modules in the spatial dimension.
[0017] Furthermore, the multimodal channel attention correction module first downsamples the color features and depth features through pooling operations of different sizes, flattens the pooling results and splices them along the second dimension to obtain multi-scale color features and multi-scale depth features, and then maps the dimensions to higher dimensions through a fully connected layer to obtain vectorized color features and depth features. The vectorized color features and depth features are respectively added with learnable modal codes and then passed into the Transformer module for global attention modeling. The calculation process of each Transformer can be expressed by the following formula:
[0018] z l =CMMHSA(LN(z t-1 ))+z l-1
[0019] z l =MLP(LN(z l ))+z l
[0020] In the formula, z l represents the input of the lth module, LN represents layer normalization, CMMHSA represents the self-attention multimodal information interaction module, and MLP represents a multilayer perceptron;
[0021] Finally, the modeled results are calculated through a multi-layer perceptron to obtain the channel attention vectors of the two modalities. After the channel attention vectors of the two modalities and the features of the two modalities are channel-multiplied, the corresponding elements are added and fused with the features of the other modalities to achieve multi-scale channel attention correction. The above operation is expressed by the following formula:
[0022] RGB msf =Concat(Flatten(msap(RGB in )), Flatten(msmp(RGB in )))
[0023] Depth msf =Concat(Flatten(msap(Depth in )), Flatten(msmp(Depth in )))
[0024] RGB tokenized , Depth tokenized =FC(RGB msf , Depth msf )
[0025]
[0026] W rgb , W depth =MLP(RGB cii , Depth cii )
[0027]
[0028]
[0029] In the formula, RGB in and Depth in Represents color features and depth features respectively, msmp and msap represent multi-scale maximum pooling and multi-scale average pooling respectively, Concat represents a merging operation, Flatten represents a flattening operation, RGB msf and Depth msf Represents multi-scale color features and multi-scale depth features, RGB tokenized and Depth tokenizedRepresents vectorized color features and depth features, RGB cii and Depth cii They represent the color features and depth features after channel information interaction, MTE rgb and MTE depth Represent the modal encoding of color features and depth features, W rgb and W depth Represent the channel attention vectors of the two modalities respectively, MLP represents the multi-layer perceptron, and Respectively represent the addition of corresponding elements and the multiplication of corresponding channels, RGB rec and Depth rec They represent the color features and depth features after multi-scale attention correction respectively.
[0030] Furthermore, the multimodal global feature aggregation module first embeds position information and modality information in the feature map, introduces position information through a depth-separable convolution with a convolution kernel size of 3×3, a step size of l×1, and a padding size of 1×1, and then adds it to the input features through a residual connection. In addition to the position encoding, a learnable modality encoding is also added to obtain color features and depth features with position information and modality information. Then, the self-attention multimodal information interaction module is used to interact with the spatial dimension and perform a residual connection with the input. At the same time, a spatial reduction module is introduced to reduce the amount of calculation through a sharing mechanism of key vectors and value vectors. Then, the color features and depth features after spatial information interaction are obtained through layer normalization. Finally, a 1×1 convolution is used to fuse the feature maps of the two modalities into a single feature map. In addition, in order to improve the robustness of the model, the original feature map is obtained through a depth-separable convolution of size 3×3 to obtain local features, and is fused with the global features through a residual connection, and then the final output is obtained through a batch normalization layer. The above calculation process can be expressed by the formula:
[0031]
[0032]
[0033]
[0034] F global =Conv 1×1 (Concat(RGB sii , Depth sii ))
[0035]
[0036] In the formula, pme represents position encoding and modality encoding, DWC3×3 Represents a depth-separable convolution with a kernel size of 3×3, RGB pme and Depth pme represents the color features and depth features after position encoding and modality encoding, SR represents the space reduction module, RGB sii and Depth sii They represent the color features and depth features after spatial information interaction, respectively. global Represents the global feature, Relu represents the Relu activation function, F local represents local features, BN represents batch normalization layer, F out Represents the final output.
[0037] Compared with the prior art, the present invention has the following beneficial effects:
[0038] 1. Effectively solve the problem of large differences in target scales in the scene;
[0039] 2. Effectively improve the accuracy of RGB-D semantic segmentation of scene targets;
[0040] 3. The Transformer-based cross-modal semantic segmentation method can simultaneously ensure the accuracy and real-time requirements of RGB-D semantic segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a self-attention multimodal information interaction module structure.
[0042] Figure 2 It is the overall structure of the RGB-D multimodal semantic segmentation model.
[0043] Figure 3 It is the multimodal channel attention correction module structure.
[0044] Figure 4 It is a multimodal global feature aggregation module structure.
[0045] Figure 5 It is the original color picture.
[0046] Figure 6 is the original depth image.
[0047] Figure 7 This is the effect picture after semantic segmentation. DETAILED DESCRIPTION
[0048] The technical solution of the present invention is described in detail below with reference to the accompanying drawings.
[0049] The RGB-D multimodal semantic segmentation method specifically includes the following steps:
[0050] S1. Design of self-attention multimodal information interaction module: as shown in the attached Figure 1 As shown in the figure, the self-attention multimodal information interaction module is mainly used to realize cross-modal information interaction between color images and depth images in two dimensions: channel dimension and spatial dimension.
[0051] First, the query vector, key vector and value vector are calculated for the input color features and depth features respectively. The acquisition of the three vectors is completed by the fully connected layer. Then, by exchanging the query vectors of the two modalities, the query vector of one modality and the transpose of the key vector of the other modality are used for matrix multiplication to calculate the self-attention matrix of each modality. The obtained self-attention matrix is then matrix multiplied with the value vector to obtain the color features and depth features after information interaction. Finally, a fully connected layer is used to obtain the final output result to realize cross-modal information interaction. The above operations are expressed by the following formula:
[0052]
[0053]
[0054] RGB ii , Depth ii =FC(Attention RGB V RGB , Attention Depth V Depth )
[0055] In the formula, Q Depth , K Depth , V Depth Represent the query vector, key vector and value vector of the color feature respectively, Q Depth , K Depth , V Depth Represent the query vector, key vector and value vector of the deep feature respectively, d head Represents the dimension of the vector, Softmax represents the Softmax activation function, Attention RGB and Attention Depth Represents the self-attention matrix of color features and depth features respectively, FC represents the fully connected layer, RGB ii and Depth ii They respectively represent the color features and depth features after information interaction.
[0056] S2. Establish a multimodal semantic segmentation model based on RGB-D: as shown in the attached Figure 2As shown in the figure, the RGB-D multimodal semantic segmentation model includes a two-stream pvt_v2 backbone network, a multimodal channel attention correction module, a multimodal global feature aggregation module and a feature pyramid decoder module; the two-stream pvt_v2 backbone network extracts the features of the color image and the depth image respectively, and passes the features to the multimodal channel attention correction module after each downsampling; the multimodal channel attention correction module performs feature correction on the color features and depth features in the channel dimension, and passes the corrected features to the multimodal global feature aggregation module; the multimodal global feature aggregation module performs feature aggregation on the corrected color features and depth features in the spatial dimension; finally, the feature pyramid decoder module is used to decode the aggregated features to realize the prediction of the two-dimensional semantic segmentation area.
[0057] As attached Figure 3 As shown in the figure, the multimodal channel attention correction module first downsamples the color features and depth features through pooling operations of different sizes, flattens the pooling results and splices them along the second dimension to obtain multi-scale color features and multi-scale depth features, and then maps the dimensions to higher dimensions through a fully connected layer to obtain vectorized color features and depth features. The vectorized color features and depth features are respectively added with learnable modal codes and then passed into the Transformer module for global attention modeling. The calculation process of each Transformer can be expressed by the following formula:
[0058] z f =CMMHSA(LN(z l-1 ))+z l-1
[0059] z l =MLP(LN(z l ))+z l
[0060] In the formula, z l represents the input of the lth module, LN represents layer normalization, CMMHSA represents self-attention multimodal information interaction module, and MLP represents multi-layer perceptron;
[0061] Finally, the modeled results are calculated through a multi-layer perceptron to obtain the channel attention vectors of the two modalities. After the channel attention vectors of the two modalities and the features of the two modalities are channel-multiplied, the corresponding elements are added and fused with the features of the other modalities to achieve multi-scale channel attention correction. The above operation is expressed by the following formula:
[0062] RGB msf =Concat(Flatten(msap(RGB in )), Flatten(msmp(RGBin )))
[0063] Depth msf =Concat(Flatten(msap(Depth in )), Flatten(msmp(Depth in )))
[0064] RGB tokenized , Depth tokenized =FC(RGB msf , Depth msf )
[0065]
[0066] W rgb , W depth =MLP(RGB cii , Depth cii )
[0067]
[0068]
[0069] In the formula, RGB in and Depth in Represents color features and depth features respectively, msmp and msap represent multi-scale maximum pooling and multi-scale average pooling respectively, Concat represents a merging operation, Flatten represents a flattening operation, RGB msf and Depth msf Represents multi-scale color features and multi-scale depth features, RGB tokenized and Depth tokenized Represents vectorized color features and depth features, RGB cii and Depth cii They represent the color features and depth features after channel information interaction, MTE rgb and MTE depth Represent the modal encoding of color features and depth features, W rgb and W depth Represent the channel attention vectors of the two modalities respectively, MLP represents the multi-layer perceptron, and Respectively represent the addition of corresponding elements and the multiplication of corresponding channels, RGB rec and Depth rec They represent the color features and depth features after multi-scale attention correction respectively;
[0070] As attached Figure 4As shown, the multimodal global feature aggregation module first embeds position information and modal information in the feature map, introduces position information through a depth-separable convolution with a convolution kernel size of 3×3, a step size of 1×1, and a padding size of 1×1, and then adds it to the input features through residual connections. In addition to position encoding, learnable modal encoding is also added to obtain color features and depth features with position information and modal information. Then, the self-attention multimodal information interaction module described in claim 1 is used to interact with the spatial dimension and perform residual connections with the input. At the same time, a spatial reduction module is introduced to reduce the amount of calculation through a sharing mechanism of key vectors and value vectors. Then, the color features and depth features after spatial information interaction are obtained through layer normalization. Finally, a 1×1 convolution is used to fuse the feature maps of the two modalities into a single feature map. In addition, in order to improve the robustness of the model, the original feature map is obtained through a depth-separable convolution of 3×3 to obtain local features, and is fused with the global features through residual connections, and then the final output is obtained through a batch normalization layer. The above calculation process can be expressed by the formula:
[0071]
[0072]
[0073]
[0074] F global =Conv 1×1 (Concat(RGB sii , Depth sii ))
[0075]
[0076]
[0077] In the formula, pme represents position encoding and modality encoding, DWC 3×3 Represents a depth-separable convolution with a kernel size of 3×3, RGB pme and Depth pme represents the color features and depth features after position encoding and modality encoding, SR represents the space reduction module, RGB sii and Depth sii They represent the color features and depth features after spatial information interaction, respectively. global Represents the global feature, Relu represents the Relu activation function, F local represents local features, BN represents batch normalization layer, F out Represents the final output.
[0078] S3. Perform RGB-D multimodal semantic segmentation model training: Input the color image and depth image into the RGB-D multimodal semantic segmentation model, and perform end-to-end training with the semantic segmentation label to select the best model. Finally, input the test data into the best model to obtain the final segmentation effect. The effect is shown in the attached figure. Figure 5-7 shown.
[0079] The entire cross-modal semantic segmentation is fully described as follows:
[0080] Step 1: Fix the resolution of color and depth images to 640×480, and perform data enhancement methods such as flipping, cropping, and scaling;
[0081] Step 2: Input the enhanced image into the two-stream pvt_v2 backbone network to extract color features and depth features respectively, and obtain feature maps of different sizes after each downsampling;
[0082] Step 3: The color features and depth features obtained after each downsampling are passed into the multimodal channel attention correction module, which performs feature correction on the color features and depth features in the channel dimension; the corrected features are then passed into the multimodal global feature aggregation module, which performs feature aggregation on the corrected color features and depth features in the spatial dimension; the aggregated features are input into the feature pyramid decoder module to realize the prediction of the two-dimensional semantic segmentation area.
Claims
1. RGB-D multimodal semantic segmentation method, characterized in that: The steps include: S1. Design a self-attention multimodal information interaction module: The self-attention multimodal information interaction module is mainly used to realize multimodal information interaction between color images and depth images in two dimensions: channel dimension and spatial dimension; S2. Establish an RGB-D multimodal semantic segmentation model: The RGB-D multimodal semantic segmentation model includes a dual-stream feature extraction backbone network, a multimodal channel attention correction module, a multimodal global feature aggregation module and a feature pyramid decoder module; the dual-stream feature extraction backbone network is used to extract features of the color image and the depth image respectively, generate feature images of different sizes, and pass the features into four multimodal channel attention correction modules after each downsampling; the multimodal channel attention correction module performs feature correction on the color features and depth features extracted by the dual-stream feature extraction backbone network in the channel dimension, and passes the corrected features into the multimodal global feature aggregation module, first downsamples the color features and depth features through pooling operations of different sizes, flattens the pooling results and splices them along the second dimension to obtain multi-scale color features and multi-scale depth features, and then maps the dimensions to higher dimensions through a fully connected layer to obtain vectorized color features and depth features, and adds learnable modal encoding to the vectorized color features and depth features and passes them into the Transformer module for global attention modeling; the multi The modal global feature aggregation module performs feature aggregation on the corrected color features and depth features output by the multimodal channel attention correction module in the spatial dimension. First, the position information and modality information are embedded in the feature map, and the position information is introduced through a depth-wise separable convolution with a convolution kernel size of 3×3, a step size of 1×1, and a padding size of 1×1. Then, the position information is added to the input features through a residual connection. In addition to the position encoding, a learnable modality encoding is also added to obtain color features and depth features with position information and modality information. Then, the self-attention multimodal information is exchanged. The interaction module performs information interaction in the spatial dimension and performs residual connection with the input. At the same time, the spatial reduction module is introduced to reduce the amount of calculation through the sharing mechanism of key vectors and value vectors. Then, the color features and depth features after spatial information interaction are obtained through layer normalization. Finally, a 1×1 convolution is used to fuse the feature maps of the two modalities into a single feature map. In addition, in order to improve the robustness of the model, the original feature map is obtained through a 3×3 depth-separable convolution to obtain local features, and then fused with the global features through residual connection, and the final output is obtained through the batch normalization layer; S3. Perform RGB-D multimodal semantic segmentation model training: input the color image, depth image and semantic segmentation true label into the RGB-D multimodal semantic segmentation model for training to obtain a trained RGB-D multimodal semantic segmentation model.
2. The RGB-D multimodal semantic segmentation method according to claim 1, characterized in that: The self-attention multimodal information interaction module is improved for multimodal data based on the original multi-head attention module; The self-attention multimodal information interaction module first calculates the query vector, key vector and value vector for the input color features and depth features respectively, wherein the acquisition of the three vectors is completed by the fully connected layer, and then the query vectors of the two modalities are exchanged, and the query vector of one modality and the transpose of the key vector of the other modality are used for matrix multiplication to calculate the self-attention matrix of each modality, and then the obtained self-attention matrix is matrix multiplied with the value vector to obtain the color features and depth features after information interaction, and finally the final output result is obtained through a fully connected layer to realize cross-modal information interaction. The above operation is expressed by the following formula: ; In the formula, Represent the query vector, key vector and value vector of the color feature respectively, Represent the query vector, key vector and value vector of the deep feature respectively, represents the dimension of the vector, express Activation function, and Represent the self-attention matrices of color features and depth features respectively, represents the fully connected layer, and They respectively represent the color features and depth features after information interaction.
3. The RGB-D multimodal semantic segmentation model according to claim 1, characterized in that: The RGB-D multimodal semantic segmentation model consists of a two-stream feature extraction backbone network, four multimodal channel attention correction modules, and four multimodal global feature aggregation modules; The calculation process of the dual-stream feature extraction backbone network is expressed by the following formula: ; In the formula, Indicates The input of the module, Representation layer normalization, represents the self-attention multimodal information interaction module, represents a multi-layer perceptron; The calculation process of the multimodal channel attention correction module is expressed by the following formula: ; In the formula, and Represent color features and depth features respectively, and Respectively represent multi-scale maximum pooling and multi-scale average pooling, Represents a merge operation. represents the flattening operation, and Represent multi-scale color features and multi-scale depth features respectively, and Represent the vectorized color features and depth features respectively, and They represent the color features and depth features after the channel information interacts, and Represent the modal encoding of color features and depth features respectively, and Represent the channel attention vectors of the two modalities respectively, represents a multi-layer perceptron, and Respectively represent the addition of corresponding elements and the multiplication of corresponding channels, and They represent the color features and depth features after multi-scale attention correction respectively; The calculation process of the multimodal global feature aggregation module is expressed by the following formula: ; In the formula, represents positional encoding and modal encoding, represents a depth-wise separable convolution with a kernel size of 3×3. and represents the color features and depth features after position encoding and modality encoding, SR represents the space reduction module, and They represent the color features and depth features after the spatial information interaction, Represents global features, express Activation function, represents local features, BN represents batch normalization layer, Represents the final output.