A semantic segmentation method for indoor RGB-D images based on wavelet transform
By adopting a wavelet transform fusion module in indoor RGB-D image semantic segmentation and combining with ResNet-50 skeleton network, the misclassification and robustness problems in indoor scene RGB image semantic segmentation are solved, and higher semantic segmentation accuracy and inference speed are achieved.
Patent Information
- Application Number
- CN202210461039.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-28
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-04-28
AI Technical Summary
The prior art has problems such as category misclassification, edge misclassification, robustness and accuracy in the semantic segmentation of indoor scenes, and it is difficult to effectively integrate RGB color images and depth information.
The indoor RGB-D image semantic segmentation method based on wavelet transform is adopted, and ResNet-50 is used as the skeleton network, and multi-scale fusion is carried out through a module combining wavelet transform and inverse transform at the connection between the encoder and the decoder, retaining low-frequency information and high-frequency contour details.
It improves the accuracy and inference speed of indoor image semantic segmentation, reduces the amount of parameters and calculation, and can better perform semantic segmentation of size and targets.
Smart Images

Figure CN114842216B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to a semantic segmentation method of indoor RGB-D images based on wavelet transform. Background Art
[0002] With the development of computer vision, semantic segmentation of images has become an important topic in this field. Each pixel in the image is classified and its label, position, and shape are predicted. The above operations provide a complete understanding of the relevant scene. Semantic segmentation is now widely used in autonomous driving, remote sensing analysis, medical image processing, etc. For indoor scene semantic segmentation, since the factors affecting the semantic segmentation of indoor scenes are relatively complex (lighting, occlusion, etc.), reducing the influence of these factors in the study of indoor scenes and improving the accuracy of semantic segmentation has become a major problem. Previous deep learning methods achieved end-to-end image semantic segmentation by processing indoor scene RGB images. However, due to the complex indoor scenes, uneven lighting, and high color and texture repetition, the indoor scene semantic segmentation based on RGB images has the characteristics of category misclassification, edge missegmentation, low robustness and accuracy. Recent studies have found that semantic segmentation based on RGB-D images can improve the segmentation effect through depth information in the scene that is less affected by conditions such as lighting. At the same time, depth information can also show the position and other relationships between objects, which is complementary to RGB color images and also has a certain auxiliary effect on semantic segmentation.
[0003] With the emergence of depth cameras such as TOF (Time-of-flight) and Kinect, it has become easier to obtain the depth information of the scene. However, it has always been a challenging problem to find an effective and high-quality way to fuse RGB color images with depth information. Jiang et al. added RGB and depth information of some intermediate layers of the encoder through parallel processing. Li et al. believed that single-layer fusion could not complement the color image and depth map well, and chose to fuse features before the final prediction. Chang et al. added depth information to the loss function and used the change of depth value at the edge of the classified object to constrain network training. However, there are still many problems with the above methods: 1) Simply fusion of depth image with color image as the fourth channel does not fully utilize the complementarity of RGB color information and depth information; 2) There are problems such as loss of multiple scale information features in different methods. However, for the problem of semantic segmentation of indoor scenes, the efficiency of multi-scale feature extraction affects the segmentation accuracy of small target objects. 3) Depth image is sparser than RGB image. It represents the depth value of each pixel in RGB image. Edge contour information, that is, high-frequency information, can be better obtained from depth image, but traditional convolution, pooling and other operations often lose high-frequency features. Therefore, it is particularly important to solve this type of problem. Summary of the invention
[0004] In view of the above problems, the present invention provides a semantic segmentation method for indoor RGB-D images based on wavelet transform, adopts ResNet-50 as the skeleton network and a fusion module based on wavelet transform. The original image is divided into four frequency sub-generations through wavelet transform, and image fusion is realized through inverse wavelet transform, which are used for the fusion of RGB color image and depth image and the context module respectively.
[0005] In order to implement the above technical solution, the present invention provides a method for semantic segmentation of indoor RGB-D images based on wavelet transform, comprising the following steps:
[0006] Step 1: Construct a convolutional neural network based on wavelet transform. The encoder network uses ResNet-50 and removes the fully connected layer. There are two branches in the network that extract RGB and depth feature information from the original image respectively. At the same time, the wavelet transform fusion module is used to perform RGB-D feature fusion.
[0007] Step 2: At the connection between the encoder and the decoder, the image features of different frequencies of the original image are multi-scale fused through a module combining wavelet transform and inverse transform;
[0008] Step 3: The above features are upsampled multiple times through the decoder network. Each module of the decoder upsamples the features twice, and better features are mapped through convolution and skip connections of the encoder, gradually restoring high-resolution images and outputting semantic segmentation results.
[0009] A further improvement is that the training process of the network models such as the encoder and decoder is as follows:
[0010] Step 1: Input the original image and the corresponding label image into the convolutional neural network. The original image includes a three-channel color image and a depth image of the corresponding scene. The label image is the label category corresponding to each pixel in the original image.
[0011] Step 2: A wavelet transform fusion module is added before each residual module in the encoder to fuse the RGB color information and the depth information. The two feature information are first decomposed into four different sub-bands of A, V, H, and D by discrete wavelets, and then the corresponding sub-bands of the RGB color image and the depth image are added, and then passed through the attention module SEModule, and finally fused and output by inverse wavelet transform; the present invention designs a discrete wavelet transform module to replace the pooling operation in ResNet. After the image enters the module, it is decomposed into four sub-bands, and then the feature information resolution obtained by splicing the sub-bands becomes half of the input, and the number of channels is 4 times the input. Then, a 1×1 convolution operation is performed to convert the number of channels into the same channel value as the input;
[0012] Step 3: Set a wavelet transform connection module at the connection between the encoder and decoder. Similar to the fusion module, it retains the different frequency information features of the original input through wavelet transform and fuses multi-scale information at the same time. The input features are decomposed into four sub-bands after discrete wavelet transform, and then fused by inverse wavelet transform through the attention module. At the same time, it is jump-connected with the input branch to aggregate information of different scales.
[0013] Step 4: Integrate depthwise separable convolution to further improve the accuracy and efficiency of semantic segmentation, increase the resolution through nearest neighbor upsampling, and use depthwise separable convolution to integrate feature information.
[0014] The beneficial effects of the present invention are:
[0015] 1. The present invention fuses image information of different frequencies through wavelet transform and inverse wavelet transform, thereby retaining low-frequency information and high-frequency contour details of the image.
[0016] 2. The pooling operation of the original network is replaced by wavelet transform, and a wavelet transform module is added at the connection between the encoder and decoder, which effectively retains the original image detail information, reduces parameters and optimizes the inference speed.
[0017] 3. The present invention can reduce the number of parameters and the amount of calculation while speeding up the reasoning speed, retaining the detailed contour information of the image, and can perform semantic segmentation of large and small objects well. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a flow chart of the present invention.
[0019] Figure 2 This is a network module designed based on wavelet transform of the present invention, in which the top is the attention module, the left side of the lower part of the figure is the wavelet transform fusion block, the middle part is the wavelet transform module (DWT block), and the right side is the wavelet transform connect module (wavelet transform connect module).
[0020] Figure 3 Schematic diagram of a decoder module of the present invention. DETAILED DESCRIPTION
[0021] In order to deepen the understanding of the present invention, the present invention will be further described in detail below in conjunction with examples. The examples are only used to explain the present invention and do not constitute a limitation on the protection scope of the present invention.
[0022] according to Figure 1 ,2 As shown in , 3, this embodiment provides a semantic segmentation method for indoor RGB-D images based on wavelet transform, which makes full use of high-frequency features that have not been utilized by previous methods, retains information such as contour details in image information, effectively improves the semantic segmentation accuracy and reasoning speed of indoor images, and reduces the number of network parameters.
[0023] First, we build a convolutional neural network based on wavelet transform, using an encoder-decoder structure network, where the encoder uses ResNet-50 as the skeleton network, and uses the designed wavelet fusion module and four convolution modules. We select the efficient and fast Haar wavelet, and use the discrete wavelet transform (DWT) of the 2D Haar wavelet transform to decompose the original image x into four sub-band images. The image size, i.e., the image resolution, becomes half of the original image. The above operation is equivalent to using four filters (f LL 、f LH 、f HL 、f HH ) Decompose the original image x to obtain x LL 、x LH 、x HL x HH The parameters of the filter are fixed for the four sub-band images, namely A low-frequency image, V vertical detail image, H horizontal detail image, and D diagonal detail image. The parameters are not updated by gradient descent operation with the back propagation of network training, and the stride is set to 2. The filter of Haar wavelet is as follows:
[0024]
[0025] The input image is x(i, j), where i is the row and j is the column, then the 2D DWT is shown in Equation 2:
[0026]
[0027] in Represents the convolution operation. The input x can be represented by convolution operations with different filters, which can be understood as downsampling with a stride of 2. Since the wavelet transform does not lose information, the wavelet transform and the inverse wavelet transform are reversible operations. For the Haar wavelet, the inverse wavelet transform can be expressed as Equation 3:
[0028]
[0029] The original image is obtained by DWT to obtain four sub-bands, and the four sub-bands can also be obtained by IWT to obtain the original image. The designed wavelet transform module is used at the codec connection, and the decoder part uses depth separable convolution and nearest neighbor upsampling modules for upsampling, and skip connections are used to prevent information loss. The features of different frequencies are extracted from the image by wavelet transform to ensure the effective use of high-frequency information;
[0030] The established wavelet transform-based convolutional neural network is trained end-to-end, using the Adam optimizer and setting the initial learning rate to 0.002. Each upsampled output in the decoder is summarized and all upsampled outputs and the total output are input into the loss function part. The obtained loss function is back-propagated to update the network parameters. The loss function is shown in formula (4);
[0031]
[0032] Among them, α t is a vector representing the parameters of each subclass; p t represents the predicted value of the current sample on the label image; γ represents the focus parameter, which is set to 2.
[0033]
[0034] In the above formula, p is the value of the feature image generated by the prediction result after softmax processing;
[0035] Input the image to be segmented into the trained network to obtain the semantic segmentation result.
[0036] The encoder uses dual branches to extract feature information of color images and depth images respectively; it contains four convolution modules, and after each convolution module, the color image and the depth image are fused through the designed wavelet transform fusion module; the wavelet transform fusion module decomposes the color image and the depth image into four sub-bands consisting of a low-frequency image, a vertical detail image, a horizontal detail image, and a diagonal detail image through 2D discrete wavelet transform, and adds and fuses the four sub-bands corresponding to the color image and the depth image respectively, and then passes through the attention module, and finally obtains the fused image through the inverse wavelet transform;
[0037] At the connection between the encoder and the decoder, the extracted features are decomposed into four sub-bands through the wavelet transform module, and restored to the original input size through the attention module and the inverse wavelet transform to aggregate the feature information of different frequencies;
[0038] After the above steps, the decoder is entered, and the resolution is increased and feature information is integrated through the integrated separable convolution module and nearest neighbor upsampling. At the same time, the encoder and decoder are skipped using 1×1 convolution to avoid information loss. After three decoder modules and two upsampling modules, the resolution is restored to the original image input size.
[0039] The predicted feature image obtained in the above steps is input into the prediction layer to obtain the output result.
[0040] In this embodiment, the training process of the network models such as the encoder and decoder is as follows:
[0041] Step 1: Input the original color image, depth image and corresponding labels into the network;
[0042] Step 2: If Figure 1 As shown in the figure, a wavelet transform fusion module is added before each residual module in the encoder to fuse the RGB color information and the depth information. The two feature information are first decomposed into four different sub-bands of A, V, H, and D by discrete wavelet, and then the corresponding sub-bands of the RGB color image and the depth image are added, and then passed through the attention module SEModule, and finally fused and output by the inverse wavelet transform. Figure 2 As shown, the present invention designs a discrete wavelet transform module to replace the pooling operation in ResNet. After the image enters the module, it is decomposed into four sub-bands, and then the resolution of the feature information obtained by splicing the sub-bands becomes half of that of the input, and the number of channels is 4 times that of the input. Then, a 1×1 convolution operation is performed to convert the number of channels into the same channel value as the input.
[0043] Step 3: The present invention designs a wavelet transform connection module at the connection between the encoder and decoder. Similar to the fusion module, we hope to retain the different frequency information characteristics of the original input through wavelet transform, and fuse the multi-scale information at the same time. The input features are decomposed into four sub-bands (A, V, H, D) after discrete wavelet transform, and then fused by inverse wavelet transform through the attention module, and at the same time jump-connected with the input branch to aggregate information of different scales. Step 4: The decoder part is as follows Figure 3As shown. The decoder of the present invention is composed of three decoder modules, in which the number of channels of the decoder module is reduced from 2048 through convolution operations as the resolution increases. In addition, depth-separable convolution is integrated. Depth-separable convolution can reduce the amount of parameters and save computational costs. These operations can further improve the accuracy and efficiency of semantic segmentation. Finally, the resolution is increased by nearest neighbor upsampling, and feature information is integrated by depth-separable convolution. It is difficult to avoid information loss in the process of decoder upsampling. We use 1×1 convolution operations to integrate the encoder RGB-D fused features through the decoder skip connection. After the three decoder modules, the image resolution is restored through two upsampling modules. At the same time, we add an output after each decoder module. The outputs of different resolutions and the final results are input to the end of the network and finally the Loss function is obtained. The Loss function selects the cross entropy function, as shown in the following formula:
[0044]
[0045] Among them, class is the label category at pixel i, andx represents the model output at pixel i, and N represents the spatial resolution of the specific output. Since there are three outputs of the decoder module, the total loss function is the sum of the four loss functions, and since the resolutions of different outputs are different, we assign different weights according to the size of the resolution, with a ratio of 1:2:3:4.
[0046] Experiments were conducted on two public datasets, NYUV2 and SUN RGB-D, using RGB color images and depth images as input. The entire model was trained on AMD Ryzen9, using the Adam optimizer to dynamically adjust the learning rate, with the initial learning rate set to 0.002. At the same time, we used transfer learning and used the ResNet50 pre-trained weights on the imagenet dataset as the initial weights, which effectively accelerated the training speed of the entire network.
[0047] Three semantic segmentation evaluation indicators are used to evaluate the experimental results of this paper, including PA, MPA and MIoU. PA is the ratio of the number of correctly classified pixels to the total number of pixels in an image, and is defined as follows:
[0048]
[0049] where p ii Indicates the number of pixels classified correctly, p ij It represents the number of pixels that belong to class i but are predicted to be class j.
[0050] MPA is the average value of the ratio of the number of correctly classified pixels in each class to the number of all pixels, and is defined as follows:
[0051]
[0052] Where k represents the number of categories.
[0053] MIoU is the average value of the ratio of the intersection and union of the two sets of true values and predicted values, and is defined as follows:
[0054]
[0055] NYUv2 and SUNRGBD are introduced into the proposed algorithm and compared with the existing algorithms. The results show that this paper outperforms the existing algorithms in terms of pixel accuracy, average pixel accuracy and average intersection-over-union ratio. All three evaluation indicators have been improved. We believe that this is due to the RGB-D fusion combined with wavelet transform, which retains more detail information and enables a better fusion effect between the two. At the same time, although only the ResNet50 encoding structure is used instead of a deeper encoding network, the edge information of the target object is more accurately judged by replacing the pooling layer and designing the encoder-decoder wavelet transform connection module. At the same time, the entire network does not need a very deep encoder structure to obtain a good segmentation effect.
[0056]
[0057] At the same time, the parameter size and inference speed of the model are tested on the NYUv2 dataset. The entire experiment is deployed on NVIDIA 1080Ti and compared with existing methods. From the results, it can be seen that the wavelet transform module we designed only requires a small amount of additional calculations and can achieve real-time inference speed.
[0058]
[0059] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A semantic segmentation method for indoor RGB-D images based on wavelet transform, It is characterized in that The following steps are involved: Step 1: Construct a convolutional neural network based on wavelet transform. The encoder network uses ResNet-50 and removes the fully connected layer. There are two branches in the encoder network that extract RGB and depth feature information from the original image respectively. At the same time, the wavelet transform fusion module is used to perform RGB-D feature fusion. Step 2: At the connection between the encoder and the decoder, the image features of different frequencies of the original image are multi-scale fused through a module combining wavelet transform and inverse transform; Step 3: The above features are upsampled multiple times through the decoder network. Each module of the decoder upsamples the features twice, and better features are mapped through convolution and skip connections of the encoder, gradually restoring high-resolution images and outputting semantic segmentation results.
2. According to the indoor RGB-D image semantic segmentation method based on wavelet transform according to claim 1, It is characterized in that The training process of the encoder and decoder network model is as follows: Step 1: Input the original image and the corresponding label image into the convolutional neural network. The original image includes a three-channel color image and a depth image of the corresponding scene. The label image is the label category corresponding to each pixel in the original image. Step 2: A wavelet transform fusion module is added before each residual module in the encoder to fuse the RGB color information and the depth information. The two feature information are first decomposed into four different sub-bands of A, V, H, and D by discrete wavelet, and then the corresponding sub-bands of the RGB color image and the depth image are added, and then passed through the attention module SEModule, and finally fused and output by inverse wavelet transform; Step 3: Set a wavelet transform connection module at the connection between the encoder and the decoder. The module is the same as the fusion module. It retains the different frequency information characteristics of the original input through wavelet transform and fuses the multi-scale information at the same time. The input features are decomposed into four sub-bands after discrete wavelet transform, and then fused by inverse wavelet transform through the attention module, while skip-connecting with the input branch to aggregate information of different scales; Step 4: Integrate depthwise separable convolution to further improve the accuracy and efficiency of semantic segmentation, increase the resolution through nearest neighbor upsampling, and use depthwise separable convolution to integrate feature information.
Citation Information
Patent Citations
Image semantic segmentation based on multi-level feature fusion and Gaussian conditional random field
CN109461157A
Image segmentation method and system based on multi-encoder convolutional neural network
CN111915612A