RGB-d image semantic segmentation method and device based on cross-data view collaborative training and computer readable medium
Through the RGB-D semantic segmentation method of collaborative training across data views, the point cloud semantic segmentation model is used to transfer 3D spatial relationships and combined with the deep spatial adapter to calibrate RGB features, which solves the problems of insufficient utilization of depth information and RGB feature noise, realizes accurate segmentation and recognition of complex indoor environments, and improves the operational capabilities of domestic robots.
Patent Information
- Application Number
- CN202510935266.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-07-08
AI Technical Summary
Existing RGB-D semantic segmentation methods cannot fully utilize depth information, resulting in segmentation errors for household robots in complex indoor environments. For example, carpet wrinkles are misidentified as obstacles or objects with similar textures are identified as similar, affecting cleaning and operation efficiency.
A cross-data view collaborative training method is adopted to transfer 3D spatial relationships through the point cloud semantic segmentation model. The RGB features are calibrated with the deep spatial adapter module to construct an RGB-D semantic segmentation model. The distillation loss and structured loss are used for collaborative training to fully utilize the depth information and accurately extract the RGB features.
It improves the accuracy of RGB-D image segmentation, helps domestic robots identify object edges and objects with similar textures in complex environments, and improves the accuracy and efficiency of cleaning and operation.
Smart Images

Figure CN120431583B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, in particular to the field of multispectral semantic segmentation, and relates to a method, device and computer-readable medium for RGB-D image semantic segmentation based on collaborative training across data views. Background Art
[0002] With the rapid development of artificial intelligence and robotics, intelligent household robots are gradually showing broad application prospects in the field of service robots. These robots are designed to complete household tasks such as cleaning, object identification and handling, and human-computer interaction, placing high demands on environmental perception and understanding capabilities. To enable robots to have stronger environmental modeling and manipulation capabilities, RGB-D semantic segmentation technology has become one of the key research and engineering implementation directions in recent years. Generally speaking, traditional semantic segmentation methods rely on stable lighting conditions and have difficulty coping with complex indoor environments. The models are often limited when dealing with objects that are occluded, have changing lighting conditions, or have unclear textures. For example, multiple objects placed in a closet and multiple pillows with similar patterns on a sofa are difficult for traditional semantic segmentation methods to accurately segment objects. To address the above challenges, RGB-D semantic segmentation technology has emerged as a key method for scene understanding that combines color information (RGB) and depth information (Depth). By fusing the texture and color features in RGB images with the geometric and spatial structure information in depth maps, RGB-D semantic segmentation can achieve fine segmentation and recognition of elements in the domestic environment such as the floor, furniture, appliances, walls, and people, thereby providing a solid perception foundation for subsequent path planning, object manipulation, and interactive decision-making.
[0003] Compared to traditional semantic segmentation methods, the core of RGB-D semantic segmentation lies in the utilization of depth information. Generally, these methods can be divided into two categories based on how they utilize depth information: 1) RGB-D semantic segmentation methods based on geometric priors; 2) RGB-D semantic segmentation methods based on feature fusion. RGB-D semantic segmentation methods based on geometric priors view depth information as a representation of geometric relationships in RGB images and use the depth prior to guide RGB feature extraction. For example, ShapeConv employs a shape-aware convolutional algorithm that decomposes depth features into a shape component and a basis component. It then introduces two learnable weights to independently coordinate these components. Finally, the reweighted combination of these two components is applied to the convolution kernel weights to effectively encode local geometric shapes. PDCNet, on the other hand, believes that pixels with more consistent geometric shapes should contribute more to the corresponding output. Building on this perspective, it proposes a pixel difference convolutional (PDC) module. This module considers local and detailed geometric information in depth data by aggregating intensity and gradient information to enhance feature representation and model generalization. RGB-D semantic segmentation methods based on feature fusion treat depth information as a distinct modality. They use a two-stream feature extractor in a non-parametric shared fashion to extract image features using either a convolutional neural network (CNN) or a transformer. To effectively fuse multimodal features, RGB-D semantic segmentation methods based on feature fusion employ a fusion module at each feature extraction stage to fuse cross-modal features. This module typically consists of an attention mechanism and a gating mechanism. For example, CMXNet uses cross-channel attention to promote cross-modal information interaction and combination, allowing for sufficient long-range context exchange before modal fusion. HDBFormer, on the other hand, proposes a Modality Information Interaction Module (MIIM), which employs a targeted fusion strategy to address the complexity difference between RGB and depth features. Furthermore, it introduces a hierarchical feature processing mechanism to explicitly divide input features into primary and secondary features to prevent the accumulation of redundant information. It also utilizes a transformer combined with large-kernel convolution to effectively integrate global and local information across modalities.
[0004] The shortcomings of the above methods are mainly reflected in the following aspects: (1) Existing RGB-D semantic segmentation methods cannot fully utilize depth information. Segmentation methods based on geometric priors often use artificial prior knowledge and therefore cannot be flexibly adjusted during the model learning process. Segmentation methods based on feature fusion are difficult to learn the potential 3D spatial position relationship and local geometric information in depth information by fusing in 2D space; (2) Existing RGB-D segmentation uses independent feature extractors to extract RGB and depth modes respectively. However, in complex scenes, RGB views are prone to confuse objects with similar textures, thereby generating noise in the RGB feature extraction stage, affecting subsequent fusion and further affecting segmentation performance. For example, insufficient utilization of depth information will cause a housekeeping robot to misjudge carpet wrinkles as obstacles and frequently detour, reducing cleaning efficiency. The RGB noise problem will cause the robot to recognize "marble floor" and "white plastic table" as the same type (similar texture), and the depth branch is not effectively corrected due to noise, resulting in segmentation errors, causing the robot to try to mop the table.
[0005] Based on the above considerations, for practical application scenarios such as smart cameras and smart household robots in complex indoor scenes, there is an urgent need to design an image semantic segmentation method that can fully utilize depth information and accurately extract RGB features. Summary of the Invention
[0006] The purpose of the present invention is to address the shortcomings of the existing technology and to develop an RGB-D image semantic segmentation method based on collaborative training across data views. The method can not only collaboratively train and learn the 3D spatial position relationship and local geometric information in the point cloud semantic segmentation model to fully utilize the depth information, but also introduce geometric depth information through the depth space adapter to improve and enhance the RGB feature representation, thereby achieving accurate segmentation of RGB-D images.
[0007] The present invention provides an RGB-D image semantic segmentation method based on cross-data view collaborative training, comprising the following steps:
[0008] Step (1) obtain a dataset containing RGB images, depth images and real labels;
[0009] Convert the RGB image and depth image into point cloud data, construct and input a point cloud feature extractor to obtain four layers of point cloud features; the RGB image and depth image are both indoor scene images;
[0010] Construct a hierarchical upsampling decoder with four-layer point cloud features as input and point-by-point semantic prediction results as output. Use a two-dimensional projection algorithm to convert the point-by-point semantic prediction results and four-layer point cloud features into a point cloud semantic prediction map and a four-layer point cloud projection feature map.
[0011] Step (2) constructing a depth feature extractor, inputting the depth image to obtain a four-layer depth feature map;
[0012] Step (3) constructing an RGB feature extractor, wherein the RGB feature extractor consists of a four-layer RGB feature extraction submodule and a depth-guided spatial adapter module;
[0013] Inputting the RGB image and the four-layer depth feature map into the RGB feature extractor to obtain a four-layer RGB feature map and a four-layer adapter RGB feature map;
[0014] Step (4) construct a four-layer RGB-D feature fusion module, the input is a four-layer depth feature map and a four-layer RGB feature map, and the output is a four-layer RGB-D fused feature;
[0015] A multi-layer perceptron is used to construct an RGB-D decoder, which takes as input the fused features of four layers of RGB-D and outputs a semantic prediction map of the RGB-D image.
[0016] Step (5) calculates the cross entropy loss between the true label and the point cloud semantic prediction map, calculates the cross entropy loss between the true label and the RGB-D image semantic prediction map, calculates the distillation loss and structured loss between the point cloud semantic prediction map and the RGB-D image semantic prediction map for collaborative training, calculates the mean square error loss between the adapter RGB feature map and its corresponding point cloud projection feature map layer by layer at the feature level, and trains the model constructed by steps (1) to (4).
[0017] Preferably, in step (1), the data set further includes camera parameters; based on the camera parameters, the RGB image and the depth image are converted into point cloud data.
[0018] Preferably, the method further comprises the following steps:
[0019] The model constructed by steps (1) to (4) specifically includes: an RGB-D semantic segmentation model composed of an RGB feature extractor, a depth feature extractor, a depth-guided spatial adapter module, a four-layer RGB-D feature fusion module, and an RGB-D decoder; and a point cloud semantic segmentation model composed of a point cloud feature extractor and a hierarchical upsampling decoder;
[0020] After the training, a trained RGB-D semantic segmentation model is obtained;
[0021] Obtain the RGB-D image to be predicted and input it into the trained RGB-D semantic segmentation model to obtain the predicted pixel-level semantic category probability. Select the semantic category with the highest probability as the final pixel-level prediction result.
[0022] Preferably, the step (1) specifically includes the following steps:
[0023] (1-1) Convert the RGB image and its corresponding depth image to obtain point cloud data;
[0024] (1-2) constructing a point transformer (PT) as the backbone network of the point cloud feature extractor to extract features from the point cloud data. The point cloud feature extractor consists of four layers of feature extraction submodules. Except for the first layer of feature extraction submodule that takes point cloud data as input, the remaining three layers of feature extraction submodules take the output of the previous layer of feature extraction submodule as input to obtain four layers of point cloud features.
[0025] (1-3) Construct a hierarchical upsampling decoder as a point cloud feature decoder. The hierarchical upsampling decoder is composed of a multi-layer perceptron and a skip connection, and uses an interpolation algorithm to upsample the four-layer point cloud features obtained in step (1-2) layer by layer. The hierarchical upsampling decoder takes the four-layer point cloud features as input and outputs a point-by-point semantic prediction result.
[0026] (1-4) Use the interpolation algorithm to restore the four-layer point cloud features to the same resolution size, and use the two-dimensional projection algorithm to project the point-by-point semantic prediction results and the four-layer point cloud features into two-dimensional space to obtain the four-layer point cloud projection feature map and the point cloud semantic prediction map.
[0027] Preferably, in step (2), constructing a deep feature extractor includes:
[0028] A vision transformer (ViT) is constructed as the backbone network of the deep feature extractor for feature extraction of depth images. The backbone network of the deep feature extractor consists of four layers of deep feature extraction sub-modules. Except for the first layer of deep feature extraction sub-module which takes the depth image as input, the other three layers of deep feature extraction sub-modules take the output of the previous layer of deep feature extraction sub-module as input to obtain four layers of deep feature maps. Each layer of depth feature map is downsampled twice based on the previous layer.
[0029] Preferably, the step (3) of constructing an RGB feature extractor includes:
[0030] This step constructs a visual converter and a depth-guided spatial adapter module to extract more robust RGB feature representations, thereby better enabling perception of indoor scenes. For example, it helps domestic robots recognize sheets and quilts with similar textures.
[0031] Construct a Vision Transformer (ViT) and a four-layer depth-guided spatial adapter as the backbone network of the RGB feature extractor to extract features from RGB images.
[0032] The Vision Transformer (ViT) includes four layers of RGB feature extraction submodules;
[0033] In the 4-layer RGB feature extraction submodule and the 4-layer depth-guided spatial adapter, except for the first layer RGB feature extraction submodule which takes the RGB image as input, the remaining RGB feature extraction submodules and depth-guided spatial adapter modules take the sum of the outputs of the two modules in the previous layer as the input of this layer; the output of each layer RGB feature extraction submodule is: RGB semantic feature map;
[0034] Except for the first layer of depth-guided spatial adapter module, which takes the RGB image and the first layer of depth feature map obtained in step (2) as input, the depth-guided spatial adapter modules of the remaining layers take the RGB feature map and depth feature map obtained by the RGB feature extraction submodule and the depth-guided spatial adapter module of the previous layer as input.
[0035] The depth-guided spatial adapter module consists of a depth-space calibration submodule and a linear gating submodule connected in series;
[0036] The depth space calibration submodule includes three branch structures of different scales; each branch structure consists of two layers of depth convolution;
[0037] The depth space calibration submodule also includes a depth convolution module, which is used to process the RGB feature map obtained in the previous layer after layer normalization, and then perform Hadamard product operations on the output with the outputs of the branch structures of three different scales and add them together.
[0038] Preferably, step (4) specifically includes the following steps:
[0039] (4-1) Construct a four-layer RGB-D feature fusion module, which consists of a cross-attention layer and a multi-layer perceptron.
[0040] Each layer's RGB-D feature fusion module takes the feature maps extracted by the corresponding layer's RGB feature extractor and depth feature extractor as input, processes them through a cross-attention layer and a multi-layer perceptron, and generates a fused feature map. This step fuses complementary information within the modality through the cross-attention layer and the multi-layer perceptron. For example, a housekeeping robot may have difficulty seeing trash in a closet in low light, but it can identify the specific location of the trash using depth information.
[0041] Use a multi-layer perceptron to build an RGB-D decoder. The input is the fused feature map. The four-layer fused features are spliced along the channel and the multi-layer perceptron is used to obtain the RGB-D semantic prediction map.
[0042] Preferably, step (5) specifically includes the following steps:
[0043] Calculate the cross entropy loss between the true label and the point cloud semantic prediction map , calculate the cross entropy loss between the true label and the RGB-D image semantic prediction map The distillation loss and structured loss are calculated between the point cloud semantic prediction map and the RGB-D image semantic prediction map. At the feature level, the mean squared error loss between the adapter RGB feature map and its corresponding point cloud projection feature map is calculated layer by layer. This constructs a total loss and performs collaborative training. This step, through the collaborative training of the point cloud semantic segmentation model and the RGB-D semantic segmentation model, transfers the point cloud semantic segmentation model's understanding of 3D spatial relationships to the RGB-D semantic segmentation model. This helps the domestic robot better utilize depth information and identify object edges in complex environments, guiding it to perform complex functions such as grasping, for example, grabbing a specific screw from a stack of screws.
[0044] The distillation loss consists of the distillation loss from the point cloud semantic prediction map to the RGB-D image semantic prediction map and the distillation loss from the RGB-D image semantic prediction map to the point cloud semantic prediction map. The difference between the point cloud semantic prediction map and the RGB-D image semantic prediction map is measured by relative entropy, and the distillation temperature is used to smooth the probability distribution to obtain the difference between the point cloud semantic prediction map and the RGB-D image semantic prediction map. The difference between RGB-D image semantic prediction map and point cloud semantic prediction map .
[0045] The calculation of the structured loss includes: first, predicting the semantic map of the RGB-D image and combining the empirical hyperparameters The RGB-D segmentation boundary is obtained, and then the gradient of the local probability distribution is obtained through the point cloud semantic prediction map and the RGB-D image semantic prediction map, and the probability distribution from the point cloud semantic prediction map to the boundary direction and the probability distribution from the RGB-D image semantic prediction map to the boundary direction are calculated; then the relative entropy between the probability distribution from the RGB-D image semantic prediction map to the boundary direction at the RGB-D segmentation boundary and the probability distribution from the point cloud semantic prediction map to the boundary direction at the corresponding spatial position is obtained. , thereby achieving the transfer of structured knowledge;
[0046] Upsample the adapter RGB feature map to unify the resolution size, and calculate the mean square error loss between the adapter RGB feature map and its corresponding point cloud projection feature map layer by layer at the feature level ,
[0047] Calculate the final loss function of the point cloud semantic segmentation model , calculate the final loss function of the RGB-D semantic segmentation model .
[0048] The application further provides a computer readable medium storing a computer program, which can be executed by a processor to implement steps included in the RGB-D image semantic segmentation method based on cross-data view collaborative training.
[0049] The application further provides an RGB-D image semantic segmentation device based on cross-data view collaborative training, comprising a processor and a memory, and the memory stores a computer program which is executed by the processor to implement steps included in the RGB-D image semantic segmentation method based on cross-data view collaborative training.
[0050] The application provides an RGB-D semantic segmentation method based on cross-data view collaborative training, which has the following characteristics: 1) using a cross-data view collaborative training method, for the problem of insufficient utilization of depth information, using distillation loss and spatial structure loss to transfer the understanding of 3D spatial relationship from the point cloud semantic segmentation model to assist the RGB-D segmentation model to fully exploit and utilize depth information; 2) using a depth-guided spatial adapter module, for the noise problem in the RGB feature extraction stage, recalibrating the RGB features through the geometric information in the depth features to achieve more accurate semantic information extraction.
[0051] The application has the following advantages for the problems of insufficient utilization of depth information and noise in RGB feature extraction in RGB-D semantic segmentation: 1) for the problem that it is difficult to learn the potential 3D spatial position relationship and local geometric information in depth information in the fusion in two-dimensional space, a cross-data view collaborative training method is proposed to use distillation loss and spatial structure loss to transfer the understanding of 3D spatial relationship from the point cloud semantic segmentation model, which solves the problem of insufficient utilization of depth information without increasing additional inference overhead; 2) for the noise problem in the RGB feature extraction process, a depth-guided spatial adapter module is proposed to realize more accurate semantic information extraction under the guidance of the geometric information in the depth features, which effectively improves the model segmentation accuracy; this lays a foundation for practical application scenarios such as intelligent cameras and intelligent home robots in complex indoor scenes. BRIEF DESCRIPTION OF DRAWINGS
[0052] Figure 1 is a flowchart of the method of the application.
[0053] Figure 2 is a schematic diagram of an RGB-D image semantic segmentation device based on cross-data view collaborative training. DETAILED DESCRIPTION
[0054] Embodiment 1:
[0055] like Figure 1 , an RGB-D image semantic segmentation method based on cross-data view collaborative training, this method first converts a given RGB-D image into point cloud data, and performs data preprocessing on RGB-D and point cloud data. Construct a point cloud feature extractor and a hierarchical upsampling decoder, use the point cloud feature extractor for feature extraction, and after extraction, the point cloud features are decoded layer by layer through the decoder, and a two-dimensional projection algorithm is used to obtain a point cloud semantic prediction map and a four-layer point cloud projection feature map; use the visual converter to construct a deep feature extractor to obtain a four-layer deep feature map; use the visual converter and the depth-guided spatial adapter module to construct an RGB feature extractor, and the depth-guided spatial adapter module uses the deep feature map to recalibrate the RGB features obtained by the feature extraction submodule to obtain a four-layer RGB feature map; construct a feature fusion module and an RGB-D decoder, and obtain RGB-D after feature fusion and decoding. The method calculates the cross-entropy loss with the ground-truth labels, the distillation loss, and the structured loss for co-training. At the feature level, the mean squared error loss is calculated between the adapter RGB feature map and its corresponding point cloud projection feature map. The method uses the stochastic gradient descent algorithm and the aforementioned losses to optimize the RGB-D semantic segmentation model, consisting of an RGB feature extractor, a deep feature extractor, a depth-guided spatial adapter module, a four-layer RGB-D feature fusion module, and an RGB-D decoder, as well as the point cloud semantic segmentation model, consisting of a point cloud feature extractor and a hierarchical upsampling decoder. The optimized RGB-D semantic segmentation model is then used to generate a semantic segmentation map for a new RGB-D image. This method utilizes cross-data-view co-training to transfer the 3D spatial relationship understanding of the point cloud semantic segmentation model, helping the RGB-D semantic segmentation model fully utilize depth information. The depth-guided spatial adapter module recalibrates RGB features using geometric information from the depth features, achieving more accurate semantic information extraction.
[0056] The following is a detailed description of each step
[0057] The method of the present invention performs the following operations in sequence on a given image data set containing camera parameters, RGB images, depth images and real labels: the RGB images and depth images are both indoor scene images;
[0058] Step (1) Use the camera parameters of the dataset to convert the given RGB image and depth image into point cloud data. After data preprocessing, construct and input the point cloud feature extractor to obtain four layers of point cloud features.
[0059] Construct a hierarchical upsampling decoder with four-layer point cloud features as input and point-by-point semantic prediction results as output. Use a two-dimensional projection algorithm to convert the point-by-point semantic prediction results and four-layer point cloud features into a point cloud semantic prediction map and a four-layer point cloud projection feature map.
[0060] Step (2) obtains the depth image in the dataset, constructs and inputs it into the depth feature extractor after preprocessing, and outputs a four-layer depth feature map;
[0061] Step (3) obtaining the RGB images in the dataset, preprocessing them, and constructing an RGB feature extractor, wherein the RGB feature extractor consists of a four-layer RGB feature extraction submodule and a depth-guided spatial adapter module;
[0062] Inputting the RGB image and the four-layer depth feature map into the RGB feature extractor to obtain a four-layer RGB feature map and a four-layer adapter RGB feature map;
[0063] Step (4) construct a four-layer RGB-D feature fusion module, the input is a four-layer depth feature map and a four-layer RGB feature map, and the output is a four-layer RGB-D fused feature;
[0064] A multi-layer perceptron is used to construct an RGB-D decoder, which takes as input the fused features of four layers of RGB-D and outputs a semantic prediction map of the RGB-D image.
[0065] Step (5) calculates the cross entropy loss between the true label and the point cloud semantic prediction map, calculates the cross entropy loss between the true label and the RGB-D image semantic prediction map, calculates the distillation loss and structured loss between the point cloud semantic prediction map and the RGB-D image semantic prediction map for collaborative training, and calculates the mean square error loss between the adapter RGB feature map and its corresponding point cloud projection feature map layer by layer at the feature level.
[0066] Step (6) using the stochastic gradient descent algorithm to optimize the RGB-D semantic segmentation model consisting of the RGB feature extractor, the depth feature extractor, the depth-guided spatial adapter module, the four-layer RGB-D feature fusion module, and the RGB-D decoder, and the point cloud semantic segmentation model consisting of the point cloud feature extractor and the hierarchical upsampling decoder, and repeating steps (1) to (5) until the model converges to obtain the trained RGB-D semantic segmentation model;
[0067] Step (7) For the new RGB-D image, input it into the trained RGB-D semantic segmentation model to obtain the predicted pixel-level semantic category probability, and select the semantic category with the highest probability as the final pixel-level prediction result.
[0068] Furthermore, step (1) is specifically:
[0069] (1-1) For RGB images And its corresponding depth image ,in and respectively, and 3 represents RGB three channels. The RGB-D image is converted into point cloud data using a conversion formula, and for the spatial position The pixel has a conversion formula as follows:
[0070]
[0071] wherein are the camera intrinsic parameters, respectively representing the focal length in the horizontal direction, the focal length in the vertical direction, the horizontal coordinate of the principal point, and the vertical coordinate of the principal point, are the pixel space indexes of the row and column respectively. After conversion by the above formula, the point cloud data is obtained, and preprocessing such as rotation, cropping, and flipping is performed to increase the sample size.
[0072] (1-2) A point transformer (PT) is constructed as a backbone network of a point cloud feature extractor to perform feature extraction on the point cloud data. The point cloud feature extractor is composed of four layers of feature extraction sub-modules. Except that the first layer of feature extraction sub-module takes the point cloud data as input, the remaining feature extraction sub-modules take the output of the previous layer as input to obtain four layers of point cloud features wherein i represents the i-th layer of the feature extractor, represents the number of point clouds in the i-th layer, is the number of feature channels in the i-th stage and the three-dimensional spatial coordinates;
[0073] (1-3) A hierarchical upsampling decoder is constructed as a point cloud feature decoder. The hierarchical upsampling decoder is composed of multiple perceptrons and jump connections in an alternating manner, and uses an interpolation algorithm to perform layer-by-layer upsampling on the point cloud features. The hierarchical upsampling decoder takes the four layers of point cloud features as input and outputs a point-by-point semantic prediction result wherein is the number of semantic categories;
[0074] (1-4) An interpolation algorithm is used to restore the four layers of point cloud features to the same resolution size to obtain , and then a two-dimensional projection algorithm is used to project the point-by-point semantic prediction result and the four layers of point cloud features into a two-dimensional space. For the point cloud at , the two-dimensional projection algorithm formula is:
[0075]
[0076] wherein is the camera intrinsic parameter, are the pixel space indexes of the projected pixel row and column respectively. After using the projection algorithm, four layers of point cloud projection feature maps and point cloud semantic prediction maps are obtained. .
[0077] Furthermore, step (2) is specifically:
[0078] (2-1) Get the depth image of the dataset And perform preprocessing such as rotation, cropping, and flipping, and build a vision transformer (ViT) as the backbone network of the deep feature extractor to extract features from the depth image. The backbone network of the deep feature extractor consists of four layers of deep feature extraction submodules. Except for the first layer of deep feature extraction submodule that takes the depth image as input, the remaining deep feature extraction submodules take the output of the previous layer as input to obtain four layers of deep feature maps. ,in , which is downsampled by a factor of two at each feature stage.
[0079] Furthermore, step (3) is specifically:
[0080] (3-1) Get the RGB image of the dataset And perform preprocessing such as rotation, cropping, and flipping;
[0081] A Mix Transformer (MiT) and a depth-guided spatial adapter module are constructed as the backbone network of the RGB feature encoder to extract features from RGB images. The RGB feature extractor consists of a four-layer RGB feature extraction sub-module and a depth-guided spatial adapter module, which consists of a depth-space calibration sub-module and a gating sub-module.
[0082] Build an RGB feature extractor, which includes a 4-layer Vision Transformer (ViT) and a depth-guided spatial adapter module.
[0083] 4-layer Vision Transformer (ViT) and Depth-guided Spatial Adapter modules. Except for the first layer of Vision Transformer which takes RGB image as input, the rest of the Vision Transformer and Depth-guided Spatial Adapter modules take the sum of the outputs of the two modules in the previous layer as input.
[0084] (3-2) Use the RGB feature extraction submodule to extract RGB semantic features and convert the RGB image As the input of the first layer RGB feature extraction submodule, and for the remaining layers , the RGB feature map of the previous layer after feature extraction submodule and depth-guided spatial adapter feature extraction As input, output RGB semantic feature map ;
[0085] (3-3) The depth-guided spatial adapter module uses depth information to calibrate the spatial structure of RGB semantic features, except for the first layer that converts the RGB image and the first layer depth feature map As input, and for the rest of the layers The depth-guided spatial adapter module extracts the RGB feature map of the previous layer after the feature extraction submodule and the depth-guided spatial adapter feature extraction. and deep feature maps As input, the depth-guided spatial adapter module first extracts deep features using a multi-scale branch module through the depth-spatial calibration submodule. The multi-scale spatial structure in is:
[0086] ,
[0087] ,
[0088] ,
[0089] in DepthwiseConvolution represents a depthwise convolution with a convolution kernel size of 1×7. Then, they represent the deep spatial structural features at different scales, and then layer normalization (LN) and deep convolution are used to enhance the RGB local feature representation. The formula is:
[0090]
[0091] in Representation layer normalization, represents a depthwise convolution with a kernel size of 5×5. In order to enhance the feature map after local feature representation, the deep spatial structure feature is combined with the RGB local feature representation to recalibrate the spatial structure. The formula is:
[0092]
[0093]
[0094]
[0095]
[0096] in Represents the RGB feature map after depth information calibration, represents the Hadamard product, represents the Gaussian Error Linear Unit (Gelu) activation function. The gating submodule dynamically models and selects feature dimensions and channels according to their importance through a linear gating mechanism. The formula is:
[0097]
[0098]
[0099] in represents the RGB gating weight, RGB feature map and RGB semantic feature map for the depth-guided adapter Perform numerical addition to obtain the output of the i-th layer RGB feature extractor , the four-layer RGB feature extractor extracts the RGB feature map .
[0100] Furthermore, step (4) is specifically:
[0101] (4-1) Construct a four-layer RGB-D feature fusion module, which consists of a cross attention layer and a multi-layer perceptron. For the i-th layer RGB-D feature fusion module, the RGB feature extractor and the depth feature extractor are extracted to obtain a feature map. and as input.
[0102] (4-2) For the i-th layer RGB-D feature fusion module, the cross-attention layer is used to exchange cross-modal features. The formula is:
[0103]
[0104]
[0105] in Linear layers for query, key, and value of RGB and Depth modalities respectively, Indicates cross attention, Represents the feature map after modal interaction.
[0106] (4-3) For the i-th layer RGB-D feature fusion module, a multi-layer perceptron is used to aggregate cross-modal features. The formula is:
[0107]
[0108] in Indicates splicing along the channel, represents a multilayer perceptron, Represents the fusion feature of the i-th layer, and the output of the four-layer RGB-D feature fusion module is the fusion feature map ;
[0109] (4-4) Use a multilayer perceptron to build an RGB-D decoder with the input being the fused feature map , the four-layer fusion features are spliced along the channel and a multi-layer perceptron is used to obtain the RGB-D semantic prediction map ;
[0110] Furthermore, step (5) is specifically:
[0111] (5-1) Calculate the true labels separately The cross entropy loss between the point cloud semantic prediction map and the RGB-D image semantic prediction map is:
[0112]
[0113]
[0114] in denote the point cloud cross entropy loss and RGB-D cross entropy loss respectively, where Indicates the vertical coordinate value of the pixel point. Indicates the horizontal axis coordinate value of the pixel point, Indicates a The semantic category categories;
[0115] (5-2) Calculate the distillation loss between the point cloud semantic prediction map and the RGB-D image semantic prediction map. The formula is:
[0116]
[0117]
[0118] in Point cloud distillation loss and RGB-D distillation loss, is the distillation temperature, is the relative entropy (Kullback-Leibler divergence, KL), and then the structured loss is calculated. First, the RGB-D segmentation boundary is obtained, and the formula is:
[0119]
[0120] in Representing coordinates Neighborhood, are the vertical and horizontal coordinate values of the pixel point, is the boundary graph, The empirical parameter is usually set to 0.1. Then the direction from the point cloud semantic prediction pixel to the boundary and the direction from the RGB-D semantic prediction map to the boundary are calculated. The formula is:
[0121]
[0122]
[0123] in The probability distribution of the point cloud and RGB-D prediction map boundaries in various directions, The vertical and horizontal coordinate values of the pixel point are used to transfer geometric structure knowledge by constraining the consistency of the boundary position direction. The formula is:
[0124]
[0125] in is the cross entropy loss, Indicates structured loss;
[0126] (5-3) Upsample the adapter RGB feature map to a uniform resolution, and calculate the mean square error loss between the adapter RGB feature map and its corresponding point cloud projection feature map layer by layer at the feature level. The formula is:
[0127]
[0128] in represents the upsampling of the spatial dimensions to align the resolution size, is the point cloud projection feature map obtained in step (2-3), is the depth-guided adapter RGB feature map obtained in step (4-3), is the L2 norm, is the mean square error loss;
[0129] (5-4) The loss value As input, the final loss function of the point cloud semantic segmentation model is calculated as , the final loss function of the RGB-D semantic segmentation model is calculated as ;
[0130] Furthermore, step (6) is specifically:
[0131] (6-1) The stochastic gradient descent algorithm is used to optimize the RGB-D semantic segmentation model consisting of an RGB feature extractor, a deep feature extractor, a depth-guided spatial adapter module, a four-layer RGB-D feature fusion module, and an RGB-D decoder, as well as the point cloud semantic segmentation model consisting of a point cloud feature extractor and a hierarchical upsampling decoder.
[0132] (6-2) Repeat steps (1) to (5) until the model converges and obtains the trained RGB-D semantic segmentation model.
[0133] Obtain the RGB-D image to be predicted and input it into the trained RGB-D semantic segmentation model to obtain the predicted pixel-level semantic category probability. Select the semantic category with the highest probability as the final pixel-level prediction result.
[0134] Example 2
[0135] like Figure 2 , an RGB-D image semantic segmentation device 910 based on cross-data view collaborative training, the RGB-D image semantic segmentation device based on cross-data view collaborative training includes: at least one processor; and, a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the RGB-D image semantic segmentation method based on cross-data view collaborative training according to one of the above embodiments.
[0136] Reference below Figure 2 , which shows a schematic diagram of the structure of an RGB-D image semantic segmentation device based on cross-data-view collaborative training suitable for implementing embodiments of the present application. The RGB-D image semantic segmentation device based on cross-data-view collaborative training in embodiments of the present application can include, but is not limited to, mobile terminals such as mobile phones, tablet computers, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 2 The RGB-D image semantic segmentation device based on cross-data view collaborative training is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0137] like Figure 2As shown, the apparatus for RGB-D image semantic segmentation based on cross-data-view collaborative training may include a CPU device 915 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 913 or programs loaded from a storage device 916 into a random access memory (RAM) 914. RAM 914 also stores various programs and data required for the operation of the apparatus for RGB-D image semantic segmentation based on cross-data-view collaborative training. The CPU device 915, ROM 913, and RAM 914 are interconnected via a bus 911. An input / output (I / O) interface 912 is also connected to the bus. Typically, the following systems can be connected to the I / O interface: input devices such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices such as a liquid crystal display (LCD), speaker, vibrator, etc.; storage device 916 such as a magnetic tape, hard disk, etc.; and communication device 917. The communication device 917 can allow the RGB-D image semantic segmentation device based on cross-data-view collaborative training to communicate wirelessly or wired with other devices to exchange data. Although the figure shows an RGB-D image semantic segmentation device based on cross-data-view collaborative training with various systems, it should be understood that it is not required to implement or have all of the illustrated systems. More or fewer systems may be implemented or have alternatively.
[0138] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 916, or installed from a ROM 913. When the computer program is executed by the CPU device 915, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.
[0139] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0140] Example 3
[0141] A computer-readable medium having computer-executable program instructions stored thereon for implementing the RGB-D image semantic segmentation method based on cross-data view collaborative training described in the above embodiment.
[0142] The computer-readable storage medium may be a variety of storage media, such as, but not limited to, a mobile hard disk, and systems, devices, or equipment made using electrical, magnetic, optical, electromagnetic, infrared, or semiconductor technologies. It may also be a combination of the above media. More specifically, examples of the computer-readable storage medium include, but are not limited to, an electrical connector containing one or more wires, a portable computer disk, a hard drive, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM, including flash memory), optical fiber media, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any combination of the above media. In this embodiment, the computer-readable storage medium refers to any tangible medium that contains or stores program code. The program code can be transmitted via a variety of media, including, but not limited to, wired transmission methods such as cables (including copper cables, coaxial cables, optical fiber cables, etc.), wireless transmission methods such as radio frequency signals (RF), infrared signals, microwave communications, Bluetooth, Wi-Fi, mobile cellular networks (such as 4G, 5G), and other wireless technologies. It may also be implemented via other suitable transmission media or combinations thereof.
[0143] In addition, the above-mentioned computer-readable storage medium can exist as a component of the RGB-D image semantic segmentation device based on cross-data view collaborative training, or as an independent storage medium without being integrated into the device.
[0144] The flowcharts and block diagrams in the accompanying drawings illustrate possible implementation architectures, functions, and operational steps of the devices, methods, and computer program products described in this application. Each process box or block may represent a module, a program, or a code unit, and these modules or codes contain one or more executable instructions that implement specific logical functions. It should be noted that in some embodiments, the order of the steps shown in the boxes in the flowchart may be adjusted, for example, certain consecutive steps may be executed in parallel, or in the opposite order, depending on the specific functional implementation requirements. At the same time, each step in the block diagram or flowchart, or a combination thereof, may be implemented by a dedicated hardware system, or by a combination of dedicated hardware and computer instructions.
[0145] The contents described in this embodiment are merely an enumeration of the implementation forms of the inventive concept. The scope of protection of the present invention should not be regarded as being limited to the specific forms described in the embodiment. The scope of protection of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.
Claims
1. An RGB-D image semantic segmentation method based on cross-data view collaborative training, characterized by: The following steps are involved: Step (1) obtain a dataset containing RGB images, depth images and real labels; Convert the RGB image and depth image into point cloud data, construct and input a point cloud feature extractor to obtain four layers of point cloud features; the RGB image and depth image are both indoor scene images; Construct a hierarchical upsampling decoder with four-layer point cloud features as input and point-by-point semantic prediction results as output. Use a two-dimensional projection algorithm to convert the point-by-point semantic prediction results and four-layer point cloud features into a point cloud semantic prediction map and a four-layer point cloud projection feature map. Step (2) constructing a depth feature extractor, inputting the depth image to obtain a four-layer depth feature map; Step (3) constructing an RGB feature extractor, wherein the RGB feature extractor consists of a four-layer RGB feature extraction submodule and a depth-guided spatial adapter module; Inputting the RGB image and the four-layer depth feature map into the RGB feature extractor to obtain a four-layer RGB feature map and a four-layer adapter RGB feature map; Step (4) construct a four-layer RGB-D feature fusion module, the input is a four-layer depth feature map and a four-layer RGB feature map, and the output is a four-layer RGB-D fused feature; A multi-layer perceptron is used to construct an RGB-D decoder, which takes as input the fused features of four layers of RGB-D and outputs an RGB-D image semantic prediction map. Step (5) calculates the cross entropy loss between the true label and the point cloud semantic prediction map, calculates the cross entropy loss between the true label and the RGB-D image semantic prediction map, calculates the distillation loss and structured loss between the point cloud semantic prediction map and the RGB-D image semantic prediction map for collaborative training, calculates the mean square error loss between the adapter RGB feature map and its corresponding point cloud projection feature map layer by layer at the feature level, and trains the model constructed by steps (1) to (4).
2. The RGB-D image semantic segmentation method based on cross-data view collaborative training according to claim 1, characterized in that In step (1), the data set also includes camera parameters; based on the camera parameters, the RGB image and the depth image are converted into point cloud data.
3. The RGB-D image semantic segmentation method based on cross-data view collaborative training according to claim 1, characterized in that The following steps are also included: The model constructed by steps (1) to (4) specifically includes: an RGB-D semantic segmentation model composed of an RGB feature extractor, a depth feature extractor, a depth-guided spatial adapter module, a four-layer RGB-D feature fusion module, and an RGB-D decoder; and a point cloud semantic segmentation model composed of a point cloud feature extractor and a hierarchical upsampling decoder; After the training, a trained RGB-D semantic segmentation model is obtained; Obtain the RGB-D image to be predicted and input it into the trained RGB-D semantic segmentation model to obtain the predicted pixel-level semantic category probability. Select the semantic category with the highest probability as the final pixel-level prediction result.
4. The RGB-D image semantic segmentation method based on cross-data view collaborative training according to claim 1, characterized in that The step (1) specifically includes the following steps: (1-1) Convert the RGB image and its corresponding depth image to obtain point cloud data; (1-2) constructing a point converter as the backbone network of the point cloud feature extractor to extract features from the point cloud data. The point cloud feature extractor consists of four layers of feature extraction submodules. Except for the first layer of feature extraction submodule that takes point cloud data as input, the other three layers of feature extraction submodules take the output of the previous layer of feature extraction submodule as input to obtain four layers of point cloud features. (1-3) Construct a hierarchical upsampling decoder as a point cloud feature decoder. The hierarchical upsampling decoder is composed of a multi-layer perceptron and a skip connection, and uses an interpolation algorithm to upsample the four-layer point cloud features obtained in step (1-2) layer by layer. The hierarchical upsampling decoder takes the four-layer point cloud features as input and outputs a point-by-point semantic prediction result. (1-4) Use the interpolation algorithm to restore the four-layer point cloud features to the same resolution size, and use the two-dimensional projection algorithm to project the point-by-point semantic prediction results and the four-layer point cloud features into two-dimensional space to obtain the four-layer point cloud projection feature map and the point cloud semantic prediction map.
5. The RGB-D image semantic segmentation method based on cross-data view collaborative training according to claim 1, characterized in that In step (2), the construction of a deep feature extractor includes: A visual converter is constructed as the backbone network of a deep feature extractor for extracting features from depth images. The backbone network of the deep feature extractor consists of four layers of deep feature extraction sub-modules. Except for the first layer of deep feature extraction sub-module which takes the depth image as input, the other three layers of deep feature extraction sub-modules take the output of the previous layer of deep feature extraction sub-module as input to obtain four layers of deep feature maps. Each layer of depth feature map is downsampled twice based on the previous layer.
6. The RGB-D image semantic segmentation method based on cross-data view collaborative training according to claim 5, characterized in that The RGB feature extractor constructed in step (3) includes: Construct a visual converter and a four-layer depth-guided spatial adapter as the backbone network of the RGB feature extractor to extract features from RGB images. The visual converter includes a four-layer RGB feature extraction submodule; In the 4-layer RGB feature extraction submodule and the 4-layer depth-guided spatial adapter, except for the first layer RGB feature extraction submodule which takes the RGB image as input, the remaining RGB feature extraction submodules take the sum of the outputs of the previous layer RGB feature extraction submodule and the depth-guided spatial adapter module as input; the output of each layer RGB feature extraction submodule is: RGB semantic feature map; Except for the first layer of depth-guided spatial adapter module, which takes the RGB image and the first layer of depth feature map obtained in step (2) as input, the depth-guided spatial adapter modules of the remaining layers take the RGB feature map and depth feature map obtained by the RGB feature extraction submodule and the depth-guided spatial adapter module of the previous layer as input. The depth-guided spatial adapter module consists of a depth-space calibration submodule and a linear gating submodule connected in series; The depth space calibration submodule includes three branch structures of different scales; each branch structure consists of two layers of depth convolution; The depth space calibration submodule also includes a depth convolution module, which is used to process the RGB feature map obtained in the previous layer after layer normalization, and then perform Hadamard product operation on the output with the output of the branch structures of three different scales and add them together.
7. The RGB-D image semantic segmentation method based on cross-data view collaborative training according to claim 1, characterized in that Step (4) specifically includes the following steps: (4-1) Construct a four-layer RGB-D feature fusion module, which consists of a cross-attention layer and a multi-layer perceptron. Each layer of RGB-D feature fusion module takes the feature map extracted by the corresponding layer RGB feature extractor and deep feature extractor as input, and processes it through the cross attention layer and multi-layer perceptron to obtain a fused feature map; A multi-layer perceptron is used to construct an RGB-D decoder. The input is a fused feature map. The four-layer fused features are concatenated along the channel and a multi-layer perceptron is used to obtain an RGB-D semantic prediction map.
8. The RGB-D image semantic segmentation method based on cross-data view collaborative training according to claim 1, characterized in that Step (5) specifically includes the following steps: Calculate the cross entropy loss between the true label and the point cloud semantic prediction map , calculate the cross entropy loss between the true label and the RGB-D image semantic prediction map , calculate the distillation loss and structured loss between the point cloud semantic prediction map and the RGB-D image semantic prediction map, calculate the mean square error loss of the adapter RGB feature map and its corresponding point cloud projection feature map layer by layer at the feature level, construct the total loss and perform collaborative training; The distillation loss consists of the distillation loss from the point cloud semantic prediction map to the RGB-D image semantic prediction map and the distillation loss from the RGB-D image semantic prediction map to the point cloud semantic prediction map. The difference between the point cloud semantic prediction map and the RGB-D image semantic prediction map is measured by relative entropy, and the distillation temperature is used to smooth the probability distribution to obtain the difference between the point cloud semantic prediction map and the RGB-D image semantic prediction map. The difference between RGB-D image semantic prediction map and point cloud semantic prediction map ; The calculation of the structured loss includes: first, predicting the semantic map of the RGB-D image and combining the empirical hyperparameters The RGB-D segmentation boundary is obtained, and then the gradient of the local probability distribution is obtained through the point cloud semantic prediction map and the RGB-D image semantic prediction map, and the probability distribution from the point cloud semantic prediction map to the boundary direction and the probability distribution from the RGB-D image semantic prediction map to the boundary direction are calculated; then the relative entropy between the probability distribution from the RGB-D image semantic prediction map to the boundary direction at the RGB-D segmentation boundary and the probability distribution from the point cloud semantic prediction map to the boundary direction at the corresponding spatial position is obtained. ; Upsample the adapter RGB feature map to unify the resolution size, and calculate the mean square error loss between the adapter RGB feature map and its corresponding point cloud projection feature map layer by layer at the feature level , Calculate the final loss function of the point cloud semantic segmentation model , calculate the final loss function of the RGB-D semantic segmentation model .
9. A computer-readable medium storing a computer program, characterized in that: The computer program can be executed by a processor to implement the steps included in the RGB-D image semantic segmentation method based on cross-data view collaborative training as described in any one of claims 1 to 8.
10. A device for RGB-D image semantic segmentation based on cross-data view collaborative training, comprising a processor and a memory, characterized in that: A computer program is stored on the memory, which is executed by the processor to implement the steps included in the RGB-D image semantic segmentation method based on cross-data view collaborative training as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Rapid indoor scene understanding method
CN113536987A
Semantic map construction method based on RGBD visual segmentation algorithm
CN119323772A