A Scene Semantic Segmentation Method and System Based on RGB-D Feature Fusion

By adopting the RGB-D feature fusion method in scene semantic segmentation, combining attention mechanism, feature refinement network and context feature processing network, the problem of low semantic segmentation accuracy in indoor scenes is solved, and higher segmentation accuracy and better environmental adaptability are achieved.

CN113888557BActive Publication Date: 2025-06-27SHANDONG NORMAL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111105921.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-22
Publication Date
2025-06-27
Estimated Expiration
2041-09-22

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problems of low semantic segmentation accuracy due to complex background, uneven lighting, weak recognition of low-level visual features and a large number of occlusions between objects in indoor scenes.

Method used

The scene semantic segmentation method based on RGB-D feature fusion is adopted, and the RGB-D feature fusion network constructed by the attention mechanism is used to fusion of RGB features and deep features, combining feature refinement networks and context feature processing networks to optimize the semantic segmentation effect.

Benefits of technology

It improves the semantic segmentation accuracy of indoor scenes, enhances spatial position sensitivity and complex environment adaptability, can effectively deal with problems such as complex background and uneven lighting, and obtains relatively accurate segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113888557B_ABST
    Figure CN113888557B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for scene semantic segmentation based on RGB-D feature fusion. First, an RGB image and a depth image of the scene to be segmented are obtained; then, the RGB image and the depth image are simultaneously input into the scene semantic segmentation model to obtain the scene semantic segmentation result. Among them, the encoder of the scene semantic segmentation model processes the RGB image and the depth image respectively using an RGB branch and a depth branch. During the processing, an RGB-D feature fusion network is used to fuse RGB features and depth features to obtain multiple low-level fusion features and a high-level fusion feature, so as to achieve the purpose of enhancing target features and suppressing the background. At the same time, the introduced feature refinement network and context feature processing network can reduce semantic information loss and optimize the semantic segmentation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of scene semantic segmentation, and particularly relates to a scene semantic segmentation method and system based on RGB-D feature fusion. Background Technique

[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.

[0003] Compared with outdoor scenes, the image structure in indoor scenes is often more complex, with problems such as uneven illumination, mutual occlusion between objects, and large similarities in the colors and structures of different objects. Most current semantic segmentation methods focus on standard outdoor conditions, and indoor semantic segmentation that has not been deeply studied is still challenging in many aspects.

[0004] The RGB images of indoor scenes will lose a large amount of spatial information in the scenes, so ordinary image segmentation methods cannot effectively segment indoor scenes. In response to this situation, researchers choose to add depth information to improve the segmentation effect. Because depth information is almost not affected by light changes and the colors and textures of objects, it can provide corresponding geometric relationships for RGB images and highlight the spatial information of objects in the scene. The features learned by this method using depth information are richer, have stronger expression ability, and are more suitable for complex indoor scenes. At the same time, depth sensor technologies such as Kinect are already quite mature, and depth images can be easily obtained, so the accuracy of segmentation can be improved by studying deep learning methods that fuse RGB image information and depth information.

[0005] However, the current methods still cannot solve the problem of low segmentation accuracy caused by complex backgrounds, uneven illumination, weak recognition ability of low-level visual features, and a large number of occlusions between objects in the scene. Summary of the Invention

[0006] In order to solve the technical problems existing in the above background technique, the present invention provides a scene semantic segmentation method and system based on RGB-D feature fusion. The RGB-D feature fusion network constructed based on the attention mechanism can perform selective weighting processing on the feature map to achieve the purpose of enhancing target features and suppressing the background; at the same time, the introduced feature refinement network and context feature processing network can reduce semantic information loss and optimize the semantic segmentation effect.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] The first aspect of the present invention provides a scene semantic segmentation method based on RGB-D feature fusion, which includes:

[0009] Obtain the RGB image and depth image of the scene to be segmented;

[0010] Input the RGB image and depth image into the scene semantic segmentation model simultaneously to obtain the scene semantic segmentation result;

[0011] Among them, the encoder of the scene semantic segmentation model processes the RGB image and depth image respectively using the RGB branch and depth branch. During the processing, an RGB-D feature fusion network is used to fuse the RGB features and depth features, obtaining multiple low-level fusion features and one high-level fusion feature.

[0012] Further, the RGB branch processes the RGB image using an RGB convolutional layer and then inputs it into multiple sequentially connected RGB branch layers;

[0013] The depth branch processes the depth image using a depth convolutional layer and then inputs it into multiple sequentially connected depth branch layers.

[0014] Further, the depth branch layers correspond one-to-one with the RGB branch layers, and the depth features output by the depth convolutional layer and the RGB features output by its corresponding RGB branch layer are fused through the RGB-D feature fusion network.

[0015] Further, the RGB-D feature fusion network is specifically:

[0016] Connect the RGB features and depth features to obtain the concatenated features;

[0017] At the same time, the RGB features, depth features, and concatenated features are respectively subjected to global average pooling and then sent to different MLP layers;

[0018] The result after adding the output elements of the MLP layer is processed by the Sigmoid function;

[0019] The processed result is respectively subjected to channel multiplication with the RGB features and depth features and then element-wise addition to obtain the output result.

[0020] Further, the multiple low-level fusion features enter the decoder after being processed by the feature refinement network;

[0021] The feature refinement network first inputs the low-level fusion features into a convolutional layer for convolution operation and then sends them to the channel attention layer for processing.

[0022] Further, the high-level fusion feature enters the decoder after being processed by the context feature processing network;

[0023] The context feature processing network performs a 1×1 dilated convolution operation, a global pooling operation, and multiple 3×3 dilated convolution operations with different dilation coefficients on the high-level fusion feature respectively.

[0024] Further, the decoder processes multiple low-level fusion features and one high-level fusion feature through an upsampling layer and an RGB-D feature fusion network to obtain the features of the scene to be segmented.

[0025] The second aspect of the present invention provides a scene semantic segmentation system based on RGB-D feature fusion, which includes:

[0026] An image acquisition module, which is configured to: acquire an RGB image and a depth image of the scene to be segmented;

[0027] A semantic segmentation module, which is configured to: input the RGB image and the depth image into the scene semantic segmentation model simultaneously to obtain a scene semantic segmentation result;

[0028] Among them, the encoder of the scene semantic segmentation model processes the RGB image and the depth image respectively by using an RGB branch and a depth branch, and an RGB-D feature fusion network is used for fusing RGB features and depth features during the processing to obtain multiple low-level fusion features and one high-level fusion feature.

[0029] The third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in a scene semantic segmentation method based on RGB-D feature fusion as described above.

[0030] The fourth aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps in a scene semantic segmentation method based on RGB-D feature fusion as described above.

[0031] Compared with the prior art, the beneficial effects of the present invention are:

[0032] The present invention provides a scene semantic segmentation method based on RGB-D feature fusion, which uses the depth information of the scene for semantic segmentation of indoor scenes, and has better spatial position sensitivity and complex environment adaptability, because the depth information is not affected by light changes, occlusion, etc., and can well retain the geometric information of the objects in the scene.

[0033] The present invention provides a method for scene semantic segmentation based on RGB-D feature fusion, which integrates scene depth information and RGB color information into a deep learning network structure model, solves the problem of a large amount of spatial information loss in scene images due to complex image structures, and the method is fast, stable, and has a high segmentation accuracy. It can be used to solve the problems of complex backgrounds, uneven illumination, weak recognition ability of low-level visual features, and a large number of occlusions between objects in the scene, and can obtain relatively accurate segmentation results, which is applicable to semantic segmentation of indoor scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments and descriptions thereof of the invention are used to explain the invention and do not constitute an improper limitation of the invention.

[0035] Figure 1 is the overall framework diagram of the scene semantic segmentation model in the first embodiment of the present invention;

[0036] Figure 2 (a) is the framework diagram of the RGB branch layer and the depth branch layer in the first embodiment of the present invention;

[0037] Figure 2 (b) is the framework diagram of the downsampling layer in the first embodiment of the present invention;

[0038] Figure 3 is the framework diagram of the traditional channel attention mechanism;

[0039] Figure 4 is the framework diagram of the RGB-D feature fusion network in the first embodiment of the present invention;

[0040] Figure 5 is the framework diagram of the feature refinement network in the first embodiment of the present invention;

[0041] Figure 6 is the framework diagram of the context feature processing network in the first embodiment of the present invention;

[0042] Figure 7 is the convolution schematic diagram corresponding to different coefficients in the first embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0044] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0045] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0046] Embodiment 1

[0047] As Figure 1 As shown, this embodiment provides a scene semantic segmentation method based on RGB-D feature fusion. This method uses ResNet-50 as the basic framework on the basis of the encoding-decoding architecture, and processes the original data in the way of separate training and gradual fusion. RGB and depth image features are extracted by two parallel and independent convolutional branches respectively. The encoder is responsible for extracting image semantic features step by step, and the decoder upsamples the extracted features. The RGB-D feature fusion module constructed based on the attention mechanism can perform selective weighting on the feature maps to enhance the target features and suppress the background; at the same time, the introduced feature refinement network and context feature processing network can reduce the loss of semantic information and optimize the semantic segmentation effect. The specific steps are as follows:

[0048] Step 1: Obtain the RGB image and depth image of the scene to be segmented.

[0049] (1) Image acquisition: Use the Kinect V2 camera to collect the RGB image and depth image Depth of the sample point at the same angle under natural light source at the same sampling point. The images of the scene to be segmented are all collected in the natural light environment (including front light and backlight conditions). The lens used for image acquisition: Kinect V2 (Microsoft). The RGB image and depth image acquisition resolutions of the Kinect V2 camera are uniformly set to 480×640, and the output format is JPG. Keep the camera angle unchanged, and collect the RGB image and depth image of the target object at the same sample point respectively. For each sample sampling point, the RGB and depth images of the target object are output simultaneously.

[0050] (2) Image processing: Randomly scale, crop and flip all input images respectively, and perform normalization operations on the RGB image and depth image respectively. Further enhance the input RGB image by randomly adjusting the hue, brightness and saturation in the HSV space. The obtained images have relatively large noise, so the images are enhanced.

[0051] Step 2: Input the RGB image and the depth image into the scene semantic segmentation model simultaneously to obtain the scene semantic segmentation result. The scene semantic segmentation model includes an encoder, a decoder, and a classifier. The encoder of the scene semantic segmentation model processes the RGB image and the depth image respectively using an RGB branch and a depth branch. During the processing, an RGB-D feature fusion network is used to fuse the RGB features and the depth features, obtaining multiple low-level fusion features and one high-level fusion feature. The high-level fusion feature enters the decoder after passing through the context feature processing network; the multiple low-level fusion features enter the decoder after being processed by the feature refinement network. The decoder processes the multiple low-level fusion features and one high-level fusion feature through an upsampling layer and an RGB-D feature fusion network to obtain the features of the scene to be segmented.

[0052] As an implementation, the overall structure of the scene semantic segmentation model is as Figure 1 shown. An encoder-decoder structure is adopted. The encoder uses ResNet-50 as the basic architecture, including five downsampling modules, and an RGB-D feature fusion module is introduced to strengthen the fusion of RGB features and depth features, extract the semantic information of the image, and send the processing result to the context feature processing network Context; the result processed by the context feature processing network is input to the decoder to upsample the image. At the same time, a feature refinement module is introduced to improve the semantic segmentation accuracy. The decoder also contains five upsampling modules. The encoder is responsible for gradually extracting the semantic features of the image, and the decoder upsamples the extracted features. The structures of the downsampling module and the upsampling module are as Figure 2 shown.

[0053] (1) The RGB branch processes the RGB image using an RGB convolutional layer and then inputs it into multiple sequentially connected RGB branch layers; the depth branch processes the depth image using a depth convolutional layer and then inputs it into multiple sequentially connected depth branch layers. The depth features output by the depth convolutional layer and the RGB features output by the RGB convolutional layer are fused through an RGB-D feature fusion network; the depth branch layers and the RGB branch layers correspond one by one, and the depth features output by the depth convolutional layer and the RGB features output by its corresponding RGB branch layer are fused through an RGB-D feature fusion network.

[0054] The encoder uses two independent convolution branches (RGB branch and depth branch) to process RGB images and depth images respectively. During the processing, the RGB-D feature fusion network is used to fuse RGB features and depth features. The RGB branch and the depth branch have the same network configuration, but because the RGB branch inputs RGB three-channel image information and the depth branch inputs single-channel depth image information, the parameters of the Conv1_D layer in the depth branch are not exactly the same as the parameters of the Conv1 layer in the RGB branch. Two independent convolution branch networks based on ResNet are used to process the input RGB image and depth image respectively. Because the color of the RGB image of the indoor scene will blur the boundaries between objects and lose a lot of spatial information in the scene, adding depth information can provide corresponding geometric relationships for the RGB image and well preserve the spatial information of objects in the scene.

[0055] As an implementation method, the number of RGB branch layers and depth branch layers is four. The RGB image and depth image are input into the network, and they are trained and gradually integrated separately, and then embedded into the RGB-D feature fusion network. After the features of the RGB image branch and the depth image branch are processed by the feature fusion module, the results of the fusion processing are processed by the RGB branch, and the original depth branch output results continue to be processed by the depth branch. This process is repeated until all data information is collected in the five-layer RGB branch, and the model encoding is completed. Specifically:

[0056] (1-1) After the RGB image and the depth image are processed by the RGB convolution layer (Conv1 layer) and the depth convolution layer (Conv1_D layer) respectively, the two results are input into the RGB-D feature fusion network RFM respectively. The result after the RGB-D feature fusion network processing is the first fusion feature, which serves as the new input of the first RGB branch layer (Layer1 layer). The output of the original depth convolution layer (Conv1_D layer) continues to be the input of the first depth branch layer (Layer1_D) layer.

[0057] (1-2) By analogy with the operation method in (1-1), the input data of the second RGB branch layer (Layer2), the third RGB branch layer (Layer3), the fourth RGB branch layer (Layer4), and the context feature processing network (Context) respectively come from: the output data of the first RGB branch layer Layer1, the second RGB branch layer Layer2, the third RGB branch layer Layer3, and the fourth RGB branch layer Layer4, and the output data of the first depth branch layer Layer1_D, the second depth branch layer Layer2_D, the third depth branch layer Layer3_D, and the fourth depth branch layer Layer4_D, after being processed and fused by the RGB-D feature fusion network (the second fusion feature, the third fusion feature, the fourth fusion feature, and the fifth fusion feature).

[0058] Finally, the encoder outputs multiple low-level fusion features (the first fusion feature, the second fusion feature, the third fusion feature, the fourth fusion feature) and a high-level fusion feature (the fifth fusion feature).

[0059] The method of separately training and gradually fusing is adopted to fuse the information of the RGB image and the depth image. The RGB-D feature fusion network is embedded in the RGB branch layer, aiming to effectively select the RGB features and the depth features, and the result after being processed by the feature fusion module is used as the input of the next layer of the RGB branch.

[0060] Among them, the structure of each RGB branch layer and each depth branch layer is as Figure 2 shown in (a), including a 1×1 convolutional layer, and a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer connected in sequence; the input data is processed by the 1×1 convolutional layer to obtain the first output, the input data is processed by the 1×1 convolutional layer, the 3×3 convolutional layer, and the 1×1 convolutional layer connected in sequence to obtain the second output, and the first output and the second output are fused to obtain the final output.

[0061] (1-3) The RGB image and the depth image have different feature distributions. Introducing the RGB-D feature fusion network RFM constructed based on the channel attention mechanism can make the network focus on the regions with richer information and filter out some unnecessary features.

[0062] The traditional construction of the channel attention mechanism is as Figure 3 shown: Assume the input feature map is globally average pooled to obtain the output one-dimensional feature map

[0063]

[0064] where, z cDenote the output result with the number of channels being \(c\), where \(H\) and \(W\) represent the height and width of the feature map respectively. Such processing allows the network to collect global information, and the subsequent operations can be written as

[0065]

[0066] Among them, represents channel multiplication, \(\sigma\) represents the Sigmoid activation function, \(U\) is the final output result, and \(g\) is the final attention tensor obtained through the Conv layer. This Conv operation can be written as

[0067] \(g = T2(ReLU(T1(z)))\)

[0068] The Conv layer here includes two different \(1\times1\) convolutional layers and a Relu layer, whose function is to mine the correlations between channels. Through the first convolution, an intermediate attention tensor can be obtained, and then through the second convolution, the final result attention tensor \(g\) can be obtained.

[0069] In this embodiment, based on the channel attention mechanism, a new RGB-D feature fusion network is constructed. The structure of the RGB-D feature fusion network of this application is as Figure 4 shown: Assume the input RGB feature map is and the depth feature map Connecting the RGB feature map and the depth feature map Depth can obtain a new feature module, namely the concatenated feature Then, simultaneously perform global average pooling on the RGB feature map RGB, the depth feature map Depth, and the concatenated feature RGB cat respectively, and send the three obtained results into different MLP layers for convolution processing. Each MLP layer (MLP1, MLP2, and MLP3) includes two different \(1\times1\) convolutional layers and a Relu layer, and its function is still to mine the correlations between channels. Note that after the second convolution processing of the convolutional layer that processes RGB cat the number of input feature channels will change from the original \(2C\) to the initial \(C\). Finally, the result after adding the output elements of the MLP layer is processed by the Sigmoid function. After performing channel multiplication on this final result with the RGB feature and the depth feature respectively and then adding the elements, the final output result can be obtained.

[0070] (2) Multiple low-level fusion features are sequentially input into the feature refinement network. The feature refinement network first performs a convolution operation on the low-level fusion features by inputting them into a convolutional layer, and then sends them to the channel attention layer for processing.

[0071] The feature refinement network aims to guide low-level features with high-level semantic features. Since there are significant differences between high-level semantic information and low-level semantic information, the direct fusion effect is relatively poor. Embedding a feature refinement module constructed based on the attention mechanism in the decoder network can improve the fusion effect of the two semantic informations.

[0072] The skip structure can be used to retain high-level features to guide low-level features. However, the features at different stages have differences, and the channel attention mechanism can effectively make up for this difference; the feature refinement network constructed by using the channel attention mechanism is embedded in the skip structure. The construction of the feature refinement network Refine is as Figure 5 shown. As can be seen from the figure, the module consists of a convolutional layer and a channel attention mechanism module. Note that since the channel sizes of the features at different stages need to be aligned before fusion, the low-level fusion features are first input into the convolutional layer for a 1×1 convolutional operation and then sent to the channel attention layer for processing. The convolutional layer plus the channel attention layer constitute the feature refinement network.

[0073] The feature refinement network adopts a skip connection structure and is embedded in the feature refinement network to perform a deep complementarity between low-level information and high-level information.

[0074] (3) High-level fusion feature context feature processing network. The context feature processing network Context performs a 1×1 dilated convolutional operation, a global pooling operation, and multiple 3×3 dilated convolutional operations with different dilation coefficients on the high-level fusion features respectively.

[0075] After all the data is aggregated, it is processed by a context feature processing network and then sent to the decoder to obtain the final prediction, which can help the features to be better propagated and fused.

[0076] When the output of the encoder is sent to the decoder, a large amount of feature information is often lost. In order to retain more feature context information, a context feature processing network is embedded before the data is sent to the decoder. The context feature processing network detects the convolutional feature layer with filters of multiple sampling rates and effective fields of view, so as to capture objects and image context at multiple scales. The construction of the context feature processing network is as Figure 6 shown. Dilated convolution injects holes into the standard convolution to increase the receptive field of the convolutional kernel. Compared with the original normal convolution, dilated convolution can make each output contain a larger range of information. It has an additional parameter called the dilation coefficient, which refers to the number of intervals of the convolutional kernel. Here, three 3×3 dilated convolutions are performed, and the dilation coefficients are 6, 12, and 18 respectively. The convolutions corresponding to different coefficients are as Figure 7As shown. In addition to these three dilated convolutions, a 1×1 dilated convolution and a global pooling operation are also performed. These five operations together form the context feature processing network. Through such an information processing method, the network can capture more context information.

[0077] (4) The decoder processes multiple low-level fusion features and a high-level fusion feature through an upsampling layer and an RGB-D feature fusion network to obtain the features of the scene to be segmented, and inputs them into the classifier. Multiple upsampling layers are connected in sequence to gradually restore the features to the original spatial resolution size.

[0078] The decoding stage of the model is responsible for upsampling the extracted features and finally outputting an output image with the same size as the input image for the segmentation result.

[0079] As an implementation, the high-level fusion feature is input into the first upsampling layer UP1. After the upsampling operation, it is fused with the fourth fusion feature through the RGB-D feature fusion network and then input into the second upsampling layer UP2; the output of the second upsampling layer is fused with the third fusion feature through the RGB-D feature fusion network and then input into the third upsampling layer UP3; the output of the third upsampling layer is fused with the second fusion feature through the RGB-D feature fusion network and then input into the fourth upsampling layer UP4; the output of the fourth upsampling layer UP4 is fused with the first fusion feature through the RGB-D feature fusion network and then input into the fifth upsampling layer UP5; the fifth upsampling layer UP5 outputs the features of the final scene to be segmented.

[0080] Among them, the structure of each upsampling layer is as Figure 2 (b) shown, including a 2×2 convolutional layer, and a 1×1 convolutional layer and a 3×3 convolutional layer connected in sequence; the input data is processed by the 2×2 convolutional layer to obtain the first output, the input data is processed by the 1×1 convolutional layer and the 3×3 convolutional layer connected in sequence to obtain the second output, and the first output and the second output are fused to obtain the final output.

[0081] (5) When training the model, a loss function is needed to measure the performance. This model uses the cross-entropy function to evaluate the model:

[0082]

[0083] Among them, P(x = k) is the probability that the pixel belongs to the correct category k; K is the number of categories; s i is the feature value of the i-th category. When the network uses the last softmax function, the formula of the cross-entropy function is:

[0084]

[0085] Three types of metrics are used to measure the performance of different semantic segmentation models, namely pixel accuracy, mean accuracy, and mean intersection over union accuracy.

[0086] 1) Pixel accuracy (PA): This is the simplest metric, which simply calculates the ratio between the number of correctly classified pixels and their total number.

[0087]

[0088] 2) Mean accuracy (MPA): A slightly improved pixel accuracy, where the ratio of correct pixels is calculated on a per-class basis and then averaged over the total number of classes.

[0089]

[0090] 3) Mean intersection over union accuracy (MIoU): This is the standard metric for segmentation purposes. The percentage of the correctly segmented area of the algorithm that intersects with the total area of the correctly segmented area and the output segmented area.

[0091]

[0092] Assume there are a total of k + 1 classes, from 0 to k, including the empty class or background. p ij is the number of pixels that actually belong to class i but are inferred to belong to class j, and p ii is the number of pixels correctly classified as class i.

[0093] To verify the functionality of the RGB-D feature fusion module in the model structure, an ablation study was conducted by comparing the original model with three defective models. The comparison results of these ablation experiments can prove that the constructed RGB-D feature fusion module indeed plays a huge role in the model. And by comparing the model of the present invention with some other common models on the dataset, the experiments show that the scene semantic segmentation model of the present invention can outperform most of the existing commonly used network structures for RGB-D semantic segmentation.

[0094] Example Two

[0095] This example provides a scene semantic segmentation system based on RGB-D feature fusion, which specifically includes the following modules:

[0096] An image acquisition module, which is configured to: acquire the RGB image and depth image of the scene to be segmented;

[0097] A semantic segmentation module, which is configured to: input the RGB image and depth image into the scene semantic segmentation model simultaneously to obtain the scene semantic segmentation result;

[0098] Among them, the encoder of the scene semantic segmentation model processes the RGB image and the depth image through the RGB branch and the depth branch respectively. During the processing, an RGB-D feature fusion network is used to fuse the RGB features and the depth features, obtaining multiple low-level fusion features and one high-level fusion feature.

[0099] It should be noted here that each module in this embodiment corresponds one by one to each step in Embodiment 1, and the specific implementation process is the same, so it will not be repeated here.

[0100] Embodiment 3

[0101] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in a method for scene semantic segmentation based on RGB-D feature fusion as described in Embodiment 1 above.

[0102] Embodiment 4

[0103] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a method for scene semantic segmentation based on RGB-D feature fusion as described in Embodiment 1 above.

[0104] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can adopt the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.

[0105] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0106] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in the process Figure 1 step or steps and / or block Figure 1 or blocks specified in the function.

[0107] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the process Figure 1 step or steps and / or block Figure 1 or blocks specified in the function.

[0108] Those of ordinary skill in the art will appreciate that all or part of the processes of the methods of the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0109] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A scene semantic segmentation method based on RGB-D feature fusion, characterized in that Including: Obtain the RGB image and depth image of the scene to be segmented; Input the RGB image and depth image into the scene semantic segmentation model simultaneously to obtain the scene semantic segmentation result; Among them, the encoder of the scene semantic segmentation model processes the RGB image and depth image respectively using the RGB branch and the depth branch. During the processing, the RGB-D feature fusion network is used to fuse the RGB features and depth features to obtain multiple low-level fusion features and one high-level fusion feature; the multiple low-level fusion features enter the decoder after being processed by the feature refinement network; the high-level fusion feature enters the decoder after being processed by the context feature processing network; the decoder processes the multiple low-level fusion features and one high-level fusion feature through the upsampling layer and the RGB-D feature fusion network to obtain the features of the scene to be segmented; The RGB-D feature fusion network is specifically: Connect the RGB feature and the depth feature to obtain the concatenated feature; Simultaneously perform global average pooling on the RGB feature, depth feature, and concatenated feature respectively and then send them to different MLP layers; The result after adding the output elements of the MLP layer is processed by the Sigmoid function; The processed result is multiplied by the RGB feature and the depth feature in channels respectively and then added element by element to obtain the output result.

2. The method for scene semantic segmentation based on RGB-D feature fusion according to claim 1, wherein The RGB branch processes the RGB image using the RGB convolutional layer and then inputs it into multiple sequentially connected RGB branch layers; The depth branch processes the depth image using the depth convolutional layer and then inputs it into multiple sequentially connected depth branch layers.

3. The method for scene semantic segmentation based on RGB-D feature fusion according to claim 2, characterized in that The depth branch layers correspond one by one to the RGB branch layers, and the depth features output by the depth convolutional layer and the RGB features output by the corresponding RGB branch layers are fused through the RGB-D feature fusion network.

4. A method for scene semantic segmentation based on RGB-D feature fusion according to claim 1, wherein The feature refinement network first inputs the low-level fusion feature into the convolutional layer for convolution operation and then sends it to the channel attention layer for processing.

5. The scene semantic segmentation method based on RGB-D feature fusion according to claim 1, characterized in that, The context feature processing network performs a 1×1 dilated convolution operation, a global pooling operation, and multiple 3×3 dilated convolution operations with different dilation coefficients on the high-level fusion feature respectively.

6. A scene semantic segmentation system based on RGB-D feature fusion, characterized in that, Including: An image acquisition module configured to: obtain the RGB image and depth image of the scene to be segmented; A semantic segmentation module configured to: input the RGB image and depth image into the scene semantic segmentation model simultaneously to obtain the scene semantic segmentation result; Among them, the encoder of the scene semantic segmentation model processes the RGB image and depth image respectively using the RGB branch and the depth branch. During the processing, the RGB-D feature fusion network is used to fuse the RGB features and depth features to obtain multiple low-level fusion features and one high-level fusion feature; the multiple low-level fusion features enter the decoder after being processed by the feature refinement network; the high-level fusion feature enters the decoder after being processed by the context feature processing network; the decoder processes the multiple low-level fusion features and one high-level fusion feature through the upsampling layer and the RGB-D feature fusion network to obtain the features of the scene to be segmented; The RGB-D feature fusion network is specifically: Connect the RGB feature and the depth feature to obtain the concatenated feature; Meanwhile, the RGB features, depth features, and concatenated features are respectively subjected to global average pooling and then fed into different MLP layers; The results after adding the output elements of the MLP layers are processed by the Sigmoid function; The processed results are respectively subjected to channel multiplication with the RGB features and depth features and then added element-wise to obtain the output result.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in a method for RGB-D feature fusion-based scene semantic segmentation according to any one of claims 1-5.

8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in a method for RGB-D feature fusion-based scene semantic segmentation according to any one of claims 1-5.

Citation Information

Patent Citations

  • Indoor scene semantic segmentation method based on convolutional neural network

    CN111563507A