A method and system for semantic segmentation of RGB-D images based on spatial semantic enhancement

CN122416038BActive Publication Date: 2026-09-08SOUTHWEAT UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610864254.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-16
Publication Date
2026-09-08
Estimated Expiration
2046-06-16

AI Technical Summary

Technical Problem

现有RGB-D图像多模态特征提取与语义分割技术中存在的以下具体问题:(1)针对现有深度图像存在高频散斑噪声和大面积空洞,导致几何信息失真、影响分割精度的问题

Benefits of technology

1、本申请串联了BM3D协同滤波与Navier-Stokes流体力学修复。BM3D精准剥离散斑噪声,Navier-Stokes方程则利用等照度线扩散原理修复大面积空洞,无需依赖大规模数据预训练即可保持物体表面的曲率连续性,从数据源头为多模态网络提供了极高纯度的几何先验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122416038B_ABST
    Figure CN122416038B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of computer vision and image data, and discloses an RGB-D image semantic segmentation method and system based on spatial semantic enhancement, which comprises the following steps: acquiring original RGB images and original depth images of a target scene collected at the same angle and synchronously, obtaining high-fidelity depth images according to the original depth images, realizing cross-modal interaction of RGB features and depth features through a multi-modal low-rank fine-tuning MMLoRA module, extracting multi-modal feature maps, obtaining spatial enhancement feature maps according to the multi-modal feature maps and parameters of a discretized state space model SSM, obtaining multi-scale fusion feature maps according to the spatial enhancement feature maps, inputting the multi-scale fusion feature maps into a decoding head to complete pixel-level dense classification, so that a target to be grabbed and obstacles are completely separated from a complex background of the target scene, and a semantic segmentation prediction map containing accurate contours of the target to be grabbed and obstacle boundaries is output. The application improves the accuracy of segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer vision and image data technology, and in particular to an RGB-D image semantic segmentation method and system based on spatial semantic enhancement. Background Technology

[0002] In unstructured scenarios, a multimodal visual perception system based on an RGB-D (color-depth) camera is the core prerequisite for achieving autonomous grasping. The following specific problems exist in existing RGB-D image multimodal feature extraction and semantic segmentation technologies: (1) The existence of high-frequency speckle noise and large-area holes in existing depth images leads to geometric information distortion and affects segmentation accuracy. (2) The limited receptive field of traditional multimodal feature extraction networks makes it impossible to effectively capture the global topological structure of irregular objects. (3) When fusing RGB features and depth features, existing networks are insufficient in perceiving spatial details such as the edges of slender objects and small parts, and cannot adaptively and dynamically allocate multimodal feature weights. Summary of the Invention

[0003] In view of this, this application provides a multimodal image segmentation method and system based on spatial semantic enhancement of RGB-D images. It can repair geometric distortion from the data preprocessing end and take into account both global receptive field and fine spatial semantic enhancement at the model end. It is particularly suitable for robot vision grasping target recognition and semantic segmentation in unstructured scenarios.

[0004] This application discloses a semantic segmentation method for RGB-D images based on spatial semantic enhancement, which includes: Step 1: Acquire the original RGB image and original depth image of the target scene from the same viewpoint and simultaneously. The original RGB image and the original depth image constitute an RGB-D image. Perform geometric restoration preprocessing on the original depth image to obtain the preprocessed depth image. Combine BM3D collaborative filtering and Navier-Stokes fluid dynamics equations to filter out speckle noise and repair missing parts in the preprocessed depth image, outputting a high-fidelity depth image. The original RGB image and the original depth image represent the target to be grasped and background obstacles in the target scene. The target to be grasped includes at least one of slender objects and irregular parts. Step 2: Construct a ViT dual-stream encoder based on the visual foundation model SAM prior, and realize cross-modal interaction between RGB features and depth features through the multimodal low-rank fine-tuning MMLoRA module to extract multimodal feature maps; Step 3: Based on the multimodal feature map and the parameters of the discretized state-space model SSM, obtain the output sequence. Based on the output sequence, projection layer, and residual connection, obtain the spatial augmentation feature map. Step 4: Convert the spatial enhancement feature map into a set of RGB multi-resolution feature maps and a set of depth multi-resolution feature maps. At each resolution scale, introduce an SE channel attention fusion unit to perform channel-level recalibration to dynamically suppress redundant noise in cross-modal interaction. Then, stitch and fuse the recalibrated feature maps at each scale to obtain a multi-scale fused feature map. Step 5: Input the multi-scale fused feature map into the decoding head to complete pixel-level dense classification, so that the target to be captured and the obstacle are completely separated from the complex background of the target scene, and output a semantic segmentation prediction map containing its precise contour and obstacle boundaries.

[0005] Further, step 1 includes: Step 11: Use the BM3D algorithm to perform geometric inpainting preprocessing on the original depth image to obtain the preprocessed depth image. Perform hard thresholding shrinkage operation in the three-dimensional transform domain to separate and remove speckle noise on the smooth surface of the preprocessed depth image, and preserve the edge geometric features of the object to the greatest extent. Step 12: For the remaining missing and void regions in the preprocessed depth image after denoising, construct a repair model based on the Navier-Stokes fluid dynamics equations. Treat the preprocessed depth image as an incompressible fluid and introduce time... Construct partial differential equations; Step 13: Based on the partial differential equation and the diffusion principle of iso-illuminance lines, guide the depth value and curvature of the edge of the damaged area to diffuse smoothly into the cavity area along the iso-illuminance lines, complete the repair of geometric and topological continuity, and output a high-fidelity depth image.

[0006] Furthermore, the expression for the partial differential equation is:

[0007] in, This is the preprocessed depth image. The gradient vector of the preprocessed depth image. for The orthogonal vectors, i.e., the directions of the isoluminescence lines, The Laplacian operator for the preprocessed depth image characterizes the smoothness of the depth.

[0008] Further, step 2 includes: Step 21: After downsampling the original RGB image and the high-fidelity depth image using PatchEmbedding, they are fed into a dual-stream encoder containing multiple stacked ViT blocks. The weights of the SAM pre-trained model are loaded, and the initial weights of the core ViT blocks are fixed. To preserve prior knowledge of images, text, and vision, only the parameters of the multimodal low-rank fine-tuning MMLoRA module are fine-tuned. Step 22: Embed a multimodal low-rank fine-tuning MMLoRA module in the bypass of the two-stream encoder to introduce a trainable low-rank decomposition matrix. and , obtained the Output features at each scale; Step 23: Apply low-rank decomposition matrices to the RGB branch and depth branch in the neural network respectively. and Independent extraction of the first The RGB modal and deep modal details of the output features at each scale are combined through learnable fusion coefficients to exchange cross-modal information and output a multimodal feature map.

[0009] Further, step 22 includes: The first layer is obtained through the linear forward calculation formula of each layer. Output features at each scale :

[0010] in, And represents the first Input features at each scale, This is a learnable scaling factor.

[0011] Further, step 3 includes: Step 31: Convert the multimodal feature map output by the encoder into a two-dimensional feature map through the input projection layer, and perform cross-scanning in four directions on the two-dimensional feature map: flatten the two-dimensional feature map into a one-dimensional sequence along the four independent directions of upper left to lower right, lower right to upper left, upper right to lower left, and lower left to upper right, to obtain a total of four one-dimensional sequences; Step 32: Introduce the parameters of the discretized state-space model SSM , and Parallel state-space modeling is performed on four one-dimensional sequences to obtain output sequences in four directions, where, As a data-dependent gating mechanism, it is used to dynamically control the preservation of current pixel features and the forgetting of historical states; Step 33: The output sequences in the four directions are spatially rearranged and added element by element. The results of the addition are fused through the output projection layer and residual connection to output a spatially enhanced feature map that strengthens the elongated object and the irregular edge structure. The elongated object is a strip-shaped target in the image with an aspect ratio greater than a set threshold. The irregular edge structure is the edge texture structure of an irregular geometric contour.

[0012] Further, step 4 includes: Step 41: Using a pyramid structure, the spatially enhanced feature map is independently expanded into an RGB multi-scale feature map set and a depth multi-scale feature map set, respectively; through convolution and deconvolution operations with different strides, the feature maps in the RGB multi-scale feature map set and the depth multi-scale feature map set are resampled into multiple-scale RGB multi-resolution feature map sets and depth multi-resolution feature map sets, respectively; the spatially enhanced feature map includes an RGB single-scale feature map and a depth single-scale feature map; Step 42: Introduce an SE channel attention fusion unit at each scale to extract the first SE channel attention fusion unit from the RGB multi-resolution feature map set and the depth multi-resolution feature map set, respectively. RGB feature map at each scale With depth feature map By adding element by element, we get the first... Multimodal joint features at various scales ; Step 43: Compress the first step using global average pooling. Multimodal joint features at various scales The generation dimension is global context descriptor Combined with the fully connected layer to generate the first Channel weight vector at each scale ; Step 44: Place the first Channel weight vector at each scale and multimodal joint features Multiply by each channel to obtain the first... Recalibration features at various scales The recalibrated features at each scale are spliced ​​together to obtain a multi-scale fused feature map.

[0013] Furthermore, the first Multimodal joint features at various scales The formula is:

[0014] in, It is in RGB mode. For depth modes; The following formula generates the first... Channel weight vector at each scale :

[0015] in, and The weight matrix for the fully connected layer, responsible for dimensionality reduction and dimensionality increase. Represents the ReLU activation function. This represents the Sigmoid activation function. For the first Multimodal global context descriptors generated by compression at various scales; No. Recalibration features at various scales The expression is: .

[0016] Furthermore, the decoding head consists of multiple convolutional and upsampling layers in the semantic segmentation prediction head network.

[0017] This application also discloses an RGB-D image semantic segmentation system based on spatial semantic enhancement, which implements the above-described method and includes: The image preprocessing module is used to acquire original RGB images and original depth images of the target scene from the same viewpoint and simultaneously collected. The original RGB images and the original depth images constitute an RGB-D image. Geometric restoration preprocessing is performed on the original depth images to obtain preprocessed depth images. Combining BM3D collaborative filtering and Navier-Stokes fluid dynamics equations, speckle noise in the preprocessed depth images is filtered out and missing parts are repaired to output a high-fidelity depth image. The original RGB images and the original depth images represent the target to be grasped and background obstacles in the target scene. The target to be grasped includes at least one of slender objects and irregular parts. The feature extraction module is used to construct a ViT dual-stream encoder based on the visual basic model SAM prior. It achieves cross-modal interaction between RGB features and depth features through the multimodal low-rank fine-tuning MMLoRA module to extract multimodal feature maps. The feature enhancement module is used to obtain the output sequence based on the multimodal feature map and the parameters of the discretized state-space model SSM, and to obtain the spatially enhanced feature map based on the output sequence, projection layer and residual connection. The feature fusion module is used to convert the spatial enhanced feature map into a set of RGB multi-resolution feature maps and a set of depth multi-resolution feature maps. At each resolution scale, an SE channel attention fusion unit is introduced to perform channel-level recalibration to dynamically suppress redundant noise in cross-modal interaction. The recalibrated feature maps at each scale are then stitched together to obtain a multi-scale fused feature map. The semantic segmentation module is used to input multi-scale fused feature maps into the decoding head to complete pixel-level dense classification, so that the target to be captured and the obstacles are completely separated from the complex background of the target scene, and outputs a semantic segmentation prediction map containing its precise contours and obstacle boundaries.

[0018] Due to the adoption of the above technical solution, this application has the following advantages: 1. This application combines BM3D collaborative filtering with Navier-Stokes hydrodynamic repair. BM3D accurately removes speckle noise, while the Navier-Stokes equations utilize the principle of iso-illuminance line diffusion to repair large-area voids. This approach maintains the curvature continuity of the object's surface without relying on large-scale data pre-training, providing highly pure geometric priors for the multimodal network from the data source.

[0019] 2. This application utilizes a ViT dual-stream encoder based on the visual foundation model SAM prior, breaking the limitations of the local receptive field in traditional CNNs and endowing the model with global topological understanding of unknown irregular objects. Combined with low-rank decomposition fine-tuning modules such as MMLoRA, dynamic interaction of cross-modal features is achieved while freezing the main weights of the large model, significantly reducing the number of training parameters.

[0020] 3. The feature enhancement module of this application utilizes a four-way cross-scan mechanism to apply the State-Space Model (SSM) to two-dimensional images. It can comprehensively aggregate all features with linear computational complexity. Figure 4 The global context information in each direction effectively solves the problem that slender objects and tiny edges are easily over-smoothed in the Transformer architecture, leading to undersegmentation.

[0021] 4. The feature fusion module in this application adaptively suppresses conflict noise in multimodal fusion through the SE attention mechanism. The overall network achieves an average intersection-to-union (MIoU) ratio of 75.46% on complex datasets such as Jacquard, and thanks to the extremely high operator parallelism of ViT and SS2D and the reparameterization during the LoRA inference period, the inference speed reaches 72.1 FPS, meeting the high-frequency real-time grasping requirements of industrial vision servo systems. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0023] Figure 1 This is a flowchart illustrating an RGB-D image semantic segmentation method based on spatial semantic enhancement, according to an embodiment of this application. Figure 2 This is a schematic diagram of the ViT dual-stream encoder based on the visual fundamental model SAM prior, according to an embodiment of this application. Figure 3This is a schematic diagram of the structure of the multimodal low-rank fine-tuning MMLoRA module in an embodiment of this application; Figure 4 This is a schematic diagram of the four-way cross-scanning mechanism of the feature enhancement module in an embodiment of this application; Figure 5 This is a schematic diagram of the feature fusion module in an embodiment of this application. Detailed Implementation

[0024] The present application will be further described in conjunction with the accompanying drawings and embodiments. The described embodiments are only a part of the embodiments of the present application, and not all of the embodiments. All other embodiments obtained by those skilled in the art should fall within the protection scope of the embodiments of the present application.

[0025] See Figure 1 This application provides an embodiment of an RGB-D image semantic segmentation method based on spatial semantic enhancement, which includes: Step 1: Acquire the original RGB image and original depth image of the target scene from the same viewpoint and simultaneously. The original RGB image and the original depth image constitute an RGB-D image. Perform geometric restoration preprocessing on the original depth image to obtain the preprocessed depth image. Combine BM3D collaborative filtering and Navier-Stokes fluid dynamics equations to filter out speckle noise and repair missing parts in the preprocessed depth image, outputting a high-fidelity depth image. The original RGB image and the original depth image represent the target to be grasped and background obstacles in the target scene. The target to be grasped includes at least one of slender objects and irregular parts. Step 2: Construct a ViT dual-stream encoder based on the visual foundation model SAM prior, and realize cross-modal interaction between RGB features and depth features through the multimodal low-rank fine-tuning MMLoRA module to extract multimodal feature maps; Step 3: Based on the multimodal feature map and the parameters of the discretized state-space model SSM, obtain the output sequence. Based on the output sequence, projection layer, and residual connection, obtain the spatial augmentation feature map. Step 4: Convert the spatial enhancement feature map into a set of RGB multi-resolution feature maps and a set of depth multi-resolution feature maps. At each resolution scale, introduce an SE (Squeeze-and-Excitation) channel attention fusion unit to perform channel-level recalibration to dynamically suppress redundant noise in cross-modal interaction. Then, stitch and fuse the recalibrated feature maps at each scale to obtain a multi-scale fused feature map. Step 5: Input the multi-scale fused feature map into the decoding head to complete pixel-level dense classification, so that the target to be captured and the obstacle are completely separated from the complex background of the target scene, and output a semantic segmentation prediction map containing its precise contour and obstacle boundaries.

[0026] Optionally, step 1 includes: Step 11: Use the BM3D algorithm to perform geometric inpainting preprocessing on the original depth image to obtain the preprocessed depth image. Perform hard thresholding shrinkage operation in the three-dimensional transform domain to separate and remove speckle noise on the smooth surface of the preprocessed depth image, and preserve the edge geometric features of the object to the greatest extent. Step 12: For the remaining missing and void regions in the preprocessed depth image after denoising, construct a repair model based on the Navier-Stokes fluid dynamics equations. Treat the preprocessed depth image as an incompressible fluid and introduce time... Construct partial differential equations; Step 13: Based on the partial differential equation and the diffusion principle of iso-illuminance lines, guide the depth value and curvature of the edge of the damaged area to diffuse smoothly into the cavity area along the iso-illuminance lines, complete the repair of geometric and topological continuity, and output a high-fidelity depth image.

[0027] Optionally, the expression for the partial differential equation is:

[0028] in, This is the preprocessed depth image. The gradient vector of the preprocessed depth image. for The orthogonal vectors, i.e., the directions of the isoluminescence lines, The Laplacian operator for the preprocessed depth image characterizes the smoothness of the depth.

[0029] Optionally, step 2 includes: Step 21: After downsampling the original RGB image and the high-fidelity depth image using PatchEmbedding, they are fed into a dual-stream encoder containing multiple stacked ViT blocks. The weights of the SAM pre-trained model are loaded, and the initial weights of the core ViT blocks are fixed. To preserve prior knowledge of images, text, and vision, only the parameters of the multimodal low-rank fine-tuning MMLoRA module are fine-tuned. Step 22: Embed a multimodal low-rank fine-tuning MMLoRA module in the bypass of the two-stream encoder to introduce a trainable low-rank decomposition matrix. and , obtained the Output features at each scale; Step 23: Apply low-rank decomposition matrices to the RGB branch and depth branch in the neural network respectively. and Independent extraction of the first The RGB modal and deep modal details of the output features at each scale are combined through learnable fusion coefficients to exchange cross-modal information and output a multimodal feature map.

[0030] Optionally, step 22 includes: The first layer is obtained through the linear forward calculation formula of each layer. Output features at each scale :

[0031] in, And represents the first Input features at each scale, This is a learnable scaling factor.

[0032] Optionally, step 3 includes: Step 31: Convert the multimodal feature map output by the encoder into a two-dimensional feature map through the input projection layer, and perform cross-scanning in four directions on the two-dimensional feature map: flatten the two-dimensional feature map into a one-dimensional sequence along the four independent directions of upper left to lower right, lower right to upper left, upper right to lower left, and lower left to upper right, to obtain a total of four one-dimensional sequences; Step 32: Introduce the parameters of the discretized state-space model SSM , and Parallel state-space modeling is performed on four one-dimensional sequences to obtain output sequences in four directions, where, As a data-dependent gating mechanism, it is used to dynamically control the preservation of current pixel features and the forgetting of historical states; Step 33: The output sequences in the four directions are spatially rearranged and added element by element. The results of the addition are fused through the output projection layer and residual connection to output a spatially enhanced feature map that strengthens the elongated object and the irregular edge structure. The elongated object is a strip-shaped target in the image with an aspect ratio greater than a set threshold. The irregular edge structure is the edge texture structure of an irregular geometric contour.

[0033] Optionally, step 4 includes: Step 41: Using a pyramid structure, the spatially enhanced feature map is independently expanded into an RGB multi-scale feature map set and a depth multi-scale feature map set, respectively; through convolution and deconvolution operations with different strides, the feature maps in the RGB multi-scale feature map set and the depth multi-scale feature map set are resampled into multiple-scale RGB multi-resolution feature map sets and depth multi-resolution feature map sets, respectively; the spatially enhanced feature map includes an RGB single-scale feature map and a depth single-scale feature map; Step 42: Introduce an SE channel attention fusion unit at each scale to extract the first SE channel attention fusion unit from the RGB multi-resolution feature map set and the depth multi-resolution feature map set, respectively. RGB feature map at each scale With depth feature map By adding element by element, we get the first... Multimodal joint features at various scales ; Step 43: Compress the first step using global average pooling. Multimodal joint features at various scales The generation dimension is global context descriptor Combined with the fully connected layer to generate the first Channel weight vector at each scale ; Step 44: Place the first Channel weight vector at each scale and multimodal joint features Multiply by each channel to obtain the first... Recalibration features at various scales The recalibrated features at each scale are spliced ​​together to obtain a multi-scale fused feature map.

[0034] Optionally, the first Multimodal joint features at various scales The formula is:

[0035] in, It is in RGB mode. For depth modes; The following formula generates the first... Channel weight vector at each scale :

[0036] in, and The weight matrix for the fully connected layer, responsible for dimensionality reduction and dimensionality increase. Represents the ReLU activation function. This represents the Sigmoid activation function. For the first Multimodal global context descriptors generated by compression at various scales; No. Recalibration features at various scales The expression is: .

[0037] Optionally, the decoding head consists of multiple convolutional and upsampling layers in a semantic segmentation prediction head network.

[0038] This application also provides an embodiment of an RGB-D image semantic segmentation system based on spatial semantic enhancement, which implements the method described in the above embodiment, and includes: The image preprocessing module is used to acquire original RGB images and original depth images of the target scene from the same viewpoint and simultaneously collected. The original RGB images and the original depth images constitute an RGB-D image. Geometric restoration preprocessing is performed on the original depth images to obtain preprocessed depth images. Combining BM3D collaborative filtering and Navier-Stokes fluid dynamics equations, speckle noise in the preprocessed depth images is filtered out and missing parts are repaired to output a high-fidelity depth image. The original RGB images and the original depth images represent the target to be grasped and background obstacles in the target scene. The target to be grasped includes at least one of slender objects and irregular parts. The feature extraction module is used to construct a ViT dual-stream encoder based on the visual basic model SAM prior. It achieves cross-modal interaction between RGB features and depth features through the multimodal low-rank fine-tuning MMLoRA module to extract multimodal feature maps. The feature enhancement module is used to obtain the output sequence based on the multimodal feature map and the parameters of the discretized state-space model SSM, and to obtain the spatially enhanced feature map based on the output sequence, projection layer and residual connection. The feature fusion module is used to convert the spatial enhanced feature map into a set of RGB multi-resolution feature maps and a set of depth multi-resolution feature maps. At each resolution scale, an SE channel attention fusion unit is introduced to perform channel-level recalibration to dynamically suppress redundant noise in cross-modal interaction. The recalibrated feature maps at each scale are then stitched together to obtain a multi-scale fused feature map. The semantic segmentation module is used to input multi-scale fused feature maps into the decoding head to complete pixel-level dense classification, so that the target to be captured and the obstacles are completely separated from the complex background of the target scene, and outputs a semantic segmentation prediction map containing its precise contours and obstacle boundaries.

[0039] This application innovatively combines BM3D collaborative filtering with the Navier-Stokes hydrodynamic repair algorithm. By utilizing the diffusion principle of the partial differential equation of iso-illuminance lines, it achieves depth map speckle denoising and large-area hole geometric curvature reconstruction without relying on additional neural network pre-training, thus solving the core pain point of missing geometric priors in multimodal networks from the data source.

[0040] The ViT dual-stream encoder designed in this application, based on the visual foundation model (SAM) prior, is the first to extract features by combining the multimodal low-rank fine-tuning (MMLoRA) mechanism under the premise of freezing the visual foundation model (SAM) parameters. It realizes dynamic weighted fusion of cross-modal features through bidirectional information flow interaction, effectively preserving large-scale general visual prior knowledge and significantly reducing the number of parameters for network training.

[0041] The feature enhancement module proposed in this application combines a state-space model with a four-way cross-scanning mechanism to serialize two-dimensional image features for long-distance context modeling, accurately capturing the boundary features of slender objects and tiny parts. This significantly suppresses the defect of excessive smoothing of local spatial details in the traditional Transformer architecture and improves the edge segmentation accuracy.

[0042] The feature fusion module constructed in this application combines a multi-resolution feature pyramid and an SE channel attention mechanism, and uses global average pooling and multilayer perceptron to dynamically generate channel weights, thereby achieving adaptive suppression and enhancement of RGB features and depth features at different resolutions, further improving the target discrimination and robustness in complex stacked scenarios.

[0043] Figure 2In this context, Layer Normalization (LN) is used to standardize the input features of the hidden layers in a neural network, helping to stabilize the training process and accelerate network convergence. Multi-head Attention (Multi-head Self-Attention) is a core component of the ViT (Visual Transformer) network, used to capture global contextual dependencies and long-distance interaction information of features. Adapter is a lightweight fine-tuning module inserted as a bypass when the weights of the main Visual Model (SAM) backbone are frozen, used to independently extract intra-modal (RGB or depth) detailed features. MLP (Multi-Layer Perceptron) is a network structure composed of fully connected layers used for non-linear transformation and feature mapping of the features output by the attention mechanism. MMAdapter (Multi-Modal Adapter) is deployed between two branches, using learnable coefficients to perform cross-modal information interaction and dynamic weighted fusion of RGB and depth features.

[0044] Figure 3 In the MLP framework, LoRA (Low-Rank Adaptation) is a low-rank fine-tuning network structure that approximates weight updates by introducing trainable low-rank decomposition matrices (A and B), independently extracting detailed features within either the RGB or deep modal with a very small number of parameters. MMLoRA (Multi-Modal Low-Rank Adaptation) is a multi-modal low-rank fine-tuning mechanism deployed between two-branch MLP layers, achieving cross-modal dynamic fusion of RGB and deep features through bidirectional information flow interaction.

[0045] Figure 4 In this context, projection is a linear mapping operation of features, corresponding to the input projection layer and output projection layer mentioned in steps 31 and 33, used for dimensionality transformation. Selection mechanism is a data-dependent gating mechanism that dynamically determines which current pixel features to retain and which historical state information to forget based on the input features. Discretization discretizes the continuous state-space model parameters to enable computation on discrete two-dimensional image sequence features. Figure 4 In the middle, x t h represents the input sequence features at the current time step. t-1 and h t These represent the hidden state at the previous time step and the current hidden state after the update, respectively; y t A represents the output sequence features at the current time step; B is the state transition matrix of the state-space model; t and C tThese are the input-dependent discretized projection matrices.

[0046] Figure 5 In the pyramid structure, convolution and deconvolution operations with different strides are used to independently expand and resample the RGB single-scale feature map and the depth single-scale feature map into sets of RGB multi-resolution feature maps and sets of depth multi-resolution feature maps with different resolutions (such as 1 / 4, 1 / 8, 1 / 16, 1 / 32). , , , as well as , , , After pyramid structure expansion, RGB feature maps and depth feature maps are generated at four different scales. SE channel attention fusion unit: At each corresponding scale, the bimodal feature maps are first added to obtain multimodal joint features, and then channel weight vectors are generated through global average pooling (squeezing) and fully connected layers (excitation). Finally, the channel-level recalibration of the feature maps is completed to dynamically suppress redundant noise in cross-modal interactions. After channel recalibration is completed at four scales, the final multi-scale fused feature map is spliced ​​and output.

[0047] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and not to limit them. Although this application has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of this application. Any modifications or equivalent substitutions that do not depart from the spirit and scope of this application should be covered within the protection scope of the claims of this application.

Claims

1. A semantic segmentation method for RGB-D images based on spatial semantic enhancement, characterized in that, include: Step 1: Acquire the original RGB image and original depth image of the target scene from the same viewpoint and simultaneously. The original RGB image and the original depth image constitute an RGB-D image. Perform geometric restoration preprocessing on the original depth image to obtain the preprocessed depth image. Combine BM3D collaborative filtering and Navier-Stokes fluid dynamics equations to filter out speckle noise and repair missing parts in the preprocessed depth image, outputting a high-fidelity depth image. The original RGB image and the original depth image represent the target to be grasped and background obstacles in the target scene. The target to be grasped includes at least one of slender objects and irregular parts. Step 2: Construct a ViT dual-stream encoder based on the visual foundation model SAM prior, and realize cross-modal interaction between RGB features and depth features through the multimodal low-rank fine-tuning MMLoRA module to extract multimodal feature maps; Step 3: Based on the multimodal feature map and the parameters of the discretized state-space model SSM, obtain the output sequence. Based on the output sequence, projection layer, and residual connection, obtain the spatial augmentation feature map. Step 4: Convert the spatial enhancement feature map into a set of RGB multi-resolution feature maps and a set of depth multi-resolution feature maps. At each resolution scale, introduce an SE channel attention fusion unit to perform channel-level recalibration to dynamically suppress redundant noise in cross-modal interaction. Then, stitch and fuse the recalibrated feature maps at each scale to obtain a multi-scale fused feature map. Step 5: Input the multi-scale fused feature map into the decoding head to complete pixel-level dense classification, so that the target to be captured and the obstacle are completely separated from the complex background of the target scene, and output a semantic segmentation prediction map containing its precise contour and obstacle boundaries; Step 2 includes: Step 21: After downsampling the original RGB image and the high-fidelity depth image using PatchEmbedding, they are fed into a dual-stream encoder containing multiple stacked ViT blocks. The weights of the SAM pre-trained model are loaded, and the initial weights of the core ViT blocks are fixed. To preserve prior knowledge of images, text, and vision, only the parameters of the multimodal low-rank fine-tuning MMLoRA module are fine-tuned. Step 22: Embed a multimodal low-rank fine-tuning MMLoRA module in the bypass of the two-stream encoder to introduce a trainable low-rank decomposition matrix. and , obtained the Output features at each scale; Step 23: Apply low-rank decomposition matrices to the RGB branch and depth branch in the neural network respectively. and Independent extraction of the first The RGB modal and deep modal details of the output features at each scale are combined through learnable fusion coefficients to exchange cross-modal information and output a multimodal feature map. Step 3 includes: Step 31: Convert the multimodal feature map output by the encoder into a two-dimensional feature map through the input projection layer, and perform cross-scanning in four directions on the two-dimensional feature map: flatten the two-dimensional feature map into a one-dimensional sequence along the four independent directions of upper left to lower right, lower right to upper left, upper right to lower left, and lower left to upper right, to obtain a total of four one-dimensional sequences; Step 32: Introduce the parameters of the discretized state-space model SSM , and Parallel state-space modeling is performed on four one-dimensional sequences to obtain output sequences in four directions, where, As a data-dependent gating mechanism, it is used to dynamically control the preservation of current pixel features and the forgetting of historical states; and These are the input-dependent discretized projection matrices; Step 33: The output sequences in the four directions are spatially rearranged and added element by element. The results of the addition are fused through the output projection layer and residual connection to output a spatially enhanced feature map that strengthens the elongated object and the irregular edge structure. The elongated object is a strip-shaped target in the image with an aspect ratio greater than a set threshold. The irregular edge structure is the edge texture structure of an irregular geometric contour.

2. The method according to claim 1, characterized in that, Step 1 includes: Step 11: Use the BM3D algorithm to perform geometric inpainting preprocessing on the original depth image to obtain the preprocessed depth image. Perform hard thresholding shrinkage operation in the three-dimensional transform domain to separate and remove speckle noise on the smooth surface of the preprocessed depth image, and preserve the edge geometric features of the object to the greatest extent. Step 12: For the remaining missing and void regions in the preprocessed depth image after denoising, construct a repair model based on the Navier-Stokes fluid dynamics equations. Treat the preprocessed depth image as an incompressible fluid and introduce time... Construct partial differential equations; Step 13: Based on the partial differential equation and the diffusion principle of iso-illuminance lines, guide the depth value and curvature of the edge of the damaged area to diffuse smoothly into the cavity area along the iso-illuminance lines, complete the repair of geometric and topological continuity, and output a high-fidelity depth image.

3. The method according to claim 2, characterized in that, The expression for the partial differential equation is: in, This is the preprocessed depth image. The gradient vector of the preprocessed depth image. for The orthogonal vectors, i.e., the directions of the isoluminescence lines, The Laplacian operator for the preprocessed depth image characterizes the smoothness of the depth.

4. The method according to claim 1, characterized in that, Step 22 includes: The first layer is obtained through the linear forward calculation formula of each layer. Output features at each scale : in, And represents the first Input features at each scale, This is a learnable scaling factor.

5. The method according to claim 1, characterized in that, Step 4 includes: Step 41: Using a pyramid structure, the spatially enhanced feature map is independently expanded into an RGB multi-scale feature map set and a depth multi-scale feature map set, respectively; through convolution and deconvolution operations with different strides, the feature maps in the RGB multi-scale feature map set and the depth multi-scale feature map set are resampled into multiple-scale RGB multi-resolution feature map sets and depth multi-resolution feature map sets, respectively; the spatially enhanced feature map includes an RGB single-scale feature map and a depth single-scale feature map; Step 42: Introduce an SE channel attention fusion unit at each scale to extract the first SE channel attention fusion unit from the RGB multi-resolution feature map set and the depth multi-resolution feature map set, respectively. RGB feature map at each scale With depth feature map By adding element by element, we get the first... Multimodal joint features at various scales ; Step 43: Compress the first step using global average pooling. Multimodal joint features at various scales The generation dimension is global context descriptor Combined with the fully connected layer to generate the first Channel weight vector at each scale ; Step 44: Place the first Channel weight vector at each scale and multimodal joint features Multiply by each channel to obtain the first... Recalibration features at various scales The recalibrated features at each scale are spliced ​​together to obtain a multi-scale fused feature map.

6. The method according to claim 5, characterized in that, The first Multimodal joint features at various scales The formula is: in, It is in RGB mode. For depth modes; The following formula generates the first... Channel weight vector at each scale : in, and The weight matrix responsible for dimensionality reduction and dimensionality increase in the fully connected layer. Represents the ReLU activation function. This represents the Sigmoid activation function. For the first Multimodal global context descriptors generated by compression at various scales; No. Recalibration features at various scales The expression is: 。 7. The method according to claim 1, characterized in that, The decoding head consists of multiple convolutional and upsampling layers in the semantic segmentation prediction head network.

8. A semantic segmentation system for RGB-D images based on spatial semantic enhancement, implementing the method described in any one of claims 1-7, characterized in that, include: The image preprocessing module is used to acquire original RGB images and original depth images of the target scene from the same viewpoint and simultaneously collected. The original RGB images and the original depth images constitute an RGB-D image. Geometric restoration preprocessing is performed on the original depth images to obtain preprocessed depth images. Combining BM3D collaborative filtering and Navier-Stokes fluid dynamics equations, speckle noise in the preprocessed depth images is filtered out and missing parts are repaired to output a high-fidelity depth image. The original RGB images and the original depth images represent the target to be grasped and background obstacles in the target scene. The target to be grasped includes at least one of slender objects and irregular parts. The feature extraction module is used to construct a ViT dual-stream encoder based on the visual basic model SAM prior. It achieves cross-modal interaction between RGB features and depth features through the multimodal low-rank fine-tuning MMLoRA module to extract multimodal feature maps. The feature enhancement module is used to obtain the output sequence based on the multimodal feature map and the parameters of the discretized state-space model SSM, and to obtain the spatially enhanced feature map based on the output sequence, projection layer and residual connection. The feature fusion module is used to convert the spatial enhanced feature map into a set of RGB multi-resolution feature maps and a set of depth multi-resolution feature maps. At each resolution scale, an SE channel attention fusion unit is introduced to perform channel-level recalibration to dynamically suppress redundant noise in cross-modal interaction. The recalibrated feature maps at each scale are then stitched together to obtain a multi-scale fused feature map. The semantic segmentation module is used to input multi-scale fused feature maps into the decoding head to complete pixel-level dense classification, so that the target to be captured and the obstacles are completely separated from the complex background of the target scene, and outputs a semantic segmentation prediction map containing its precise contours and obstacle boundaries.

Citation Information

Patent Citations

  • Transparent object mechanical arm grabbing method and system based on semantic geometry fusion perception

    CN121157036A

  • Camera and laser radar calibration method and system based on multistage feature fusion

    CN121810815A