DSSMNet-based binocular speckle structured light parallax estimation method
Through the end-to-end stereo matching learning framework of DSSMNet and the improved loss function, combined with the dense connection of spatial pyramid and CBAM attention mechanism, the accuracy problem of disparity estimation in speckle textureless areas is solved, and high-precision disparity estimation effect is achieved.
Patent Information
- Application Number
- CN202410436846.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-11
- Publication Date
- 2025-10-21
AI Technical Summary
Existing deep learning networks have weak prediction effects when processing disparity estimation of structured light images in speckle and textureless areas, and mainly focus on monocular disparity estimation, lacking effective binocular disparity estimation methods.
An end-to-end stereo matching learning framework based on DSSMNet is adopted, combined with the dense connection structure of the spatial pyramid and the CBAM attention mechanism, and trained using the improved Smooth L1-Sobel loss function to generate multi-scale feature representation and adaptive weighted feature extraction to improve the accuracy of disparity estimation.
High-precision disparity estimation is achieved in speckle-texture-free areas, which enhances feature expression and edge perception capabilities, reduces the influence of noise and redundant information, and improves the accuracy of disparity estimation.
Smart Images

Figure CN120823253A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer three-dimensional vision technology, and more specifically, relates to a binocular speckle structured light parallax estimation method based on DSSMNet. Background Art
[0002] Three-dimensional object reconstruction is one of the most important and challenging problems in computer vision. Structured light 3D measurement technology is a key method for optical 3D measurement. Its advantages include high speed and precision, and it holds considerable application prospects in areas such as robot guidance, virtual reality, human-computer interaction, cultural heritage conservation, robotic vision, and biomedicine. Estimating the disparity of an object from binocular speckle structured light images is a key component of 3D measurement technology.
[0003] In recent years, deep learning has made significant progress in computer vision and has been successfully applied to disparity estimation tasks. Many deep learning networks have brought significant improvements to stereo matching, resulting in the emergence of numerous excellent stereo matching networks. Among them, MC-CNN first uses a convolutional neural network to calculate the matching cost, and then uses traditional algorithms for the remaining steps. PSMNet improves matching accuracy through spatial pyramid pooling in the feature extraction module, concatenating feature maps from different levels as the final feature map. ACVNet proposes a novel cost volume construction method that generates attention weights based on correlation cues from patch matching, thereby enhancing the representational power of the cost volume. However, existing deep learning networks have some shortcomings in binocular disparity estimation from structured light images, particularly in areas with speckle and textureless regions. Furthermore, current methods for structured light disparity estimation primarily focus on monocular disparity estimation, such as PCTNet. To improve disparity estimation accuracy in areas with speckle and textureless regions, this study proposes a binocular speckle structured light disparity estimation method based on DSSMNet.
[0004] In view of this, the present invention aims to propose a binocular speckle structured light disparity estimation method based on DSSMNet, which is used to complete the disparity estimation task of binocular speckle structured light images, and achieves high-precision disparity estimation results in the matching problem of speckle textureless areas.
[0005] To achieve the above object, the present invention proposes a binocular speckle structured light disparity estimation method based on DSSMNet, comprising the following steps:
[0006] S1: Develop an end-to-end stereo matching learning framework for binocular speckle structured light disparity estimation without any post-processing;
[0007] S2: In the feature extraction stage, a spatial pyramid dense connection structure is established. The spatial pyramid dense connection structure is used to extract features from the structured light image, generating a feature map with multi-scale feature representation and rich feature information. The CBAM attention mechanism is then used to select the most useful information from the feature map.
[0008] S3: Use the simulated binocular speckle structured light image to train the DSSMNet disparity estimation network model, and use the improved Smooth L1-Sobel loss function to update the DSSMNet disparity estimation network during the training process to obtain a trained network model;
[0009] S4: Use the trained network model to estimate the disparity of the simulated binocular speckle structured light image.
[0010] Furthermore, the densely connected structure of the spatial pyramid in step S2 is specifically:
[0011] The SPP pyramid module and the dense connection module in DenseNet are processed in parallel, the outputs of the two modules and the output of the first backbone part are spliced, 1×1 and 3×3 convolution operations are performed, and the dimensions are spliced again. In this splicing process, the idea of the improved SPP module and the SPC module are combined.
[0012] Furthermore, the CBAM attention mechanism in step S2 is specifically as follows:
[0013] The combination of channel attention and spatial attention, first through the channel attention mechanism, the input feature map is subjected to global average pooling and global maximum pooling to obtain a one-dimensional feature vector respectively; the two feature vectors pass through a weight-sharing MLP, and then the weights are added, and finally the channel attention mechanism M is obtained through the sigmoid activation function. c The above process can be expressed as:
[0014]
[0015] Among them, σ represents the sigmoid activation function, W0 and W1 represent two convolution operations respectively, and represent average pooling and maximum pooling respectively.
[0016] Then, through the spatial attention mechanism, the enhanced feature map obtained by the channel attention mechanism is subjected to global average pooling and global maximum pooling in the channel dimension to obtain two feature maps. The two feature maps are then concat, and a 7×7 convolution kernel is used for convolution operation. Finally, the spatial attention vector M is obtained through the sigmoid activation function. sFinally, the importance of the position is multiplied by the enhanced feature map element by element to obtain the final feature map. The above process can be expressed as:
[0017]
[0018] Among them, f 7×7 Represents a convolution operation with a convolution kernel size of 7×7.
[0019] Furthermore, the SPP pyramid module is specifically:
[0020] Four fixed-size average pooling blocks of 64×64, 32×32, 16×16, and 8×8 are used to compress the features into four scales, followed by 1×1 convolution dimensionality reduction and upsampling operations.
[0021] Furthermore, the densely connected structure is specifically:
[0022] It consists of three dense blocks, each of which generates features of different resolutions. The dense blocks are connected together through Transition, and the outputs of the three dense blocks are upsampled and spliced into the final feature map.
[0023] Furthermore, the SPP-improved SPPCSPC structure is specifically:
[0024] The pooled outputs are concatenated, passed through 1×1 and 3×3 convolutional layers, and then concatenated again.
[0025] Furthermore, the dense block is specifically:
[0026] The features in each dense block have the same size but different number of channels. Dense connections are used within the block, that is, each layer takes all previous feature maps as input. Each group of layers in the block consists of a 1×1 convolution and a 3×3 convolution, specifically BN-ReLU-Conv(1×1)-BN-ReLU-Conv(3×3).
[0027] Furthermore, the Transition is specifically:
[0028] It is composed of a 2×2 average pooling and a 1×1 convolutional layer, specifically BN+ReLU+1×1 Conv+2×2AvgPooling.
[0029] Furthermore, the improved Smooth L1-Sobel loss function in step S3 is specifically:
[0030] The combination of the Smooth L1 loss function and the edge Sobel loss function, the Smooth L1 loss function can be defined as:
[0031]
[0032] Where x is the difference between the true value and the predicted value;
[0033] In the edge Sobel loss function, a 3×3 Sobel operator is used to perform convolution operations in the x and y directions respectively to obtain the horizontal and vertical gradients of the image and realize the calculation of gradient loss;
[0034] Combining the Smooth L1 loss function and the edge Sobel loss function can be expressed as:
[0035] SmoothL1-Sobel=w1×SmoothL1+w2×Sobel
[0036] Among them, w1 and w2 are the weights of Smooth L1 and Sobel respectively.
[0037] Compared with the existing technology, the binocular speckle structured light disparity estimation method based on DSSMNet described in the present invention has the following advantages:
[0038] (1) The present invention combines the spatial pyramid module, which can extract features from different scales and generate fixed-size feature vectors, with the dense connection module, which improves the ability of feature transfer through dense connections. The two modules process in parallel, can extract features at different scales at the same time, and can make full use of dense connections to enhance the transferability of features, thereby better capturing multi-scale information in the image.
[0039] (2) The present invention applies the CBAM attention mechanism that combines channel attention and spatial attention in the feature extraction stage. By adaptively weighting the features, useful channel information in the feature representation is extracted, thereby reducing the impact of noise and redundant information and enhancing the expressiveness of the features.
[0040] (3) In terms of loss function, the present invention uses a loss function that combines the Smooth L1 loss function and the edge Sobel loss function. The Smooth L1 loss function has better smoothness when processing outliers, which can effectively reduce the influence of outliers. The edge Sobel loss function is also used to enhance the model's perception of image edges, making the generated results closer to the real edges. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The accompanying drawings, which constitute part of the present invention, are provided to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are provided to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:
[0042] Figure 1 This is a flow chart of a binocular speckle structured light disparity estimation method based on DSSMNet of the present invention;
[0043] Figure 2 This is the structural diagram of the densely connected spatial pyramid structure of the present invention (excluding CBAM attention);
[0044] Figure 3 CBAM attention mechanism diagram of the present invention;
[0045] Figure 4 This is the structural diagram of the spatial pyramid dense connection structure of the present invention (including CBAM attention);
[0046] Figure 5 The speckle structured light image inputted by the present invention;
[0047] Figure 6 is the disparity map label of the present invention;
[0048] Figure 7 A disparity map predicted by the present invention; DETAILED DESCRIPTION
[0049] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0050] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and the best implementation method.
[0051] The present invention proposes a binocular speckle structured light disparity estimation method based on DSSMNet. The present invention is described in more detail below with reference to the accompanying drawings and specific embodiments.
[0052] In this embodiment, the following steps are included:
[0053] Step 1: Prepare the dataset and divide it into training set and test set in the ratio of 8:2. Figure 4 The speckle pattern shown, the label image is attached Figure 5 The disparity map shown.
[0054] Step 2: Input the binocular speckle structured light image, pass through three 3×3 convolutions and four basic residual blocks, and then pass through the attached Figure 2 The spatial pyramid dense connection structure shown in the figure mainly includes the spatial pyramid module and the dense connection module. Figure 3 The CBAM attention mechanism shown combines channel attention and spatial attention.
[0055] The spatial pyramid dense connection structure uses the SPP pyramid module and the dense connection module in DenseNet for simultaneous parallel processing, splicing the outputs of the two modules with the output of the first backbone part, then performing 1×1 and 3×3 convolution operations, and then splicing in dimension again. In this splicing process, the idea of the improved SPP module SPPCSPC module is combined.
[0056] The CBAM attention mechanism includes channel attention and spatial attention.
[0057] Channel attention: First, the input feature layer x is compressed through maximum pooling and average pooling operations, so that it becomes a feature map of size (1, 1, C) in the channel dimension, where C is the number of channels. Secondly, the compressed feature map is passed to the shared fully connected layer (MLP). This fully connected layer has two convolutional layers, which compress the number of channels to 1 / reduction times the original number of channels and then expand it back to the original number of channels, where reduction is the ratio of reducing the channel dimension. The output of the fully connected layer is activated by the ReLU activation function. The sigmoid operation is performed on the activated result to limit the result between 0 and 1. Finally, the channel attention weight is multiplied by the input feature layer x to obtain the weighted feature output.
[0058] Spatial Attention: First, the output x of the channel attention is max-pooled and average-pooled to convert it into a feature map of size (1, H, W) in the channel dimension, where H and W are the height and width of the input feature layer x, respectively. Secondly, the results of max-pooling and average-pooling are concatenated along the channel dimension to form a feature map of size (2, H, W). This feature map is then passed through a convolutional layer to compress it into a feature map of size (1, H, W). The compressed feature map is activated using the sigmoid function, limiting the result to between 0 and 1. Finally, the weights of the spatial attention are multiplied by the weighted feature output to obtain the final output features.
[0059] Step 3: The present invention uses an Nvidia RTX 3090 (24GB) graphics card for experiments. The model code is based on Pytorch. The batch size used for training is 3, the initial learning rate is 0.001, the total training rounds are 100, and the optimizer uses Adam. The loss function used in the entire training network is an improved Smooth L1-Sobel loss function, which is a combination of the Smooth L1 loss function and the Sobel edge loss function with a weight of 1:0.0015.
[0060] Step 4: Test the performance of the model on the test set and output the corresponding disparity map. The output disparity map is as shown in the attached figure. Figure 6The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.
Claims
1. A binocular speckle structured light disparity estimation method based on DSSMNet, characterized by: The steps include: S1: Develop an end-to-end stereo matching learning framework for binocular speckle structured light disparity estimation without any post-processing; S2: In the feature extraction stage, a spatial pyramid dense connection structure is established. The spatial pyramid dense connection structure is used to extract features from the structured light image, generating a feature map with multi-scale feature representation and rich feature information. The CBAM attention mechanism is then used to select the most useful information from the feature map. S3: Use the simulated binocular speckle structured light image to train the DSSMNet disparity estimation network model, and use the improved Smooth L1-Sobel loss function to update the DSSMNet disparity estimation network during the training process to obtain a trained network model; S4: Use the trained network model to estimate the disparity of the simulated binocular speckle structured light image.
2. The method for binocular speckle structured light disparity estimation based on DSSMNet according to claim 1, characterized in that: The spatial pyramid dense connection structure is specifically as follows: the SPP pyramid module and the dense connection module in DenseNet are processed in parallel, the outputs of the two modules and the output of the first backbone part are spliced, 1×1 and 3×3 convolution operations are performed, and dimensional splicing is performed again. In this splicing process, the idea of the improved SPP module and the SPC module is combined.
3. The method for binocular speckle structured light disparity estimation based on DSSMNet according to claim 1, characterized in that: The CBAM attention mechanism is specifically: a combination of channel attention and spatial attention. First, the input feature map is subjected to global average pooling and global maximum pooling through the channel attention mechanism to obtain two one-dimensional feature vectors; The two feature vectors pass through a weight-sharing MLP, then the weights are added, and finally the channel attention mechanism M is obtained through the sigmoid activation function. c The above process can be expressed as: Among them, σ represents the sigmoid activation function, W0 and W1 represent two convolution operations respectively, and Represents average pooling and maximum pooling respectively; Then, through the spatial attention mechanism, the enhanced feature map obtained by the channel attention mechanism is subjected to global average pooling and global maximum pooling in the channel dimension to obtain two feature maps. The two feature maps are then concat-operated and convolved using a 7×7 convolution kernel. Finally, the spatial attention vector M is obtained through the sigmoid activation function. s Finally, the importance of the position is multiplied by the enhanced feature map element by element to obtain the final feature map. The above process can be expressed as: Among them, f 7×7 Represents a convolution operation with a convolution kernel size of 7×7.
4. The method for binocular speckle structured light disparity estimation based on DSSMNet according to claim 2, characterized in that: The SPP pyramid module specifically uses four fixed-size average pooling blocks of 64×64, 32×32, 16×16, and 8×8 to compress features into four scales, and then performs 1×1 convolution dimensionality reduction and upsampling operations.
5. The method for binocular speckle structured light disparity estimation based on DSSMNet according to claim 2, characterized in that: The dense connection structure is specifically composed of three dense blocks, each dense block generates features of different resolutions, and the dense blocks are connected together through Transition. The outputs of the three dense blocks are then upsampled and spliced into the final feature map.
6. The method for binocular speckle structured light disparity estimation based on DSSMNet according to claim 2, characterized in that: The SPP-improved SPPCSPC structure is specifically as follows: after splicing the pooled outputs, the outputs are passed through 1×1 and 3×3 convolutional layers and then spliced again.
7. The method for binocular speckle structured light disparity estimation based on DSSMNet according to claim 5, characterized in that: The dense blocks are specifically as follows: the features in each dense block have the same size, but different numbers of channels. Dense connections are used within the block, that is, each layer takes all previous feature maps as input, and each group of layers in the block consists of a 1×1 convolution and a 3×3 convolution, specifically BN-ReLU-Conv(1×1)-BN-ReLU-Conv(3×3).
8. The method for binocular speckle structured light disparity estimation based on DSSMNet according to claim 5, characterized in that: The Transition is specifically composed of a 2×2 average pooling and a 1×1 convolutional layer, specifically BN+ReLU+1×1Conv+2×2AvgPooling.
9. The method for binocular speckle structured light disparity estimation based on DSSMNet according to claim 1, characterized in that: The improved Smooth L1-Sobel loss function is specifically a combination of the Smooth L1 loss function and the edge Sobel loss function. The Smooth L1 loss function can be defined as: Where x is the difference between the true value and the predicted value; In the edge Sobel loss function, a 3×3 Sobel operator is used to perform convolution operations in the x and y directions respectively to obtain the horizontal and vertical gradients of the image and realize the calculation of gradient loss; Combining the Smooth L1 loss function and the edge Sobel loss function can be expressed as: Smooth L1-Sobel=w1×Smooth L1+w2×Sobel Among them, w1 and w2 are the weights of Smooth L1 and Sobel respectively.