A Stereo Matching Method Based on Mamba Cost Volume
By introducing residual visual mamba layer and 3D attention module into the cost body building module of the stereo matching network, the problem of poor robustness of the prior art when dealing with occlusion, weak texture and reflective areas is solved, and higher matching accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510300947.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-14
AI Technical Summary
Existing deep learning-based stereo matching networks are poorly robust when dealing with occlusion, weak textures and reflective areas, making it difficult to obtain accurate and reliable results.
A stereo matching method based on Mamba cost body is proposed. By introducing residual visual Mamba layer and 3D attention module into the cost body construction module, the relationship between deep and shallow features extracted during feature extraction is fully explored, and combined with local enhancement, the cost body provides richer context information.
The prediction ability of the stereo matching network in occluded and weak textured areas is improved, and the overall matching accuracy and robustness are enhanced.
Smart Images

Figure CN119832045B_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to the field of computer vision technology. More specifically, this application relates to a stereo matching method based on Mamba cost volume. Background Art
[0002] Stereo matching is a key task in computer vision technology and is often applied to different field scenarios such as autonomous driving, 3D reconstruction, and augmented reality. Stereo matching extracts depth information from two or more images with different perspectives and then restores the three-dimensional structure of the scene. Its basic principle is to use the disparity, that is, the displacement of the same object under different perspectives, to infer the three-dimensional structure of the scene.
[0003] Currently, stereo matching methods are divided into traditional stereo matching methods and deep learning-based stereo matching methods. Traditional stereo matching methods are divided into four steps: matching cost calculation, cost aggregation, disparity calculation, and disparity refinement. However, traditional stereo matching methods usually rely on manually designed basic features, such as pixel gray value differences or Census transforms, when calculating the matching cost. This limits their ability to extract global information and integrate multi-scale features. Therefore, they have poor robustness when dealing with occlusion, weak texture, and reflective regions and it is difficult to obtain accurate and reliable results in practical applications. The deep learning network architecture for stereo matching can generally be divided into four parts: feature extraction, cost volume construction, cost aggregation, and disparity regression. Since convolutional neural networks have powerful feature learning capabilities, they can significantly improve the prediction accuracy of the network and at the same time show great potential in solving the problem that the above traditional stereo matching methods have poor robustness when dealing with occlusion, weak texture, and reflective regions and it is difficult to obtain accurate and reliable results in practical applications.
[0004] As an example, three deep learning-based stereo matching networks in the prior art are given, namely PSMNet (Pyramid Stereo Matching Network), GwcNet (Group-wise Cost Volumes Network), and ACVNet (Attention-Cost Volume Network). Among them, PSMNet integrates global information into the concatenated cost volume, enabling the model to capture the matching relationships between distant pixels, thereby improving its performance in complex scenes, especially when dealing with occlusions and texture-rich regions. However, the generation of the cost volume in PSMNet relies on the traditional concatenation method. Although this method can effectively capture global information, it may ignore the fine-grained relationships between local pixels. Therefore, in some special cases, such as regions with scarce textures, PSMNet may face certain challenges. GwcNet provides a similarity measure for the cost volume by performing dot product operations on grouped features, improving the efficiency and accuracy of cost volume generation. However, this grouped dot product operation relies more on mathematical experience to measure similarity and may not be able to fully capture the potential relationships in the features. In particular, for those regions where there are subtle connections between local and global features, this method may not be able to fully exploit these latent clues, thus affecting its performance in weak texture or complex scenes. ACVNet constructs attention weights for enhancing the concatenated cost volume by generating group correlation cost volumes based on adaptive patch matching, improving the accuracy of the model. However, the group correlation cost volume uses mathematical experience to measure this similarity and cannot fully utilize the latent clues in the features. These clues may be the relationships between this pixel and its surrounding pixels, or the connections between pixels that are farther apart.
[0005] Based on the above content, it can be known that although the current deep learning-based stereo matching networks have made many innovations in improving accuracy and dealing with complex scenes, there is still room for improvement in fully exploring the potential feature relationships in images. In particular, there is still potential in how to utilize the local relationships between pixels and how to establish stronger connections between farther pixels. If such information can be fully explored, it will be beneficial to improve the prediction ability of the network in occlusion and weak texture regions, thereby improving the overall matching accuracy and robustness.
[0006] In view of this, there is an urgent need to provide a stereo matching method based on Mamba cost volume to improve the matching accuracy and robustness of stereo matching. Summary of the Invention
[0007] In order to solve at least one or more of the above-mentioned technical problems, the present application proposes a stereo matching method based on Mamba cost volume in multiple aspects.
[0008] The present application provides a stereo matching method based on a Mamba cost volume, including: obtaining a left-view image and a right-view image to be processed, where the left-view image and the right-view image are images of the same object from different perspectives; inputting the left-view image and the right-view image into an improved stereo matching network to obtain a target predicted disparity map; where the improved stereo matching network at least includes a feature extraction module, a cost volume construction module, an aggregation module, and a disparity regression module connected in series in sequence, and the feature extraction module is also connected in series with the disparity regression module; the cost volume construction module includes an initial cost volume construction module, a residual vision Mamba layer, a 3D attention module and a normalization module respectively connected in series with the output of the residual vision Mamba layer, a first fusion module connected in series with the output of the 3D attention module and the output of the normalization module, and a second fusion module connected in series with the first fusion module.
[0009] Optionally, the left-view image and the right-view image share the weight parameters of the feature extraction module, and the feature extraction module is a pre-trained Efficient-B3 network; the left feature maps and right feature maps output by the 4th convolutional layer, 5th convolutional layer, 6th convolutional layer, and 8th convolutional layer in the feature extraction module are used as the inputs of the cost volume construction module connected in series with the feature extraction module, and the left feature maps and right feature maps output by the 4th convolutional layer in the feature extraction module are used as the inputs of the disparity regression module connected in series with the feature extraction module; the sizes of the left feature maps and right feature maps output by the 4th convolutional layer, 5th convolutional layer, 6th convolutional layer, and 8th convolutional layer are 1 / 2, 1 / 4, 1 / 8, and 1 / 16 of the left-view image and the right-view image respectively, and the number of channels are 24, 32, 48, and 132 respectively.
[0010] Optionally, the initial cost volume construction module at least sequentially performs the following steps: removing the pixel values of the left i columns and right i columns in a group of input left feature maps and right feature maps, where i is the disparity level of the left feature maps and the right feature maps; splicing the left feature maps and right feature maps after pixel value removal in the channel dimension in an interleaved superposition manner to obtain an initial cost volume; raising the dimension of the initial cost volume so that the dimension of the initial cost volume after dimension raising meets the input dimension requirements of the residual vision Mamba layer, where the input dimension requirements of the residual vision Mamba layer are four-dimensional, and the output number of channels of the residual vision Mamba layer is C / 4.
[0011] Optionally, the residual vision Mamba layer at least sequentially performs the following steps: performing layer normalization on the initial cost volume output by the initial cost volume construction module; inputting the layer-normalized initial cost volume into the Mamba model; adding learnable parameters to the initial cost volume output by the Mamba model pixel by pixel; performing layer normalization on the initial cost volume after pixel-by-pixel addition; and inputting the layer-normalized initial cost volume into a linear layer.
[0012] Optionally, the 3D attention module at least sequentially performs the following steps: performing max pooling and average pooling on the initial cost volume output by the residual vision Mamba layer respectively; adding the max-pooled initial cost volume and the average-pooled initial cost volume pixel by pixel; performing first Sigmoid normalization on the initial cost volume after pixel-by-pixel addition; performing first pixel-by-pixel multiplication on the initial cost volume output by the residual vision Mamba layer and the initial cost volume after the first Sigmoid normalization; performing at least 3 3D convolutions on the initial cost volume after the first pixel-by-pixel multiplication, where the kernel sizes of the 3 3D convolutions are 1x7x7, 7x1x1, and 7x7x7 in sequence, acting on the spatial, depth, and overall dimensions respectively; performing second Sigmoid normalization on the initial cost volume after at least 3 3D convolutions; and performing second pixel-by-pixel multiplication on the initial cost volume after the second Sigmoid normalization and the initial cost volume after the first pixel-by-pixel addition.
[0013] The normalization module at least sequentially performs the following steps: performing layer normalization on the initial cost volume output by the residual vision Mamba layer; weighting the initial cost volume output by the residual vision Mamba layer based on learnable parameters; and performing residual connection on the weighted initial cost volume and the layer-normalized initial cost volume.
[0014] Optionally, the first fusion module at least sequentially performs the following steps: adding the initial cost volume output by the normalization module and the initial cost volume output by the 3D attention module pixel by pixel; supplementing i columns of 0s on the right side of the initial cost volume after pixel-by-pixel addition to obtain a cost volume corresponding to a set of left and right feature maps; where the cost volumes output by the first fusion module are respectively of size of 、of size of and of size of ,where the is obtained by taking as input a set of feature maps output by the 5th convolutional layer in the feature extraction module, and the is obtained by taking as input a set of feature maps output by the 6th convolutional layer in the feature extraction module. It is obtained by taking a set of feature maps output by the 8th convolutional layer in the feature extraction module as input.
[0015] Optionally, the second fusion module at least sequentially performs the following steps: fusing the output of the first fusion module with and according to a preset fusion expression to obtain an initial fusion cost volume; fusing the initial fusion cost volume with according to the preset fusion expression to obtain a fusion cost volume; where the preset fusion expression is:
[0016]
[0017] where, represents 3D transposed convolution, Conv represents 3D convolution, x1 and x2 respectively represent the number of channels of the output channels, represents group normalization and activation function, represents a concatenation operation in the channel dimension.
[0018] Optionally, the aggregation module includes a 3D residual block and a cost aggregation module based on partial 3D convolution connected in series in sequence, where the cost aggregation module based on partial 3D convolution at least includes a first hourglass network based on partial 3D convolution, a second hourglass network based on partial 3D convolution, and a third hourglass network based on partial 3D convolution connected in series in sequence; where,
[0019] The 3D residual block at least sequentially performs the following steps: performing 3D convolution on the fusion cost volume output by the cost volume construction module; performing a residual connection on the fused cost volume after 3D convolution and the fusion cost volume output by the cost volume construction module;
[0020] Each of the first hourglass network based on partial 3D convolution, the second hourglass network based on partial 3D convolution, and the third hourglass network based on partial 3D convolution at least includes a first 3D convolution, a first partial 3D convolution, a second 3D convolution, a second partial 3D convolution, a third 3D convolution, and a fourth 3D convolution connected in series in sequence, where the third 3D convolution is residually connected to the first partial 3D convolution, and the fourth 3D convolution is residually connected to the first 3D convolution.
[0021] Optionally, the disparity regression module at least includes a first disparity regression module, a second disparity regression module, a third disparity regression module, and a fourth disparity regression module; where, each of the first disparity regression module, the second disparity regression module, the third disparity regression module, and the fourth disparity regression module includes an initial disparity generation module and a disparity refinement module connected in series in sequence;
[0022] The initial disparity generation module at least sequentially performs the following steps: performing at least 2 times of 3D convolution on the input; performing trilinear interpolation on the input after 3D convolution; performing probability weighted regression on the input after trilinear interpolation to obtain an initial predicted disparity map, where the probability of the initial predicted disparity map is calculated by the Sigmod function;
[0023] The disparity refinement module at least sequentially performs the following steps: interpolating the input left feature map and right feature map to the original view size; calculating the difference between the interpolated left feature map and the interpolated right feature map to obtain the left-right feature error; concatenating the left-right feature error and the input initial predicted disparity map in the channel dimension to obtain a first concatenated result; respectively performing convolution operations on the first concatenated result using 3 2D dilated convolutions, where the dilation rates of the 2D convolution are 1, 2, and 4 respectively;
[0024] Concatenating each output obtained after the convolution operation in the channel dimension to obtain a second concatenated result; applying a channel attention module to the second concatenated result, and performing a residual connection between the obtained result and the initial predicted disparity map to obtain the target predicted disparity map;
[0025] The output of the 3D residual block in the aggregation module is connected in series with the initial disparity generation module of the first disparity regression module, the output of the first partial 3D convolutional hourglass network in the aggregation module is connected in series with the initial disparity generation module of the second disparity regression module, the output of the second partial 3D convolutional hourglass network in the aggregation module is connected in series with the initial disparity generation module of the third disparity regression module, and the output of the third partial 3D convolutional hourglass network in the aggregation module is connected in series with the initial disparity generation module of the fourth disparity regression module.
[0026] Optionally, after the step of inputting the left view image and the right view image into the improved stereo matching network to obtain the target predicted disparity map, it further includes: calculating the loss between the target predicted disparity map and the ground truth disparity map; using the loss as a constraint during the training of the stereo matching network.
[0027] Through the stereo matching method based on the Mamba cost volume provided above, in the embodiments of the present application, a residual visual Mamba layer and a 3D attention module are introduced into the cost volume construction module, so as to fully explore the relationship between the deep and shallow features extracted during the feature extraction process through the residual visual Mamba layer, and combine the local enhancement method to provide richer context information for the cost volume, so that the stereo matching network can more easily handle the difficult matching problems in the pathological regions, thereby achieving the purpose of improving the matching accuracy and robustness of the stereo matching network. Brief Description of the Drawings
[0028] By referring to the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present application will become readily understandable. In the drawings, several embodiments of the present application are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0029] Figure 1 It is a flowchart of the steps of a stereo matching method based on Mamba cost volume for some embodiments of the present application;
[0030] Figure 2 It is a schematic diagram of the architecture of a stereo matching network for some embodiments of the present application;
[0031] Figure 3 It is a schematic diagram of the network architecture of the cost volume construction module in the stereo matching network for some embodiments of the present application;
[0032] Figure 4 It is a schematic diagram of the network architecture of the hourglass network based on partial 3D convolution and partial 3D convolution in the stereo matching network for some embodiments of the present application;
[0033] Figure 5 It is a schematic diagram of the network architecture of the disparity refinement module in the stereo matching network for some embodiments of the present application;
[0034] Figure 6 It is a qualitative comparison result graph of the present application on the SceneFlow dataset;
[0035] Figure 7 It is a comparison result graph of the predicted disparity of the present application on the KITTI2015 dataset;
[0036] Figure 8 It is a qualitative comparison result graph of the present application on the KITTI2012 test set. Detailed Embodiments
[0037] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present application.
[0038] It should be understood that the terms "including" and "comprising" used in the specification and claims of the present application indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0039] It should also be understood that the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification and claims of this application, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms. It should be further understood that the term "and / or" used in the specification and claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0040] As used in this specification and the claims, the term "if" can be interpreted as "when...", "once" or "in response to determining" or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]" according to the context.
[0041] The specific embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0042] Embodiment 1
[0043] Refer to Figure 1 , Figure 1 which is a flowchart of the steps of a stereo matching method based on a Mamba cost volume for some embodiments of this application. In some embodiments, a stereo matching method based on a Mamba cost volume provided by this application at least includes the following steps S10 and S20.
[0044] Step S10: Obtain a left-view image and a right-view image to be processed, where the left-view image and the right-view image are images of the same object at different perspectives;
[0045] It should be noted that this application does not limit the means of obtaining the left-view image and the right-view image, and those skilled in the art can select different means according to specific needs. For example, in some embodiments, the left-view image and the right-view image to be processed can be obtained by a stereo camera.
[0046] In some embodiments, the maximum disparity range of the left-view image and the right-view image can be set to 192.
[0047] Step S20: Input the left-view image and the right-view image into an improved stereo matching network to obtain a target predicted disparity map.
[0048] In some embodiments, refer to Figure 2 ,Figure 2 The figure is a schematic diagram of the architecture of the stereo matching network according to some embodiments of the present application. As Figure 2 shown, the improved stereo matching network at least includes a feature extraction module, a cost volume construction module, an aggregation module, and a disparity regression module connected in series in sequence, wherein the feature extraction module is also connected in series with the disparity regression module.
[0049] Further, referring to Figure 3 , Figure 3 The figure is a schematic diagram of the network architecture of the cost volume construction module in the stereo matching network according to some embodiments of the present application. As Figure 3 shown, the cost volume construction module includes an initial cost volume construction module, a residual visual mamba layer, a 3D attention module and a normalization module respectively connected in series with the output of the residual visual mamba layer, a first fusion module connected in series with the output of the 3D attention module and the output of the normalization module, and a second fusion module connected in series with the first fusion module. It should be noted that the second fusion module is not shown in Figure 3 .
[0050] In the technical solution provided in this embodiment, by introducing a residual visual mamba layer and a 3D attention module in the cost volume construction module, the relationship between the deep and shallow features extracted during the feature extraction process is fully explored through the residual visual mamba layer, and in combination with the local enhancement method, richer context information is provided for the cost volume, so that the stereo matching network can more easily handle the difficult matching problem in the pathological area, thereby achieving the purpose of improving the matching accuracy and robustness of the stereo matching network.
[0051] Next, a more detailed description of the specific implementation manner of the improved stereo matching network will be given, but these descriptions do not mean that only this implementation manner is feasible. In other words, these specific technical details or implementation methods are only used to help understand the content of the technical improvement, and will not limit the scope of the patent or impose constraints on the implementation manner of the patent.
[0052] Embodiment 2
[0053] In some embodiments, the left and right view images input share the weight parameters of the feature extraction module. Further, in some embodiments, the feature extraction module may be a pre-trained Efficient-B3 network.
[0054] As an example, the Efficient-B3 network can be pre-trained for image classification using the ImageNet dataset to obtain the pre-trained Efficient-B3 network, and then used as the feature extraction module of the stereo matching network.
[0055] In some embodiments, the left and right feature maps output by the 5th convolutional layer, the 6th convolutional layer, and the 8th convolutional layer in the feature extraction module serve as the inputs to the cost volume construction module connected in series with the feature extraction module, while the left and right feature maps output by the 4th convolutional layer in the feature extraction module serve as the inputs to the disparity regression module connected in series with the feature extraction module.
[0056] In some specific embodiments, the sizes of the left and right feature maps output by the 4th convolutional layer, the 5th convolutional layer, the 6th convolutional layer, and the 8th convolutional layer in the feature extraction module are 1 / 2, 1 / 4, 1 / 8, and 1 / 16 of the left and right view images respectively, and the numbers of channels are 24, 32, 48, and 132 respectively.
[0057] In the technical solution provided in this embodiment, by using the pre-trained Efficient-B3 network as the feature extraction module and setting the weight parameters of the feature extraction module to be shared by the input left and right views, the potential information in the left and right views can be effectively extracted.
[0058] Embodiment III
[0059] In some embodiments, as Figure 3 shown, the cost volume construction module includes an initial cost volume construction module connected in series in sequence, a residual vision mamba layer, a 3D attention module and a normalization module respectively connected in series with the output of the residual vision mamba layer, a first fusion module connected in series with the output of the 3D attention module and the output of the normalization module, and a second fusion module connected in series with the first fusion module.
[0060] According to the foregoing description, the inputs to the cost volume construction module are the left and right feature maps output by the 5th convolutional layer, the 6th convolutional layer, and the 8th convolutional layer in the feature extraction module. It can be understood that the left and right feature maps output by each convolutional layer are a set of feature maps and are regarded as a set of inputs to the cost volume construction module. In the cost volume construction module, the left and right feature maps of each set of inputs need to go through the initial cost volume construction module, the residual vision mamba layer, the 3D attention module, the normalization module, and the first fusion module in the cost volume construction module to obtain the cost volume corresponding to each set of left and right feature maps. Finally, each set of cost volumes will be input into the second fusion module in the cost volume construction module for fusion to obtain the fused cost volume. As Figure 2 shown, the three cost volumes in the cost volume construction module are the cost volumes corresponding to each set of left and right feature maps output by the first fusion module, Figure 2 shown, the fused cost volume in the cost volume construction module is the fused cost volume obtained by fusing the three cost volumes through the second fusion module.
[0061] In some embodiments, three groups of parallel initial cost volume construction modules, residual visual mamba layers, 3D attention modules, normalization modules, and first fusion modules may be set in the cost volume construction module, and each group is used to process a group of left feature maps and right feature maps of the input to obtain the cost volume corresponding to the group of left feature maps and right feature maps. The second fusion module is connected in series with the first fusion modules in the three groups to fuse the cost volumes output by each group of the first fusion modules to generate a fused cost volume. Such a setting can process each group of inputs in parallel at the same time, thereby improving the matching rate.
[0062] In other embodiments, only one set of initial cost body construction modules, residual visual mamba layer, 3D attention module, normalization module and first fusion module may be set in the cost body construction module, and the second fusion module is connected in series with the first fusion module. Such a setting requires processing a set of left feature maps and right feature maps input in sequence, and after the processing of the previous set of left feature maps and right feature maps is completed, the next set of left feature maps and right feature maps are processed, and finally all cost bodies are fused in the second fusion module to obtain a fused cost body. It should be noted that when implementing the technical solution of this embodiment, those skilled in the art can flexibly select the corresponding solution according to the content and teaching disclosed in this embodiment.
[0063] Next, each module in the cost body construction module will be described in detail, but this does not limit the present application.
[0064] Specifically, in some embodiments, the initial cost volume construction module at least performs the following steps S21 to S23 in sequence:
[0065] Step S21: remove pixel values in the left i columns and the right i columns of the input left feature map and right feature map, where i is the disparity level of the left feature map and the right feature map;
[0066] Step S22: splicing the left feature map and the right feature map after the pixel values are removed in a staggered superposition manner in the channel latitude to obtain an initial cost volume;
[0067] Step S23: The initial cost volume is upscaled so that the latitude of the upscaled initial cost volume meets the input latitude requirement of the residual visual mamba layer.
[0068] In some embodiments, the input latitude requirement of the residual visual mamba layer can be set to four dimensions, and the number of output channels of the residual visual mamba layer can be set to C / 4. Therefore, in the above step S23, the initial cost volume is upgraded to four dimensions.
[0069] As an example, assume that the size of the input left feature map and right feature map is HxWxC, and the disparity level is i. Remove the pixel values of the left i columns and the right i columns in the left and right feature maps. Then, splice the left and right feature maps in the channel latitude in an interlaced superposition manner to obtain the size The initial cost Finally, the initial cost body Upgrade it to four dimensions and output it to the residual visual Mamba layer in series with it.
[0070] After the initial cost volume is obtained by the initial cost volume construction module, the initial cost volume is input into the residual visual mamba module for further processing. Figure 3 As shown, the residual visual mamba module at least performs the following steps S24 to S28:
[0071] Step S24: performing layer normalization on the initial cost volume output by the initial cost volume construction module;
[0072] Step S25: inputting the layer-normalized initial cost volume into the Mamba model;
[0073] Step S26: adding the learnable parameters to the initial cost volume output by the Mamba model pixel by pixel;
[0074] Step S27: performing layer normalization on the initial cost volume after pixel-by-pixel addition;
[0075] Step S28: Input the layer-normalized initial cost volume into the linear layer.
[0076] After being processed by the above-mentioned residual visual Mamba module, the initial cost volume can be enhanced to fully explore the relationship between the deep and shallow features extracted during the feature extraction process.
[0077] The initial cost volume processed by the above-mentioned residual visual Mamba module is input into the 3D attention module and the normalization module in the cost volume construction module for further processing. Finally, the initial cost volumes processed by these two branches are fused by the first fusion module in the cost volume construction module to obtain a set of cost volumes corresponding to the left feature map and the right feature map.
[0078] In some embodiments, Figure 3 As shown, the first branch 3D attention module at least performs the following steps S29 to S35 in sequence:
[0079] Step S29: performing maximum pooling and average pooling on the initial cost volume output by the residual visual mamba layer respectively;
[0080] Step S30: Add the initially obtained cost volume after max pooling and the initially obtained cost volume after average pooling pixel by pixel;
[0081] Step S31: Perform the first Sigmoid normalization on the initially obtained cost volume after pixel-by-pixel addition;
[0082] Step S32: Perform the first pixel-by-pixel multiplication on the initially obtained cost volume output by the residual vision Mamba layer and the initially obtained cost volume after the first Sigmoid normalization;
[0083] Step S33: Perform at least 3 3D convolutions on the initially obtained cost volume after the first pixel-by-pixel multiplication, where the kernel sizes of the 3 3D convolutions are 1x7x7, 7x1x1, and 7x7x7 in sequence, acting on the spatial, depth, and overall dimensions respectively;
[0084] Step S34: Perform the second Sigmoid normalization on the initially obtained cost volume after at least 3 3D convolutions;
[0085] Step S35: Perform the second pixel-by-pixel multiplication on the initially obtained cost volume after the second Sigmoid normalization and the initially obtained cost volume after the first pixel-by-pixel addition.
[0086] In some embodiments, as Figure 3 shown, the second branch normalization module at least sequentially performs the following steps S36 to S39:
[0087] Step S37: Perform layer normalization on the initially obtained cost volume output by the residual vision Mamba layer;
[0088] Step S38: Weight the initially obtained cost volume output by the residual vision Mamba layer based on learnable parameters;
[0089] Step S39: Perform a residual connection on the weighted initially obtained cost volume and the layer-normalized initially obtained cost volume.
[0090] After the initially obtained cost volume output by the residual vision Mamba layer is processed by the 3D attention module and the normalization module respectively, the initially obtained cost volume output by the 3D attention module and the initially obtained cost volume output by the normalization module are input into the first fusion module for fusion to obtain a cost volume corresponding to a set of left feature maps and right feature maps. In some embodiments, as Figure 3 shown, the first fusion module at least sequentially performs the following steps S40 and S41:
[0091] Step S40: Add the initially obtained cost volume output by the normalization module and the initially obtained cost volume output by the 3D attention module pixel by pixel;
[0092] Step S41: Supplement i columns of 0 on the right side of the initial cost volume after pixel-by-pixel addition to obtain a cost volume corresponding to a set of left feature maps and right feature maps.
[0093] As an example, based on the foregoing example, the initial cost volume output by the initial cost volume construction module is . After being processed by the above Step S40 of the first fusion module, the initial cost volume is obtained. Then, at the above Step S41, supplement columns of 0 on the right side of the initial cost volume to finally form a cost volume with a shape of at the current scale. Where is the number of disparity levels .
[0094] It can be understood that through the above steps of the first fusion module, a cost volume corresponding to a set of left feature maps and right feature maps can be obtained, that is, the cost volume in the cost volume construction module as shown in Figure 2 . It can be understood that since the inputs of the cost volume construction module are the left feature maps and right feature maps output by the 5th convolutional layer, 6th convolutional layer, and 8th convolutional layer in the feature extraction module respectively, and the left feature maps and right feature maps output by each convolutional layer are used as a set of inputs to the cost volume construction module, the cost volume construction module finally obtains 3 cost volumes.
[0095] Specifically, since the sizes of the left feature maps and right feature maps output by the 5th convolutional layer, 6th convolutional layer, and 8th convolutional layer, which are the inputs of the initial cost volume construction module, are 1 / 4, 1 / 8, and 1 / 16 of the left view image and right view image respectively, the cost volumes output by the first fusion module are respectively of , of , and of . Where is obtained by using a set of feature maps output by the 5th convolutional layer in the feature extraction module as the input, is obtained by using a set of feature maps output by the 6th convolutional layer in the feature extraction module as the input, is obtained by using a set of feature maps output by the 8th convolutional layer in the feature extraction module as the input.
[0096] All the cost volumes output by the first fusion module will be input into the second fusion module in the cost volume construction module for fusion to obtain a fused cost volume. In some embodiments, the second fusion module at least sequentially performs the following Step S42 and Step S43:
[0097] Step S42: According to a preset fusion expression, the Fuse with to obtain an initial fused cost volume;
[0098] Step S43: According to a preset fusion expression, fuse the initial fused cost volume with to obtain a fused cost volume;
[0099] where the preset fusion expression is:
[0100]
[0101] where, represents 3D deconvolution, Conv represents 3D convolution, x1 and x2 respectively represent the number of channels of the output channels, represents group normalization and activation function, represents a concatenation operation in the channel dimension.
[0102] As an example, the cost volumes , and after passing through the second fusion module, finally obtain a fused cost volume with a size of , where 16 is twice that of , and 48 is the disparity range.
[0103] In the technical solution provided in this embodiment, by introducing a residual visual mamba layer and a 3D attention module in the cost volume construction module, the relationship between the deep and shallow features extracted during the feature extraction process is fully explored through the residual visual mamba layer, and in combination with the local enhancement method, more abundant context information is provided for the cost volume, so that the stereo matching network can more easily handle the difficult matching problem in the pathological area, thereby achieving the purpose of improving the matching accuracy and robustness of the stereo matching network.
[0104] Embodiment 4
[0105] As Figure 2 shown, after the cost volume construction module outputs the fused cost volume, the fused cost volume is input into the aggregation module for cost volume aggregation. In some embodiments, as Figure 2 shown, the aggregation module includes a 3D residual block and a cost aggregation module based on partial 3D convolution connected in series in sequence. The cost aggregation module based on partial 3D convolution at least includes a first hourglass network based on partial 3D convolution, a second hourglass network based on partial 3D convolution, and a third hourglass network based on partial 3D convolution connected in series in sequence. In this embodiment, by replacing some ordinary 3D convolutions in the cost aggregation process with partial 3D convolutions in the aggregation module, the number of parameters and the computational amount of the stereo matching network are reduced, and the matching accuracy is improved.
[0106] Specifically, in some embodiments, the 3D residual block at least sequentially performs the following steps S44 and S45:
[0107] Step S44: Perform 3D convolution on the fused cost volume output by the cost volume construction module;
[0108] Step S45: Perform a residual connection on the fused cost volume after 3D convolution and the fused cost volume output by the cost volume construction module.
[0109] As Figure 2 shown, the output obtained after the fused cost volume passes through the 3D residual block is denoted as output 0.
[0110] Next, the output of the 3D residual block will be further input into the hourglass network based on partial 3D convolution connected in series therewith for processing. Specifically, in some embodiments, referring to Figure 4 , Figure 4 is a schematic diagram of the network architecture of the hourglass network based on partial 3D convolution and partial 3D convolution in the stereo matching network of some embodiments of the present application. As Figure 4 shown, the first hourglass network based on partial 3D convolution, the second hourglass network based on partial 3D convolution, and the third hourglass network based on partial 3D convolution each at least include a first 3D convolution, a first partial 3D convolution, a second 3D convolution, a second partial 3D convolution, a third 3D convolution, and a fourth 3D convolution connected in series in sequence, wherein the third 3D convolution is in residual connection with the first partial 3D convolution, and the fourth 3D convolution is in residual connection with the first 3D convolution. As Figure 2 shown, the outputs of the first hourglass network based on partial 3D convolution, the second hourglass network based on partial 3D convolution, and the third hourglass network based on partial 3D convolution are respectively denoted as output 1, output 2, and output 3.
[0111] In the technical solution provided in this embodiment, by using partial 3D convolution to replace some ordinary 3D convolutions in the cost aggregation process, while reducing the number of parameters and the amount of computation of the stereo matching network, it also has a relatively accurate matching effect.
[0112] Embodiment 5
[0113] In some embodiments, the disparity regression module at least includes a first disparity regression module, a second disparity regression module, a third disparity regression module, and a fourth disparity regression module. Among them, the first disparity regression module, the second disparity regression module, the third disparity regression module, and the fourth disparity regression module all include an initial disparity generation module and a disparity refinement module connected in series in sequence. Further, the output of the 3D residual block in the aggregation module is connected in series with the initial disparity generation module of the first disparity regression module, the output of the first partial 3D convolutional hourglass network is connected in series with the initial disparity generation module of the second disparity regression module, the output of the second partial 3D convolutional hourglass network is connected in series with the initial disparity generation module of the third disparity regression module, and the output of the third partial 3D convolutional hourglass network is connected in series with the initial disparity generation module of the fourth disparity regression module. According to such a setting, the output 0 of the 3D residual block in the aggregation module, the outputs 1, 2, and 3 of each partial 3D convolutional hourglass network can be correspondingly input into the initial disparity generation modules of the first disparity regression module, the second disparity regression module, the third disparity regression module, and the fourth disparity regression module to generate an initial predicted disparity map.
[0114] Specifically, in some embodiments, the initial disparity generation module at least sequentially performs the following steps S46 to S48:
[0115] Step S46: Perform at least 2 times of 3D convolution on the input;
[0116] Step S47: Perform trilinear interpolation on the input after 3D convolution;
[0117] Step S48: Perform probability weighted regression on the input after trilinear interpolation to obtain an initial predicted disparity map, where the probability of the initial predicted disparity map is calculated by the Sigmod function.
[0118] Further, after the initial disparity generation module outputs the initial predicted disparity map, the initial predicted disparity map is input into the disparity refinement module connected in series with it, and the left feature map and the right feature map output by the 4th convolutional layer in the feature extraction module are also input into the disparity refinement module to refine the initial predicted disparity map to obtain a target predicted disparity map.
[0119] In some embodiments, with reference to Figure 5 , Figure 5 is a schematic diagram of the network architecture of the disparity refinement module in the stereo matching network of some embodiments of the present application. As Figure 5 shown, the disparity refinement module at least sequentially performs the following steps S49 to S54:
[0120] Step S49: Interpolate the input left feature map and right feature map to the original view size;
[0121] Step S50: Calculate the difference between the interpolated left feature map and the interpolated right feature map to obtain the left - right feature error;
[0122] Step S51: Concatenate the left - right feature error and the input initial predicted disparity map in the channel dimension to obtain a first concatenation result;
[0123] Step S52: Perform convolution operations on the first concatenation result using 3 2D dilated convolutions respectively, where the dilation rates of the 2D dilated convolutions are 1, 2, and 4 respectively;
[0124] Step S53: Concatenate each output obtained after the convolution operation in the channel dimension to obtain a second concatenation result;
[0125] Step S54: Apply a channel attention module to the second concatenation result, and perform a residual connection between the obtained result and the initial predicted disparity map to obtain the target predicted disparity map.
[0126] Therefore, each disparity regression module finally obtains a target predicted disparity map. As Figure 2 shown, the target predicted disparity map output by the first disparity regression module is denoted as target predicted disparity map 0, the target predicted disparity map output by the second disparity regression module is denoted as target predicted disparity Figure 1 map, the target predicted disparity map output by the third disparity regression module is denoted as target predicted disparity Figure 2 map, and the target predicted disparity map output by the fourth disparity regression module is denoted as target predicted disparity Figure 3 .
[0127] In the technical solution provided in this embodiment, the target predicted disparity map generated by combining the feature error and the initial prediction error map can enhance the robustness of the stereo matching network in pathological regions such as specular reflection and occlusion.
[0128] Embodiment Six
[0129] In some embodiments, after the step of inputting the left view image and the right view image into the improved stereo matching network to obtain the target predicted disparity map, the following steps S55 and S56 are further included:
[0130] Step S55: Calculate the loss between the target predicted disparity map and the ground - truth disparity map;
[0131] Step S56: Use the loss as a constraint during the training of the stereo matching network.
[0132] In the technical solution provided in this embodiment, by calculating the Loss is used as a constraint during the training of the stereo matching network, so that the model parameters of the stereo matching network can be continuously updated during training, achieving a stereo matching network with high matching accuracy and robustness with only a small number of training batches.
[0133] According to the above embodiments, a stereo matching method based on the Mamba cost volume provided by the present application has at least the following beneficial effects:
[0134] (1) The present application introduces the Mamba model to fully explore the relationship between deep and shallow features extracted during the feature extraction process, and combines local enhancement methods to provide richer context information for the cost volume, enabling the model to more easily handle the difficult matching problems in pathological regions.
[0135] (2) The present application uses partial convolutions to replace some ordinary 3D convolutions in the cost aggregation process, achieving a more accurate effect while reducing the number of parameters and computational complexity.
[0136] (3) The present application designs a disparity refinement method that combines feature error and initial disparity, which can enhance the robustness of the model in pathological regions such as specular reflection and occlusion.
[0137] (4) A stereo matching method based on the Mamba cost volume proposed by the present application can achieve satisfactory prediction results with only a small number of training batches.
[0138] Embodiment Seven
[0139] This embodiment will verify the effect of the stereo matching network proposed in the above embodiments.
[0140] In this embodiment, the data in the SceneFlow and KITTI datasets are used to experiment with the stereo matching method based on the Mamba cost volume proposed by the present application. The SceneFlow dataset is a large-scale synthetic dataset for training and evaluating visual tasks such as stereo matching and depth estimation, which consists of computer-generated high-quality images. The KITTI dataset is collected from real road scenes in Karlsruhe, Germany and its surrounding areas. The data is recorded by multiple sensors installed on a car.
[0141] In this embodiment, the stereo matching method based on the Mamba cost volume proposed by the present application is compared with a series of advanced stereo matching methods in the prior art in terms of the EPE metric, and the corresponding computational complexity and number of parameters are also shown.
[0142] Table 1 shows the comparison results on the SceneFlow test set. As can be seen from Table 1, compared with PCWNet which also uses transformation operations in the disparity refinement stage, although the number of parameters of this application is 16.19% more than that of PCWNet, the computational cost is significantly reduced by 65.59%. In addition, the EPE index of this application is also 3.85% lower than that of PCWNet. Compared with the training strategy of ACVNet-fast which has three stages with 64 batches in each stage, we have carried out a total of 18 batches of training. Moreover, when the number of parameters and the computational cost are both lower than that of ACVNet-fast, the stereo matching method based on Mamba cost volume of this application has better accuracy.
[0143] Table 1. Comparison Results on the SceneFlow Test Set
[0144]
[0145] Furthermore, Figure 6 is the qualitative comparison result diagram of this application on the SceneFlow dataset. As Figure 6 shown, (a) represents the input image, and (b), (c), and (d) represent the disparity maps and error maps of GwcNet, ACVNet, and this application respectively. As Figure 6 can be seen, the method of this application has better prediction effect in pathological regions than GwcNet. Especially in regions without texture or with weak texture, it can generate smoother and more accurate results. Compared with ACVNet, the method of this application is more accurate in predicting object edges. At the same time, it shows that a cost volume containing rich information can use fewer parameters and have high accuracy in the aggregation process.
[0146] Table 2 shows the verification results on the KITTI2015 dataset. As can be seen from Table 2, when the number of parameters is only 0.52M more than that of AANet, the method of this application is superior to most networks in the prior art in terms of the D1all index of all pixels. For example, it is 6.12% higher than PGNet and 2.13% higher than CFNet. This result shows that the method of this application can achieve relatively accurate prediction effect in practical applications.
[0147] Table 2. Verification Results on the KITTI2015 Dataset
[0148]
[0149] Figure 7 is the comparison result diagram of the predicted disparity of this application on the KITTI2015 dataset. As Figure 7As shown, (a) represents the input image, and (b), (c), and (d) represent the disparity maps and error maps of GwcNet, ACVNet, and the present application respectively. From left to right are the visualization results of the details, holes, weak texture, and occlusion regions, and the corresponding regions are marked with dashed boxes. Benefiting from the powerful long-distance modeling ability of the Mamba cost volume and the targeted optimization of the error region by the disparity optimization module, compared with GwcNet, the method of the present application has significantly improved prediction effects in these pathological regions, and the prediction effects in the hole region and weak texture exceed ACVNet.
[0150] Table 3 shows the quantitative comparison results of the reflection region in the KITTI2012 dataset. It can be seen from Table 3 that the matching method based on the Mamba cost volume of the present application can well predict the disparity values of the reflection region, exceeding ACVNet in the px-3 metrics of all pixels and non-occluded pixels, and the number of parameters is 37.94% less than it. The computational cost of the method of the present application is 31.14% less than that of GwcNet, and the number of parameters decreases by 30.8% compared with GWcNet. Although the stacked hourglass network used in the cost aggregation module of the present application is the same as that of GwcNet, the addition of some 3D convolutions in the present application reduces the number of parameters and computational cost, and Mamba reduces the quadratic complexity to linear during modeling, reducing the computational overhead. In the disparity refinement module, the method of the present application uses 2D convolutions, which also avoids the increase in the number of parameters. In addition, the batch size set during the training of our stereo matching network is 2. Except for GANet and GCNet which are 1, the batch sizes of the remaining models in Table 3 during training are greater than 10. Although a smaller batch size is vulnerable to noise, the stereo matching network of the present application still achieves an effect comparable to that of ACVNet.
[0151] Table 3. Quantitative comparison results of the reflection region in the KITTI2012 dataset
[0152]
[0153] Figure 8 This is the qualitative comparison result graph of the present application on the KITTI2012 test set. As Figure 8 shown, (a) represents the input image, and (b), (c), and (d) represent the disparity maps and error maps of GwcNet, ACVNet, and the present application respectively. From left to right are the occlusion, reflection, edge, and hole regions, and the corresponding regions are marked with dashed boxes. It can be Figure 8 seen that the present application shows better performance in the pathological regions. For example, in the second figure, the effect of the reflection region on the roof is significantly better than that of GwcNet and ACVNet, and the error in the right hole region in the fourth figure is also significantly improved.
[0154] Although several embodiments of the present application have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many changes, variations, and alternative methods may occur to those skilled in the art without departing from the spirit and scope of the present application. It should be understood that various alternatives to the embodiments of the present application described herein may be employed in practicing the present application. The appended claims are intended to define the scope of the present application and thus cover equivalents or alternatives within the scope of these claims.
Claims
1. A stereo matching method based on Mamba cost volume, characterized in that: include: Acquire a left-view image and a right-view image to be processed, wherein the left-view image and the right-view image are images of the same object at different viewing angles; Inputting the left view image and the right view image into an improved stereo matching network to obtain a target predicted disparity map; The improved stereo matching network at least comprises a feature extraction module, a cost volume construction module, an aggregation module and a disparity regression module connected in series, and the feature extraction module is also connected in series with the disparity regression module; The cost volume construction module includes an initial cost volume construction module, a residual visual mamba layer, a 3D attention module and a normalization module connected in series with the output of the residual visual mamba layer, respectively, a first fusion module connected in series with the output of the 3D attention module and the output of the normalization module, and a second fusion module connected in series with the first fusion module; The initial cost volume construction module at least performs the following steps in sequence: Remove the pixel values in the left i columns and the right i columns of a set of input left feature maps and right feature maps, where i is the disparity level of the left feature map and the right feature map; The left feature map and the right feature map after the pixel values are removed are spliced in the channel latitude in an interlaced superposition manner to obtain an initial cost volume; The initial cost body is upscaled so that the latitude of the upscaled initial cost body meets the input latitude requirement of the residual visual mamba layer, wherein the input latitude requirement of the residual visual mamba layer is four-dimensional, and the number of output channels of the residual visual mamba layer is C / 4; The residual visual mamba layer at least performs the following steps in sequence: Performing layer normalization on the initial cost volume output by the initial cost volume construction module; Inputting the layer-normalized initial cost volume into the Mamba model; Adding the learnable parameters to the initial cost volume output by the Mamba model pixel by pixel; Performing layer normalization on the initial cost volume after pixel-by-pixel addition; Input the layer-normalized initial cost volume into a linear layer; The 3D attention module performs at least the following steps in sequence: Performing maximum pooling and average pooling on the initial cost volume output by the residual visual mamba layer respectively; Adding the initial cost volume after maximum pooling and the initial cost volume after average pooling pixel by pixel; Performing a first Sigmoid normalization on the initial cost volume after pixel-by-pixel addition; Performing a first pixel-by-pixel multiplication on the initial cost volume output by the residual visual mamba layer and the initial cost volume after the first sigmoid normalization; Perform at least three 3D convolutions on the initial cost volume after the first pixel-by-pixel multiplication, wherein the convolution kernel sizes of the three 3D convolutions are 1x7x7, 7x1x1, and 7x7x7, respectively, acting on the spatial, depth, and overall dimensions; Perform a second Sigmoid normalization on the initial cost volume after at least three 3D convolutions; Perform a second pixel-by-pixel multiplication on the initial cost volume after the second sigmoid normalization and the initial cost volume after the first pixel-by-pixel multiplication; The normalization module at least performs the following steps in sequence: Performing layer normalization on the initial cost volume output by the residual visual mamba layer; Weighting the initial cost volume output by the residual visual mamba layer based on the learnable parameters; The weighted initial cost volume and the layer-normalized initial cost volume are residually connected.
2. The method according to claim 1, characterized in that The left view image and the right view image share the weight parameters of the feature extraction module, and the feature extraction module is a pre-trained Efficient-B3 network; the left feature map and the right feature map output by the 4th convolutional layer, the 5th convolutional layer, the 6th convolutional layer and the 8th convolutional layer in the feature extraction module are used as the input of the cost volume construction module connected in series with the feature extraction module, and the left feature map and the right feature map output by the 4th convolutional layer in the feature extraction module are used as the input of the disparity regression module connected in series with the feature extraction module; the sizes of the left feature map and the right feature map output by the 4th convolutional layer, the 5th convolutional layer, the 6th convolutional layer and the 8th convolutional layer are 1 / 2, 1 / 4, 1 / 8 and 1 / 16 of the left view image and the right view image, respectively, and the number of channels is 24, 32, 48 and 132, respectively.
3. The method according to claim 1, characterized in that The first fusion module at least performs the following steps in sequence: Adding the initial cost volume output by the normalization module to the initial cost volume output by the 3D attention module pixel by pixel; After pixel-by-pixel addition, i columns of 0 are added to the right of the initial cost volume to obtain a set of cost volumes corresponding to the left feature map and the right feature map; wherein the cost volumes output by the first fusion module are of size of , size is of and size is of , where is obtained by taking a set of feature maps output by the 5th convolutional layer in the feature extraction module as input, It is obtained by taking a set of feature maps output by the 6th convolutional layer in the feature extraction module as input. It is obtained by taking a set of feature maps output by the 8th convolutional layer in the feature extraction module as input.
4. The method according to claim 3, characterized in that The second fusion module at least performs the following steps in sequence: The output of the first fusion module is converted into and Perform fusion to obtain an initial fusion cost body; According to the preset fusion expression, the initial fusion cost volume and the The fusion is performed to obtain a fusion cost body; wherein the preset fusion expression is: ; in, represents 3D deconvolution, Conv represents 3D convolution, x1 and x2 represent the number of channels of the output channel, represents group regularization and Activation function, Represents a concatenation operation on the channel dimension.
5. The method according to claim 4, characterized in that The aggregation module includes a 3D residual block and a cost aggregation module based on partial 3D convolution connected in series, wherein the cost aggregation module based on partial 3D convolution includes at least a first hourglass network based on partial 3D convolution, a second hourglass network based on partial 3D convolution, and a third hourglass network based on partial 3D convolution connected in series; wherein, The 3D residual block at least performs the following steps in sequence: Performing 3D convolution on the fused cost volume output by the cost volume construction module; Performing residual connection on the fused cost volume after 3D convolution and the fused cost volume output by the cost volume construction module; The first hourglass network based on partial 3D convolution, the second hourglass network based on partial 3D convolution, and the third hourglass network based on partial 3D convolution all include at least a first 3D convolution, a first partial 3D convolution, a second 3D convolution, a second partial 3D convolution, a third 3D convolution, and a fourth 3D convolution connected in series, wherein the third 3D convolution is residually connected to the first partial 3D convolution, and the fourth 3D convolution is residually connected to the first 3D convolution.
6. The method according to claim 5, characterized in that The disparity regression module at least includes a first disparity regression module, a second disparity regression module, a third disparity regression module and a fourth disparity regression module; wherein the first disparity regression module, the second disparity regression module, the third disparity regression module and the fourth disparity regression module each include an initial disparity generation module and a disparity refinement module connected in series in sequence; The initial disparity generation module at least performs the following steps in sequence: Perform at least 2 3D convolutions on the input; Perform trilinear interpolation on the input after 3D convolution; Performing probability weighted regression on the input after trilinear interpolation to obtain an initial predicted disparity map, wherein the probability of the initial predicted disparity map is calculated by a Sigmod function; The parallax refinement module at least performs the following steps in sequence: Interpolate the input left and right feature maps to the original view size; Calculate the difference between the interpolated left feature map and the interpolated right feature map to obtain a left-right feature error; Splicing the left and right feature errors and the input initial predicted disparity map in a channel dimension to obtain a first splicing result; Using three 2D dilated convolutions to perform convolution operations on the first splicing results respectively, wherein the dilation rates of the 2D dilated convolutions are 1, 2, and 4 respectively; Each output obtained after the convolution operation is concatenated in the channel dimension to obtain a second concatenated result; Applying a channel attention module to the second splicing result, and performing a residual connection between the obtained result and the initial predicted disparity map to obtain the target predicted disparity map; The 3D residual block output in the aggregation module is connected in series with the initial disparity generation module of the first disparity regression module, the first partial 3D convolution-based hourglass network output in the aggregation module is connected in series with the initial disparity generation module of the second disparity regression module, the second partial 3D convolution-based hourglass network output in the aggregation module is connected in series with the initial disparity generation module of the third disparity regression module, and the third partial 3D convolution-based hourglass network output in the aggregation module is connected in series with the initial disparity generation module of the fourth disparity regression module.
7. The method according to claim 1, characterized in that After the step of inputting the left view image and the right view image into an improved stereo matching network to obtain a target predicted disparity map, the step further includes: Calculate the target predicted disparity map and the real disparity map loss; The The loss serves as a constraint during the training of the stereo matching network.
Citation Information
Patent Citations
Binocular vision stereo matching method based on dense multi-scale information fusion
CN115641285A
Lunar navigation image three-dimensional terrain reconstruction method based on deep learning
CN115984494A