Remote sensing image stereo matching method and system based on Mama model interpretation cost body

Through the remote sensing image stereo matching method based on the Mamba model, the selective state space model and multi-scale feature interaction mechanism are used to solve the problem of occlusion and large-span geographic modeling in remote sensing image stereo matching, and more efficient remote sensing image stereo matching is achieved.

CN120298723APending Publication Date: 2025-07-11BEIHANG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510463579.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing remote sensing image stereo matching technology faces multiple occlusions, low-texture area matching failures, local details and kilometer-level geographical structure spatial consistency in complex remote sensing scenarios, and the 3D convolution method is difficult to model large-span terrain and global continuity, resulting in an intensification of the matching error between occlusion boundaries and small terrain.

Method used

The method of interpreting the cost body based on the Mamba model is adopted, and the parallax probability distribution map is generated through multi-level and multi-scale semantic feature extraction, combined with the selective state space model and the multi-scale feature interaction mechanism, which overcomes the limitations of traditional methods and the local receptive field limitations of 3D convolution, and realizes adaptive modeling of global dependence and local details.

Benefits of technology

It improves the accuracy and efficiency of stereo matching of remote sensing images, can more accurately identify key information, improve the multi-grained feature characterization ability, and solves the problems of occlusion area robustness and large-span geographic modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298723A_ABST
    Figure CN120298723A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image stereo matching method and system based on a Mama model interpretation cost body, and the method comprises the steps: S1, extracting multi-level and multi-scale semantic features from input left and right remote sensing stereo image pairs, and outputting a stereo image pair feature map; s2, aligning the left and right stereo image pair feature maps on the parallax dimension to construct a three-dimensional cost body; s3, introducing a Mama framework based on a selective state space model, interpreting the cost body in combination with a multi-scale feature interaction mechanism, and generating a parallax probability distribution diagram; s4, mapping the probability distribution into a continuous disparity value, and outputting a disparity map; according to the method, adaptive feature screening and global modeling are realized through a dynamic weight selection mechanism of the state space model, the local receptive field limitation of 3D convolution is overcome, and the multi-granularity feature representation capability of the model is improved by promoting interaction and fusion among features of different scales by virtue of a cross-scale information interaction mechanism, so that the robustness of the model is improved. And the overall interpretation accuracy and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional reconstruction and machine learning, and particularly relates to remote sensing image analysis and binocular stereo matching. More specifically, it relates to a remote sensing image stereo matching method and system based on interpreting cost volume by the Mamba model. Background Art

[0002] Remote sensing image stereo matching, as a core link in three-dimensional reconstruction, aims to accurately match pixels in images from different perspectives to obtain disparity information. Currently, existing remote sensing image stereo matching technologies can be divided into two categories: traditional non-deep learning methods and end-to-end methods based on deep learning.

[0003] Traditional methods generally adopt a non-end-to-end architecture that first extracts features and then selects through a classifier. These methods construct matching costs based on manually designed features and optimize the disparity space consistency through dynamic programming or graph cut algorithms. Although traditional methods have good effects in some simple scenarios, due to their heavy dependence on prior assumptions and heuristic strategies, when facing problems such as multi-occlusion and matching failures in low-texture regions in complex remote sensing scenarios, it is difficult to balance the spatial consistency between local details and kilometer-level geographical structures, resulting in limitations in their generalization ability and accuracy.

[0004] Deep learning methods generally adopt an end-to-end multi-layer network architecture. By extracting features through a convolutional neural network, deep learning methods can construct a three-dimensional cost volume and use 3D convolutional layers to gradually aggregate context information to predict pixel-level matching probabilities. Typical frameworks such as PSMNet and GwcNet achieve cost volume regularization by stacking 3D convolutions. Although their local receptive fields can suppress noise interference, limited by the fixed-size convolutional kernels, it is difficult to model large-span features such as ridge lines and rivers, and they face rigid constraints on video memory and computing resources in gigapixel-level remote sensing data.

[0005] Generally speaking, traditional methods have significant limitations in terms of the generalization of remote sensing scenarios and the robustness of occluded areas due to their dependence on artificial features and heuristic optimization strategies. Although deep learning methods based on 3D convolutions partially overcome the above problems, they still face two inherent defects: one lies in the contradiction between local context and global dependence. The limited receptive field of the 3D convolutional kernel results in its ability to only model short-range disparity correlations and difficulty in capturing the global continuity of large-span geographical structures, leading to local jumps and breaks in the disparity map. The other lies in the lack of multi-scale information fusion channels. Existing methods compress the disparity dimension through cascaded downsampling operations, resulting in the loss of detailed information of shallow high-resolution features and exacerbating the matching errors at occlusion boundaries and small features.

[0006] Therefore, how to overcome the long-range dependence of the existing architecture on remote sensing images and improve the representation ability of multi-granularity features through a new global modeling paradigm and an efficient computing framework is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0007] In view of this, the present invention provides a remote sensing image stereo matching method and system based on interpreting cost volume by the Mamba model to solve some of the technical problems mentioned in the background art.

[0008] In order to achieve the above object, the present invention adopts the following technical solutions:

[0009] A remote sensing image stereo matching method based on interpreting cost volume by the Mamba model, comprising the following steps:

[0010] S1. Extract multi-level and multi-scale semantic features from the input left and right remote sensing stereo image pairs, and output stereo image pair feature maps;

[0011] S2. Align the left and right stereo image pair feature maps in the disparity dimension to construct a three-dimensional cost volume;

[0012] S3. Introduce the Mamba architecture based on the selective state space model, and combine the multi-scale feature interaction mechanism to interpret the cost volume to generate a disparity probability distribution map;

[0013] S4. Map the probability distribution to continuous disparity values and output a disparity map.

[0014] Preferably, the specific content of step S1 includes:

[0015] S11. Input the left and right two remote sensing stereo images, expand the input channels through the first layer of convolution, and activate through batch normalization and the ReLU function to reduce the spatial resolution and extract basic features;

[0016] S12. Through multiple stacked residual blocks, gradually increase the number of channels and adjust the resolution. Each residual block contains two 3×3 convolutions, and the shallow and deep features are fused through skip connections to alleviate gradient disappearance, obtaining different levels of feature maps;

[0017] S13. Concatenate the different levels of feature maps obtained during the multi-layer convolution process by channels to obtain deep semantic features with a high number of channels for constructing a grouped correlation cost volume; if multi-feature fusion is enabled, generate lightweight low-channel shallow detail features through additional convolution for constructing a concatenated cost volume.

[0018] Preferably, the specific content of step S2 includes:

[0019] S21. Slide and match the left image features and the right image features, that is, the reference features and the target features, in the horizontal direction;

[0020] S22. Divide the feature channels into groups and calculate the correlation independently for each group;

[0021] S23. For each disparity, perform a dot product of the left - image features and the right - image features at each disparity point, and calculate the mean for each group to obtain a grouped correlation cost volume;

[0022] S24. If multi - feature fusion is enabled, for each disparity, concatenate the left - image features and the features of the right - image after corresponding translation along the channels to form a concatenated cost volume, and concatenate the grouped correlation cost volume and the concatenated cost volume along the channel dimension.

[0023] Preferably, the specific content of step S3 includes:

[0024] S31. Coarse processing and multi - scale compression of the cost volume: Compress the number of channels of the cost volume through multiple hourglass - shaped 3D convolutions, initially extract low - dimensional global features, and then perform a series of hierarchical compressions of downsampling;

[0025] S32. Selective scanning and channel attention: Reshape the multi - scale features into a sequence, model the global dependencies of the cost volume through a selective state - space model, and introduce channel attention for processing;

[0026] S33. Perform cross - scale interaction, integrate global disparity information and local details through multi - scale feature fusion and upsampling, and generate a disparity probability distribution map.

[0027] Preferably, the specific content of step S32 includes:

[0028] S321. Feature serialization: Flatten the height dimension and width dimension in each scale cost volume into a one - dimensional sequence, and at the same time merge the channel dimension and the disparity dimension;

[0029] S322. Selective scanning: The selective state - space model MambaIR models the global dependencies of the cost volume along the spatial dimension, captures the laws of disparity continuity, and uses an input - dependent parameter generation method and a dynamically adjusted state transition matrix to adaptively adjust the memory retention strategy according to the input content;

[0030] S323. MambaIR introduces local enhancement and channel attention to improve the ordinary Mamba.

[0031] Preferably, in step S322, the calculation process of the selective state - space model is as follows:

[0032] Denote the input feature sequence as For each input vector at position k∈[1, L] in it Generate dynamic parameters through a linear layer:

[0033]

[0034] According to the generated Δ k Calculate the discretization matrix:

[0035]

[0036] where B is the batch size, L is the sequence length, F is the feature dimension, N is the hidden state dimension, A is a learnable diagonal matrix, is the projection matrix, W proj is the projection matrix, which is a weight matrix, is a fully connected layer, and Δ is the sampling interval;

[0037] Use the prefix product algorithm to calculate the hidden state h at all positions in parallel k and output:

[0038] y k = C k h k + Dx k

[0039] where y k is the output corresponding to a state space branch or scan direction, and D is a learnable skip connection matrix;

[0040] Finally, restore the output sequence to the original scale.

[0041] Preferably, the specific content of step S323 is: obtain channel statistics through global average pooling, generate channel weights through two layers of fully connected layers, perform feature weighting and output.

[0042] Preferably, the specific method of step S33 is:

[0043] Fuse the outputs of the three hierarchical Mamba branches:

[0044] X k-1 = X k-1 + Up(X k )

[0045] That is, add the low-resolution features after transposed convolutional upsampling to the high-level features;

[0046] Reduce the number of channels to 1 through 3D convolution to generate a disparity probability distribution map.

[0047] Preferably, the method for step S34 to output a disparity map is:

[0048] Apply Softmax to the output probability map along the disparity dimension to obtain the normalized probability;

[0049] Calculate the weighted expectation:

[0050]

[0051] Among them, d is a discrete disparity candidate value, and P(d) is the probability value of the d-th disparity;

[0052] Finally, an output disparity map is obtained.

[0053] A remote sensing image stereo matching system for interpreting cost volume based on the Mamba model, based on the remote sensing image stereo matching method for interpreting cost volume based on the Mamba model, includes: a feature extraction module, a cost volume construction module, a MambaClassif module, and a disparity map output module;

[0054] The feature extraction module is used to extract multi-level and multi-scale semantic features from the input left and right remote sensing stereo image pairs and output a stereo image pair feature map;

[0055] The cost volume construction module is used to align the left and right stereo image pair feature maps in the disparity dimension to construct a three-dimensional cost volume;

[0056] The MambaClassif module is used to introduce the Mamba architecture based on the selective state space model, combine the multi-scale feature interaction mechanism to interpret the cost volume, and generate a disparity probability distribution map;

[0057] The disparity map output module is used to map the probability distribution to continuous disparity values and output a disparity map.

[0058] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a remote sensing image stereo matching method and system for interpreting cost volume based on the Mamba model. In the feature extraction stage, through multi-level feature fusion, the model can not only capture the semantic information of a large-scale scene but also retain local details, providing complementary feature representations for the cost volume; in the cost volume construction stage, the framework reduces the computational complexity by using group correlation while retaining the matching discrimination, and also introduces patch-based features to supplement the original information; in the cost volume interpretation stage, through the dynamic weight selection mechanism of the state space model, adaptive feature screening and global modeling are realized, overcoming the local receptive field limitation of 3D convolution, and with the help of the cross-scale information interaction mechanism, by promoting the interaction and fusion between different-scale features, the key information in the cost volume can be more accurately identified, improving the multi-granularity feature representation ability of the model and enhancing the overall interpretation accuracy and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0060] Figure 1 Flowchart of a remote sensing image stereo matching method based on Mamba model to interpret the cost volume provided by the present invention;

[0061] Figure 2 Schematic diagram of the overall framework of the Mamba interpretation network algorithm provided by the present invention;

[0062] Figure 3 Schematic diagram of the calculation process of the selective state space model provided by the present invention. Detailed implementation manners

[0063] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0064] An embodiment of the present invention discloses a remote sensing image stereo matching method based on Mamba model to interpret the cost volume, as Figure 1 , including the following steps:

[0065] S1. Extract multi-level and multi-scale semantic features from the input left and right remote sensing stereo image pairs, and output the stereo image pair feature maps, providing a basis for constructing the cost volume subsequently. The features include both local details and global context information to cope with complex terrain and occlusion scenarios;

[0066] S2. Align the left and right stereo image pair feature maps in the disparity dimension to construct a three-dimensional cost volume for quantifying the matching degree under different disparity hypotheses;

[0067] S3. Introduce the Mamba architecture based on the selective state space model, and combine the multi-scale feature interaction mechanism to interpret the cost volume and generate a disparity probability distribution map;

[0068] S4. Map the probability distribution to continuous disparity values and output the disparity map.

[0069] To further implement the above technical solution, the specific content of step S1 includes:

[0070] S11. Input the left and right remote sensing stereo images, expand the input channels through the first layer of convolution, and activate them through batch normalization and the ReLU function to reduce the spatial resolution and extract the basic features;

[0071] S12. Through multiple stacked residual blocks, gradually increase the number of channels and adjust the resolution. Each residual block contains two 3×3 convolutions, and the shallow and deep features are fused through skip connections to alleviate the vanishing gradient, obtaining feature maps of different levels;

[0072] In this embodiment, the sizes of the input left and right remote sensing stereo images are [B, C, H, W], where B is the batch size, C is the number of channels, and H and W are the image height and width. The first layer of convolution is to expand the input channels from C to 32 using a 3×3 convolution with a stride of 2; the number of channels is gradually increased to 32→64→128 through multiple stacked residual blocks;

[0073] S13. Concatenate the feature maps of different levels obtained during the multi-layer convolution process by channels to obtain deep semantic features with a high number of channels for constructing the grouped correlation cost volume; if multi-feature fusion is enabled, generate lightweight shallow detail features with a low number of channels through additional convolution for constructing the concatenated cost volume;

[0074] The obtained deep semantic features with a high number of channels are gwc_feature with 320 channels, and the shallow detail features with a low number of channels are lightweight concat_feature with 12 channels;

[0075] In the feature extraction stage, through multi-level feature fusion, the model can capture both the semantic information of the large-scale scene and retain local details, providing complementary feature representations for the cost volume.

[0076] To further implement the above technical solution, the specific content of step S2 includes:

[0077] S21. Slide and match the left image features and the right image features, that is, the reference features and the target features, in the horizontal direction;

[0078] The disparity search range is [x_disp, d_disp], covering negative disparity to positive disparity;

[0079] S22. Divide the feature channels into num groups groups, and calculate the correlation independently for each group;

[0080] S23. For each disparity, perform a point-by-point dot product of the left image features and the right image features, and calculate the mean by group to obtain the grouped correlation cost volume;

[0081] Specifically, for each disparity d, the i-th column of the left image features is multiplied element-wise with the (i±d)-th column of the right image features, and the mean value is calculated for each group to obtain a cost volume of size [B, num groups , D, H, W], where D is the number of disparities;

[0082] S24. If multi-feature fusion is enabled, for each disparity, the left image features and the correspondingly translated right image features are concatenated along the channel dimension to form a concatenated cost volume. The grouped correlation cost volume and the concatenated cost volume are concatenated along the channel dimension to enhance feature diversity;

[0083] Specifically, for each disparity d, the left image features and the correspondingly translated right image features are directly concatenated along the channel dimension to form a cost volume of size [B, 2C, D, H, W].

[0084] During the construction of the cost volume, the framework uses grouped correlation to reduce the computational complexity while retaining the matching discrimination, and also introduces concatenated features to supplement the original information.

[0085] To further implement the above technical solution, as Figure 2 , the specific content of step S3 includes:

[0086] The input cost volume has a size of [B, C, D, H, W], where C = the number of groups + the concatenated channels, and D = the disparity range;

[0087] S31. Cost volume rough processing and multi-scale compression: The number of channels C of the cost volume is compressed to 32 through multiple hourglass-shaped 3D convolutions to initially extract low-dimensional global features. At this time, the size of the cost volume is [B, 32, D, H, W], and then a series of downsampling-based hierarchical compressions are performed;

[0088] Specifically, for the first scale (original resolution): directly retain [B, 32, D, H, W];

[0089] For the second scale Compress it to

[0090] For the third scale Compress it again to

[0091] Through hierarchical compression, the model can gradually focus on different granularity disparity patterns;

[0092] S32. Selective scanning and channel attention: Reshape the multi-scale features into a sequence, model the global dependencies of the cost volume through a selective state space model, and introduce channel attention for processing;

[0093] S33. Perform cross-scale interaction. Through multi-scale feature fusion and upsampling, integrate global disparity information and local details to generate a disparity probability distribution map.

[0094] To further implement the above technical solution, the specific content of step S32 includes:

[0095] S321. Feature serialization: Flatten the height dimension and width dimension in each scale cost volume into a one-dimensional sequence, and at the same time merge the channel dimension and disparity dimension;

[0096] Taking the process of processing the cost volume at the first scale, i.e., the original resolution, as an example:

[0097] Flatten the height dimension and width dimension in [B, 32, D, H, W] into a one-dimensional sequence, and at the same time merge the channel dimension and disparity dimension to obtain [B, L, F]. Among them, L = H×W is the sequence length, and F = C×D is the feature dimension. The specific operation is as follows:

[0098]

[0099] That is, regard all channel and disparity information of each pixel point (h, w) as a feature vector of one time step to form a sequence with a length of L = H×W;

[0100] S322. Selective scanning: The selective state space model MambaIR models the global dependencies of the cost volume along the spatial dimension (H×W), captures the law of disparity continuity, and adaptively adjusts the memory retention strategy according to the input content using an input-dependent parameter generation method and a dynamically adjusted state transition matrix;

[0101] S323. MambaIR introduces local enhancement and channel attention to improve ordinary Mamba.

[0102] To further implement the above technical solution, as Figure 3 , in step S322, the calculation process of the selective state space model is as follows:

[0103] Denote the input feature sequence as For the input vector at each position k∈[1, L] among them Generate dynamic parameters through a linear layer:

[0104]

[0105] That is, Δ k , B k , C k are all uniformly generated by the input x k through a single fully connected layer;

[0106] According to the generated Δk Calculate the discretized matrix:

[0107]

[0108] Among them, B is the batch size, L is the sequence length, F is the feature dimension, N is the hidden state dimension, A is a learnable diagonal matrix, is the projection matrix, W proj is the projection matrix, which is a weight matrix, is a fully connected layer, and Δ is the sampling interval;

[0109] Δ is a scalar with a length of 1, B and C are one-dimensional vectors with a length of N, so W proj has a row dimension of (1 + 2N); it should be noted that although formally B and C are one-dimensional vectors (with a length of N), rather than two-dimensional matrices in traditional control theory, they are extended to the corresponding spatial dimensions through broadcasting or operations with other tensors during the calculation process; therefore, there is no need to exist in the two-dimensional form of a traditional matrix, but to perform state update and output mapping functions more lightly and efficiently;

[0110] Use the prefix product algorithm to calculate the hidden state h at all positions in parallel k , and output:

[0111] y k = C k h k + Dx k

[0112] Among them, y k is the output corresponding to a state space branch or scan direction (combining various factors of the internal state (attenuated by A), input-state mapping B, state-output mapping C, and direct feedback D), and D is a learnable skip connection matrix;

[0113] Finally, restore the output sequence [B, L, F] to the original scale [B, 32, D, H, W].

[0114] The selective state space model uses an input-dependent parameter generation method and a dynamically adjusted state transition matrix, enabling it to adaptively adjust the memory retention strategy according to the input content, effectively taking into account both global context association and local detail extraction.

[0115] To further implement the above technical solution, the specific content of step S323 is: obtain channel statistics through global average pooling, generate channel weights through two fully connected layers, perform feature weighting and output;

[0116] In this embodiment, for the [B, C, D, H, W] output by the previous part, first obtain channel statistics through global average pooling

[0117]

[0118] Then, the channel weights are generated through two fully connected layers.

[0119] α = σ(W2·ReLU(W1·s))

[0120] where σ is the Sigmoid function, r is the compression ratio. Finally, feature weighting is performed to output X out = α·X, suppressing low-response channels.

[0121] To further implement the above technical solution, the specific method of step S33 is as follows:

[0122] Fuse the outputs of the three hierarchical Mamba branches:

[0123] X k-1 = X k-1 + Up(X k )

[0124] That is, upsample the low-resolution features through transposed convolution and add them to the high-level features. For example:

[0125]

[0126] Reduce the number of channels to 1 through 3D convolution to generate a disparity probability distribution map [B, 1, D, H, W].

[0127] In this embodiment, in step S3, the Mamba architecture based on the Selective State Space Model (Selective SSM) is introduced, combined with the multi-scale feature interaction mechanism, to innovatively replace the traditional 3D convolution operation. Compared with the traditional cost volume processing method, the multi-scale Mamba algorithm framework proposed by the present invention effectively models the long-range spatio-temporal dependencies in the cost volume by introducing a sequence modeling method with global perception ability without increasing the computational complexity; it breakthroughly solves two major technical bottlenecks existing in the cost volume parsing of traditional convolutional neural networks: one is to achieve adaptive feature screening and global modeling through the dynamic weight selection mechanism of the state space model, overcoming the local receptive field limitation of 3D convolution; the other is to improve the multi-granularity feature representation ability of the model by means of the cross-scale information interaction mechanism.

[0128] To further implement the above technical solution, the method for outputting the disparity map in step S34 is as follows:

[0129] Apply Softmax along the disparity dimension to the output probability map to obtain the normalized probability;

[0130] Calculate the weighted expectation:

[0131]

[0132] where d is a discrete disparity candidate value, and P(d) is the probability value of the d-th disparity;

[0133] Here, the weighted summation actually performs a "soft" argmax operation. By taking the weighted average of all possible disparity values, the discrete candidate values are converted into a continuous value. The obtained disparity value is the expected value of each candidate item in the mathematical sense, which can better represent the continuous disparity, thereby realizing sub-pixel level disparity regression;

[0134] Finally, the output disparity map is [B, 1, H, W].

[0135] In another embodiment, two common remote sensing image stereo matching datasets, WHU-Stereo and US3D datasets, are used for comparative experimental verification, where:

[0136] WHU-Stereo: The image category is panchromatic map. The number of images in this dataset is 1757, and the image size is 1024×1024. Among them, the dataset is divided into 1220 training images, 122 validation images, and 415 test images;

[0137] US3D: The image category is RGB, the number of images is 4292, and the image size is 1024×1024. Among them, the dataset is divided into 3838 training images and 454 test images.

[0138] The experimental results are shown in the following table. Table 1 shows the stereo matching performance of this algorithm combined with the latest model on the WHU-Stereo dataset; Table 2 shows the stereo matching performance of this algorithm combined with the latest model on the US3D dataset; It can be seen that the combination of the latest model and this algorithm has improved the performance on these two remote sensing stereo matching datasets.

[0139] Table 1 Performance of the present invention on the WHU-Stereo dataset:

[0140]

[0141]

[0142] Table 2 Performance of the present invention on the US3D dataset:

[0143]

[0144] A remote sensing image stereo matching system for interpreting cost volume based on the Mamba model, based on a remote sensing image stereo matching method for interpreting cost volume based on the Mamba model, includes: a feature extraction module, a cost volume construction module, a MambaClassif module, and a disparity map output module;

[0145] The feature extraction module is used to extract multi-level and multi-scale semantic features from the input left and right remote sensing stereo image pairs and output the stereo image pair feature maps;

[0146] The cost volume construction module is used to align the left and right stereo image pair feature maps in the disparity dimension and construct a three-dimensional cost volume;

[0147] The MambaClassif module is used to introduce the Mamba architecture based on the selective state space model, combine the multi-scale feature interaction mechanism to interpret the cost volume, and generate a disparity probability distribution map;

[0148] The disparity map output module is used to map the probability distribution to continuous disparity values and output a disparity map.

[0149] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements a remote sensing image stereo matching method for interpreting cost volume based on the Mamba model.

[0150] A processing terminal includes a memory and a processor. A computer program that can run on the processor is stored in the memory. When the processor executes the computer program, it implements a remote sensing image stereo matching method for interpreting cost volume based on the Mamba model.

[0151] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, please refer to the description in the method part.

[0152] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A remote sensing image stereo matching method for interpreting cost volume based on the Mamba model, characterized in that, It includes the following steps: S1. Extract multi-level and multi-scale semantic features from the input left and right remote sensing stereo image pairs, and output the stereo image pair feature maps; S2. Align the left and right stereo image pair feature maps in the disparity dimension to construct a three-dimensional cost volume; S3. Introduce the Mamba architecture based on the selective state space model, and combine the multi-scale feature interaction mechanism to interpret the cost volume and generate a disparity probability distribution map; S4. Map the probability distribution to continuous disparity values and output the disparity map.

2. A remote sensing image stereo matching method based on interpreting a cost volume by using the Mamba model according to claim 1, wherein The specific content of step S1 includes: S11. Input the left and right two remote sensing stereo images, expand the input channels through the first layer of convolution, and activate through batch normalization and the ReLU function to reduce the spatial resolution and extract the basic features; S12. Through multiple stacked residual blocks, gradually increase the number of channels and adjust the resolution. Each residual block contains two 3×3 convolutions, and the shallow and deep features are fused through skip connections to alleviate the vanishing gradient, obtaining feature maps of different levels; S13. Concatenate the feature maps of different levels obtained in the multi-layer convolution process along the channels to obtain deep semantic features with a high number of channels for constructing a grouped correlation cost volume; if multi-feature fusion is enabled, generate lightweight low-channel shallow detail features through additional convolution for constructing a concatenated cost volume.

3. A remote sensing image stereo matching method for interpreting cost volume based on the Mamba model according to claim 1, characterized in that, The specific content of step S2 includes: S21. Slide and match the left image features and the right image features, that is, the reference features and the target features, in the horizontal direction; S22. Divide the feature channels into groups and calculate the correlation independently for each group; S23. For each disparity, perform a dot product of the left image features and the right image features at each disparity point and calculate the mean by group to obtain a grouped correlation cost volume; S24. If multi-feature fusion is enabled, for each disparity, concatenate the left image features and the features of the right image after corresponding translation along the channels to form a concatenated cost volume, and concatenate the grouped correlation cost volume and the concatenated cost volume along the channel dimension.

4. A remote sensing image stereo matching method based on Mamba model to interpret cost volume according to claim 1, characterized in that The specific content of step S3 includes: S31. Coarse processing and multi-scale compression of the cost volume: Compress the number of channels of the cost volume through multiple hourglass-shaped 3D convolutions, initially extract low-dimensional global features, and then perform a series of downsampling hierarchical compressions; S32. Selective scanning and channel attention: Reshape the multi-scale features into a sequence, model the global dependencies of the cost volume through the selective state space model, and introduce channel attention for processing; S33. Perform cross-scale interaction, integrate the global disparity information and local details through multi-scale feature fusion and upsampling, and generate a disparity probability distribution map.

5. A method for stereo matching of remote sensing images based on interpreting cost volume using the Mamba model according to claim 4, characterized in that The specific content of step S32 includes: S321. Feature serialization: Flatten the height dimension and width dimension in each scale cost volume into a one-dimensional sequence, and at the same time merge the channel dimension and the disparity dimension; S322. Selective scanning: The selective state space model MambaIR models the global dependencies of the cost volume along the spatial dimension, captures the rules of disparity continuity, and uses an input-dependent parameter generation method and a dynamically adjusted state transition matrix to adaptively adjust the memory retention strategy according to the input content; S323. MambaIR introduces local enhancement and channel attention to improve the ordinary Mamba.

6. A method for stereo matching of remote sensing images based on interpreting cost volume by Mamba model according to claim 5, characterized in that, In step S322, the calculation process of the selective state space model is as follows: Denote the input feature sequence as For each input vector at position k ∈ [1, L] among them Generate dynamic parameters through a linear layer: According to the generated Δ k Calculate the discretization matrix: Among them, B is the batch size, L is the sequence length, F is the feature dimension, N is the hidden state dimension, A is a learnable diagonal matrix, is the projection matrix, W proj is the projection matrix, which is a weight matrix, is a fully connected layer, and Δ is the sampling interval; Use the prefix product algorithm to calculate the hidden state h at all positions in parallel k , and output: y k = C k h k + Dx k where y k is the output corresponding to a state space branch or scan direction, and D is a learnable skip connection matrix; Finally, restore the output sequence to the original scale.

7. A method for stereo matching of remote sensing images based on interpreting cost volume with the Mamba model according to claim 5, characterized in that, The specific content of step S323 is: obtain channel statistics through global average pooling, generate channel weights through two fully connected layers, perform feature weighting and output.

8. A remote sensing image stereo matching method for interpreting a cost volume based on the Mamba model according to claim 4, characterized in that, The specific method of step S33 is: Fuse the outputs of the three hierarchical Mamba branches: X k-1 = X k-1 + Up(X k ) That is, upsample the low-resolution features by transposed convolution and add them to the high-level features; Reduce the number of channels to 1 through 3D convolution to generate a disparity probability distribution map.

9. A remote sensing image stereo matching method for interpreting a cost volume based on the Mamba model according to claim 4, characterized in that The method for outputting a disparity map in step S34 is: Apply Softmax to the output probability map along the disparity dimension to obtain a normalized probability; Calculate the weighted expectation: where d is a discrete disparity candidate value, and P(d) is the probability value of the d-th disparity; Finally, output the disparity map.

10. A remote sensing image stereo matching system for interpreting cost volume based on the Mamba model, characterized in that, A remote sensing image stereo matching method based on Mamba model to interpret cost volume according to any one of claims 1-9, comprising: a feature extraction module, a cost volume construction module, a MambaClassif module, and a disparity map output module; The feature extraction module is used to extract multi-level and multi-scale semantic features from the input left and right remote sensing stereo image pairs and output stereo image pair feature maps; The cost volume construction module is used to align the left and right stereo image pair feature maps in the disparity dimension to construct a three-dimensional cost volume; The MambaClassif module is used to introduce a Mamba architecture based on a selective state space model, combine a multi-scale feature interaction mechanism to interpret the cost volume, and generate a disparity probability distribution map; The disparity map output module is used to map the probability distribution to continuous disparity values and output a disparity map.

Citation Information

Cited By

  • Stereo matching method of high-resolution stereo satellite panchromatic image pair

    CN120495285A

  • A stereo matching method for high-resolution stereo satellite panchromatic image pairs

    CN120495285B