Satellite image cascade matching method and system based on double-branch context awareness

The multi-scale main features and context features of satellite images are extracted through a dual-branch network, and the context-aware cost body is constructed and refined. This solves the problem of insufficient parallax estimation accuracy of satellite remote sensing images in complex scenarios, achieving efficient three-dimensional reconstruction effect.

CN120298730APending Publication Date: 2025-07-11EAST CHINA NORMAL UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510433552.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The parallax estimation accuracy of existing satellite remote sensing images in low-texture areas, repeated texture areas and severely changing parallax areas, affecting the accuracy of three-dimensional information extraction.

Method used

The multi-scale main features and context features of satellite images are extracted separately by a dual-branch network, a context-aware cost body is constructed, the initial disparity map is optimized through a multi-scale aggregation strategy, and the gradient information is combined for fine processing, and the feature extraction capability is enhanced by using lightweight convolutional neural networks and residual modules.

Benefits of technology

It significantly improves the parallax estimation accuracy of satellite remote sensing images, adapts to the matching needs of large parallax ranges and low-texture areas, reduces computing resource consumption, enhances the matching and discrimination capabilities of key areas, alleviates the matching ambiguity between occlusion boundaries and discontinuous areas of parallax, and achieves efficient three-dimensional reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298730A_ABST
    Figure CN120298730A_ABST
Patent Text Reader

Abstract

The invention relates to the field of remote sensing image processing and computer vision, and discloses a satellite image cascade matching method and system based on double-branch context awareness, and the method comprises the following steps: constructing a heterogeneous double-branch feature extraction network, extracting multi-scale detail features through a main feature network, and capturing global semantic information through a context coding network; fusing multi-scale main features and context features to construct a group-related cost body, and adaptively enhancing key region matching response in combination with a channel incentive mechanism; optimizing the cost body by adopting a cross-scale information transfer and multi-dimensional attention fusion strategy, and generating a multi-scale initial disparity map; image edge geometric features and semantic contexts are fused, and high-resolution parallax details are recovered through a multi-stage residual decoder. According to the method, the problem of matching fuzziness of a traditional method in weak texture, repeated structure and parallax abrupt change areas of a satellite image is solved, and the three-dimensional reconstruction precision and the edge detail integrity in a complex scene are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing and computer vision, and specifically to a cascaded matching method and system for satellite images based on dual-branch context awareness. Background Art

[0002] With the development of high-resolution satellite remote sensing technology, the extraction of three-dimensional information based on satellite stereo images has been widely applied in fields such as topographic mapping, urban planning, and disaster assessment.

[0003] However, due to the fact that satellite remote sensing images are easily affected by factors such as illumination, perspective differences, and sensor errors during the shooting process, imperfect correction between images is caused, further resulting in insufficient parallax estimation accuracy of existing methods in low-texture areas, repetitive texture areas, and areas with drastic parallax changes.

[0004] Therefore, there is an urgent need for a stereo matching method that can effectively fuse context information and improve the accuracy and robustness of parallax estimation. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention provides a cascaded matching method and system for satellite images based on dual-branch context awareness. By introducing multi-scale context feature extraction, cost volume construction and aggregation, and context-guided disparity refinement modules, the disparity estimation performance of high-resolution satellite remote sensing images is significantly improved.

[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A cascaded matching method for satellite images based on dual-branch context awareness, including the following steps:

[0007] Extract the multi-scale main features and context features of the satellite image through a dual-branch network respectively;

[0008] Fuse the main features and context features to construct a context-aware cost volume;

[0009] Optimize the cost volume using a multi-scale aggregation strategy and calculate the initial disparity map;

[0010] Refine the disparity map based on gradient information and context features.

[0011] Preferably, in the step of extracting the multi-scale main features and context features of the satellite image through a dual-branch network respectively:

[0012] The main features are extracted through a lightweight convolutional neural network, and the network structure includes inverted residual modules;

[0013] The context features are extracted through an encoder structure, including residual blocks and downsampling layers.

[0014] Preferably, the lightweight convolutional neural network is MobileNetV2 and pre-trained on the ImageNet dataset.

[0015] Preferably, the steps of fusing the main feature and the context feature to construct the context-aware cost volume include:

[0016] Concatenate the main feature and the context feature in the channel dimension, and calculate the population correlation cost volume based on the concatenated feature;

[0017] Generate channel excitation weights using the left image feature, and perform spatial attention enhancement on the population correlation cost volume through the weights;

[0018] Perform regularization processing on the enhanced cost volume to obtain the optimized context-aware cost volume.

[0019] Preferably, the population correlation cost volume is calculated by the following formula:

[0020]

[0021] where C gwc (g, d, h, w) represents the correlation value at the g-th group, candidate disparity d, and spatial position (h, w) in the population correlation cost volume; N c represents the total number of channels of the concatenated feature map ; N g represents the number of groups for population correlation calculation; represents the concatenated feature value of the left image at the position (h, w); represents the concatenated feature value of the right image at the position (h, w - d).

[0022] Preferably, the steps of optimizing the cost volume and calculating the initial disparity map using the multi-scale aggregation strategy include:

[0023] Gradually transfer the context-aware cost volumes of different scales from low resolution to high resolution, and superimpose the semantic information of the low-resolution cost volume on the high-resolution cost volume through upsampling operation;

[0024] Introduce an attention mechanism in the disparity dimension, calculate the global attention vector and the local attention vector respectively, and dynamically adjust the matching weights of different positions and disparity ranges through weighted fusion.

[0025] Preferably, the steps of refining the disparity map based on gradient information and context feature include:

[0026] Calculate the gradient map of the left image through the gradient operator, and concatenate the initial disparity map and the gradient map in the channel dimension to obtain the edge feature;

[0027] Construct a dual-branch coding structure, where the first branch processes the edge features and the second branch multiplexes the extracted context features;

[0028] Fuse the dual-branch features through a multi-level residual decoder, and upsample step by step to generate a refined disparity map.

[0029] Preferably, the gradient operator is a Sobel operator, which is used to extract the horizontal and vertical gradients of the image. After the gradient map is spliced with the initial disparity map, it is input into the first branch.

[0030] Preferably, the method further includes:

[0031] Supervise and train the multi-scale initial disparity map and the refined disparity map using a weighted smooth L1 loss function, and the loss function satisfies:

[0032]

[0033] where n is the total number of output levels, λ i is the weight coefficient of the i-th level, d gt is the ground truth disparity, d i is the predicted disparity map of the i-th level, The function is defined as:

[0034]

[0035] The present invention also provides a satellite image cascade matching system based on dual-branch context awareness, including:

[0036] A dual-branch feature extraction module for obtaining main features and context features;

[0037] A cost volume construction module for implementing population correlation and channel excitation calculation;

[0038] A multi-scale aggregation module containing cross-scale information transfer and attention fusion units;

[0039] A disparity optimization module integrating gradient calculation and a dual-branch encoder.

[0040] The present invention provides a satellite image cascade matching method and system based on dual-branch context awareness.

[0041] It has the following beneficial effects:

[0042] 1. The present invention constructs a two-branch architecture of a lightweight main feature network and a global context encoding network, which can effectively capture high-level semantic information while retaining the underlying texture details of satellite images, and solves the problem of insufficient feature representation ability of traditional single-branch networks in complex scenarios. The lightweight design of the main feature network significantly reduces the consumption of computing resources, while the context encoding network models long-range spatial dependence relationships through residual block stacking, enabling the network to adapt to the matching requirements of large parallax ranges and low-texture regions, and taking into account the advantages of both efficiency and accuracy.

[0043] 2. The present invention introduces a cost volume construction strategy of channel excitation weights and group correlation calculation, enabling the network to autonomously focus on key matching regions in the image (such as edges and texture jump regions), and suppressing false matches caused by illumination differences or repetitive structures. By fusing local detail features and global context information, the discrimination ability of the cost volume in low-texture regions (such as bare land and water surfaces) and repetitive pattern regions (such as regularly arranged building groups) is enhanced, significantly reducing the matching ambiguity.

[0044] 3. Based on cross-scale information transfer and multi-dimensional attention mechanism, the present invention dynamically fuses semantic information of different resolutions in the cost aggregation stage, effectively alleviating the matching ambiguity in occluded boundaries and disparity discontinuity regions. The attention mechanism enhances the matching signals in high-confidence regions through weighting, while suppressing noise interference, enabling the disparity estimation result to achieve natural reconstruction in smooth transition regions while maintaining the edge sharpness.

[0045] 4. The present invention extracts the image gradient features through the Sobel operator and combines the two-branch encoding-decoding structure to strengthen the geometric edge alignment ability in the disparity refinement stage. The gradient features provide the prior of pixel-level intensity changes, which are complementary to the context semantic features, enabling the network to accurately repair the local errors in the initial disparity map caused by occlusion or weak texture, especially suitable for the detail reconstruction of complex structures such as dense building facades and vegetation-covered areas in urban regions.

[0046] 5. From multi-scale feature extraction, cost volume construction to disparity refinement, the present invention adopts a cascaded processing flow to gradually optimize the matching results. This design makes full use of the multi-resolution characteristics of satellite images, quickly covering a large parallax range at the low-resolution stage and gradually refining the local accuracy at the high-resolution stage, thus meeting the dual requirements of wide-area coverage and high-resolution reconstruction of satellite images while ensuring the computational efficiency, and is applicable to practical application scenarios such as remote sensing mapping and disaster monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is one of the schematic diagrams of the method flow of the present invention;

[0048] Figure 2 is the other schematic diagram of the method flow of the present invention;

[0049] Figure 3 Schematic diagram of the network structure of the multi-scale information transmission strategy of the present invention;

[0050] Figure 4 Schematic diagram of the network structure of the multi-dimensional attention fusion of the present invention;

[0051] Figure 5 Disparity map predicted by the trained context-aware cascaded stereo matching network of the present invention on the WHU-Stereo dataset;

[0052] Figure 6 Disparity map predicted by the trained context-aware cascaded stereo matching network of the present invention on the US3D dataset;

[0053] Figure 7 Schematic diagram of the system structure of the present invention.

[0054] Among them, 10, dual-branch feature extraction module; 20, cost volume construction module; 30, multi-scale aggregation module; 40, disparity optimization module. Detailed implementation manners

[0055] Next, in combination with the accompanying drawings of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0056] Please refer to the atta Figure 1 - atta Figure 6 , the present invention provides a satellite image cascaded matching method based on dual-branch context awareness, which solves the matching problems of satellite images in low-texture, repetitive texture and disparity mutation regions by fusing multi-scale main features and context information and combining a gradient-guided disparity refinement strategy.

[0057] As Figure 1 shown, the satellite image cascaded matching method based on dual-branch context awareness may include the following steps:

[0058] S1. Extract the multi-scale main features and context features of the satellite image through a dual-branch network respectively;

[0059] S2. Construct a context-aware cost volume based on the spliced features;

[0060] S3. Optimize the cost volume and calculate the initial disparity map by using a multi-scale aggregation strategy;

[0061] S4. Refine the disparity map based on gradient information and context features.

[0062] The following is a detailed description of each step in the method of the present invention, comprehensively elaborating on the specific implementation principles, technical details, and processes for each step.

[0063] For step S1, in this embodiment, the dual-branch multi-scale feature extraction process of step S1 is implemented by constructing a heterogeneous feature extraction network, aiming to capture the underlying detail features and high-level semantic context information of satellite images respectively. The specific implementation is as follows:

[0064] The main feature network adopts a lightweight convolutional neural network architecture, and the preferred implementation is the MobileNetV2 model based on the inverted residual structure. This network is pre-trained on the ImageNet dataset to enhance its basic feature extraction ability.

[0065] The MobileNetV2 network contains multiple inverted residual modules, and each module is composed of a 1×1 convolutional upsampling layer, a 3×3 depthwise separable convolutional layer, and a linear bottleneck layer connected in sequence. Among them, the depthwise separable convolution decomposes the standard convolution into a depth convolution and a pointwise convolution, significantly reducing the computational complexity; the linear bottleneck layer is used to prevent the ReLU activation function from destroying low-dimensional features and retain more effective information.

[0066] In the feature extraction stage, the input left and right satellite images are respectively subjected to five downsampling operations with a stride of 2, reducing the original resolution to 1 / 32, and then gradually restored to 1 / 4, 1 / 8, and 1 / 16 scales through the upsampling blocks with skip connections, and multi-scale main features are output and (s = 4, 8, 16). The skip connection effectively alleviates the gradient disappearance problem and retains multi-scale details by concatenating the feature maps in the downsampling stage with the upsampling results along the channels.

[0067] The context encoding network is implemented by an encoder structure composed of residual blocks and downsampling layers, and its design goal is to extract the global semantic context information of the image. Each residual block contains two 3×3 standard convolutional layers, and the input and output are added through a skip connection to solve the degradation problem in the training of deep networks.

[0068] The downsampling layer uses a 3×3 convolutional operation with a stride of 2 to gradually reduce the resolution of the input image to 1 / 16 scale. During this process, the network outputs multi-scale context features and (s = 4, 8, 16), where the low-resolution features contain richer global semantic information, while the high-resolution features retain local details.

[0069] The context features and the main features are complementary in the deep layers of the network. The former focuses on the semantic associations of the scene, while the latter focuses on local texture details.

[0070] To fuse the complementary information of the main features and the context features, in this embodiment, a channel dimension concatenation operation is performed at the output stage of the dual-branch network. Specifically, for the feature maps of each scale, the main features and the context features are concatenated in the channel dimension to obtain the fused features The right figure features are generated in the same way. The mathematical expression of the concatenation operation is:

[0071]

[0072] where N F and N C represent the number of channels of the main features and the context features respectively, and H s and W s are the spatial resolutions of each scale. Through this operation, the fused features contain both low-level details and high-level semantic information, providing a more comprehensive feature expression basis for the subsequent cost volume construction.

[0073] It should be noted that the pre-trained weights of the main feature network are fine-tuned on satellite image data to adapt to the imaging characteristics of the remote sensing scene; while the context encoding network is trained from scratch to avoid feature bias caused by the difference in the pre-trained data distribution. This design enhances the adaptability to the unique patterns of satellite images (such as large parallax ranges and repetitive ground object structures) while ensuring the generality of the features.

[0074] For step S2, in this embodiment, the context-aware cost volume construction process of step S2 is achieved by fusing multi-source features and an adaptive attention mechanism, aiming to enhance the matching robustness of the cost volume to complex satellite image scenes. The specific implementation is as follows:

[0075] The population-related cost volume calculation is based on the channel grouping strategy. The concatenated left and right figure features and are divided into N g groups, and each group contains N c channels. For each group of features, the dot product of the feature vectors at the left figure position (h, w) and the right figure position (h, w - d) is calculated pixel by pixel along the disparity dimension d within the preset disparity search range to generate the initial population-related cost volume C gwc . Its mathematical expression is defined as:

[0076]

[0077] where · represents the vector dot product operation, that is, the corresponding eigenvalues of channels within the group are multiplied and then summed, and then normalized by the number of channels within the group. Through this grouping strategy, the network can reduce memory occupancy while retaining channel diversity, adapting to the matching requirements of large parallax ranges in satellite images. channels are multiplied and then summed, and then normalized by the number of channels within the group. Through this grouping strategy, the network can reduce memory occupancy while retaining channel diversity, adapting to the matching requirements of large parallax ranges in satellite images.

[0078] Furthermore, to enhance the matching response of the cost volume in key regions, a channel excitation mechanism is introduced to perform spatial attention weighting on C gwc . Specifically, the left image stitching features are used to generate a spatial weight map: the activation values at each position of the feature map are normalized to the [0,1] interval through the Sigmoid function to obtain the weight matrix W s . The weight is multiplied element-wise with the group-related cost volume to achieve adaptive spatial enhancement:

[0079]

[0080] where σ represents the Sigmoid activation function, and ⊙ is the element-wise multiplication operation. Through this operation, the network can autonomously focus on regions with rich texture or significant edges, suppressing low-confidence matching noise, especially suitable for local feature blurring scenarios caused by illumination differences in satellite images.

[0081] On this basis, a factorized 3D convolutional network is used to regularize the enhanced cost volume C′ cte,s to aggregate cross-parallax and cross-space context information. The regularization network decomposes the standard 3D convolution into two serial computational stages:

[0082] First, a 1×1×K d one-dimensional convolution is performed along the parallax dimension to capture the correlation between different parallax levels;

[0083] Subsequently, a K h ×K w ×1 two-dimensional convolution is performed in the spatial dimension to model the local spatial context.

[0084] This decomposition strategy reduces the number of parameters of the original 3D convolution from K d ×K h ×K w ×C in ×C out to (K d ×C in ×C mid )+(K h ×K w ×C mid ×C out ), where C midis the number of intermediate channels, thus reducing the computational cost while ensuring the feature aggregation ability. The ReLU activation function and the batch normalization layer are inserted between convolutional layers to avoid gradient vanishing and accelerate convergence.

[0085] It should be noted that the disparity search range d is dynamically adjusted according to the scale of the input image: a larger disparity range is set at the low-resolution scale (such as 1 / 16) to cover the wide-range depth changes, and the range is reduced at the high-resolution scale (such as 1 / 4) to achieve fine adjustment. This design adapts to the wide disparity distribution characteristics brought by the large baseline distance of satellite images, avoiding the efficiency degradation caused by redundant disparity search in the high-resolution stage.

[0086] Through the above-mentioned group-related calculations, channel excitation enhancement, and factorization regularization operations, the multi-scale context-aware cost volume C is finally generated. cte,s . This cost volume integrates local matching costs, spatial attention weights, and cross-disparity context information, providing a highly discriminative matching signal for subsequent disparity regression, especially having significant improvement potential for repetitive texture regions (such as regularly arranged building groups) and occlusion boundary regions (such as terrain mutation zones).

[0087] For step S3, in this embodiment, the multi-scale cost aggregation and initial disparity calculation process in step S3 are implemented through a cross-scale information transfer and multi-dimensional attention fusion strategy, aiming to optimize the discriminability of the matching cost volume and generate a high-precision initial disparity estimate. The specific implementation method is as follows:

[0088] As Figure 3 shown, the multi-scale information transfer strategy is designed based on a cascaded architecture, guiding the semantic information at the low-resolution scale to the matching process at the high-resolution scale step by step. Specifically, starting from the lowest resolution scale (such as 1 / 16), the regularized context-aware cost volume C cte,s is upsampled using the bilinear interpolation algorithm to align its spatial size with the adjacent higher resolution scale (such as 1 / 8). The upsampled low-resolution cost volume and the cost volume at the current scale are added element-wise to achieve cross-scale semantic information fusion. This operation can be expressed as:

[0089] C agg,s = Upsample(C cte,s-1 ) + C cte,s

[0090] where s represents the current scale index, and Upsample(·) represents the bilinear upsampling operation. Through this strategy, the global scene structure information captured in the low-resolution stage is effectively transferred to the high-resolution stage to assist in solving the matching ambiguity problem caused by local texture loss.

[0091] Furthermore, as Figure 4As shown, the multi-dimensional attention fusion mechanism introduces parallel attention branches in the disparity dimension to dynamically adjust the matching weights of different positions and disparity candidates. In specific implementation, for the fusion cost volume C at each scale agg,s , the global statistics and local context features are calculated along the disparity dimension respectively:

[0092] 1. Global attention branch:

[0093] Global average pooling is performed on the cost volume along the disparity dimension to generate a spatial attention map Its mathematical expression is:

[0094]

[0095] This branch focuses on the average matching response within the entire disparity search range, reflecting the overall matching confidence of different spatial positions in the scene.

[0096] 2. Local attention branch:

[0097] A 3×3 depthwise separable convolutional layer is used to extract local features from the cost volume to generate a spatial-disparity joint attention map The convolutional operation slides in the spatial dimension while keeping the disparity dimension unchanged, thereby modeling the correlation of disparity values within the local neighborhood.

[0098] The global and local attention maps are fused into the original cost volume through broadcast multiplication and weighted summation operations:

[0099]

[0100] Among them, represents element-wise multiplication broadcast along the spatial dimension. Through this design, the network can adaptively enhance the matching signals in high-confidence regions while suppressing the noise interference in occluded or weak-texture regions.

[0101] In the disparity regression stage, the probability distribution is calculated along the disparity dimension for the cost volume C optimized by attention att,s In a preferred implementation, the Softmax function is used to normalize the matching cost, and the initial disparity estimate of each pixel is obtained through the calculation of the expected value:

[0102]

[0103] This operation converts discrete disparity candidates into continuous-value predictions, effectively improving the sub-pixel level disparity estimation accuracy.

[0104] It should be noted that the disparity search range D sDynamic adjustment according to scale: Set a larger disparity range at the low-resolution stage (e.g., 1 / 16 scale) to cover the potential depth changes of wide-baseline satellite images; gradually narrow the range at the high-resolution stage (e.g., 1 / 4 scale) to focus on local disparity refinement. This design significantly reduces the computational complexity while ensuring the matching accuracy.

[0105] Through the above cross-scale information transfer, multi-dimensional attention fusion, and probabilistic disparity regression operations, this embodiment finally outputs a multi-scale initial disparity map, providing a reliable basic disparity estimate for the subsequent refinement module. The cascaded aggregation strategy makes full use of the complementarity of multi-scale features and is especially suitable for regions with discontinuous disparities caused by terrain undulation, building occlusion, etc. in satellite images.

[0106] For step S4, in this embodiment, the gradient-guided disparity refinement process in step S4 is achieved by fusing geometric edge features and semantic context information, aiming to correct the local errors of the initial disparity map and restore the detail accuracy. The specific implementation is as follows:

[0107] The gradient map is generated by extracting the edge features of the image based on the Sobel operator. The convolution kernels in the horizontal and vertical directions are respectively defined as:

[0108]

[0109] Apply and convolution operations to the left image respectively to obtain the horizontal gradient map Gx and the vertical gradient map Gy. The two are concatenated along the channel dimension to form a 2-channel gradient map G = Concat(G x , G y ). The gradient map provides a geometric structure prior for disparity refinement by capturing the regions of intensity mutation in the image (such as building contours, terrain boundaries).

[0110] Furthermore, a dual-branch encoder structure is constructed to fuse edge features and context information:

[0111] Edge feature branch: Concatenate the initial disparity map (highest resolution scale, such as 1 / 4) and the gradient map G in the channel dimension to obtain the input feature Extract the edge-guided feature F edge through three 3×3 convolutional layers. Each layer is followed by batch normalization and the ReLU activation function, gradually increasing the number of channels to align with the context feature.

[0112] Context feature branch: Reuse the 1 / 4-scale feature output by the context encoding network Adjust the channel dimension through a 1×1 convolutional layer to obtain the context feature This branch retains the global semantic association information of the scene, assisting in correcting the parallax estimation errors caused by occlusion or weak texture.

[0113] The dual-branch features are gradually fused through a multi-level residual decoder and restored to high resolution. The decoder consists of four levels of upsampling units, and the operations at each level are as follows:

[0114] 1. Transposed convolution upsampling: Perform a 2-fold transposed convolution operation on the input features, and the spatial resolution is increased to 2 times that of the previous level.

[0115] 2. Skip connection fusion: Add the edge feature F dege and the context feature F ctx element-wise to generate the fused feature F fuse = F edge + F ctx .

[0116] 3. Residual convolution refinement: Apply a 3×3 convolution, batch normalization, and ReLU activation to the fused feature, and then add it to the feature before upsampling through a skip connection to suppress the noise introduced by upsampling.

[0117] Finally, after four levels of upsampling operations, the decoder outputs a full-resolution refined parallax map The multi-level residual connection design ensures that the low-level geometric details and high-level semantic information are gradually fused during the decoding process, effectively avoiding the edge blurring problem caused by multiple upsamplings.

[0118] It should be noted that the parallel processing mechanism of the edge feature branch and the context feature branch has clear complementarity: the edge features strengthen the local consistency of the parallax mutation region, while the context features provide the global structural constraints of the scene. The two achieve feature interaction through the addition operation, enabling the network to locally optimize the occlusion boundaries and thin structure regions (such as wires, vegetation) while retaining the overall distribution of the initial parallax.

[0119] Through the above gradient-guided dual-branch encoding and multi-level residual decoding process, this embodiment can significantly improve the detail integrity and geometric consistency of the satellite image parallax map in complex scenes, and is particularly suitable for the 3D reconstruction task of dense building groups in urban areas.

[0120] In this embodiment, the loss function design and the model training process are realized based on a multi-scale supervision and adaptive optimization strategy to ensure the balanced learning of the network for feature levels at different training stages. The specific implementation is as follows:

[0121] In a preferred implementation, a weighted smooth L1 loss function is used to jointly supervise the multi-scale initial parallax map and the refined parallax map. The mathematical expression of the loss function is:

[0122]

[0123] Wherein:

[0124] d gt represents the true value of parallax (Ground Truth).

[0125] d i represents the predicted parallax map at the i-th level, including three multi-scale initial parallax maps (i = 1, 2, 3) and the final refined parallax map (i = 4).

[0126] λ i is the loss weight coefficient for each level. The weights from low to high levels are λ1, λ2, λ3, λ4, which are 0.5, 0.7, 1, 1 respectively.

[0127] Smooth L1 loss function is defined as:[[]]

[0128]

[0129] This function uses the L2 norm constraint when the error is small to improve the convergence stability, and switches to the L1 norm when the error is large to reduce the interference of outliers, thus balancing the robustness and accuracy requirements in the training process.

[0130] In the preferred implementation, the model training is based on the satellite remote sensing stereo matching dedicated datasets US3D and WHU-Stereo. Among them, the US3D dataset contains 2139 pairs of stereo images, and the parallax search range is set to [-96, 96]; the WHU-Stereo dataset contains 1757 pairs of images, and the parallax range is set to [-128, 64]. The following strategies are adopted in the training stage:

[0131] 1. Data preprocessing

[0132] Normalize the input image intensity values to the range [0, 1] to eliminate the brightness deviation caused by sensor differences.

[0133] Apply asymmetric data augmentation to the left and right images independently, including random contrast adjustment (scaling factor range [0.5, 1.5]), gamma correction (range [0.8, 1.2]), and brightness offset (range [-0.2, 0.2]) to simulate the illumination changes in satellite imaging.

[0134] Random vertical flipping augmentation with a probability set to 0.5 to improve the generalization ability of the model to the terrain undulation direction.

[0135] 2. Optimizer configuration

[0136] The AdamW optimizer is adopted, with its momentum parameters set as β1 = 0.9, β2 = 0.999, and the weight decay coefficient as 0.01 to prevent overfitting.

[0137] The initial learning rate is set to 1×10 -4 , and it is dynamically adjusted based on the cosine annealing strategy, with the minimum learning rate reduced to 1×10 -6 to balance the convergence speed and stability.

[0138] 3. Training Parameters

[0139] The batch size is set to 4 to adapt to the GPU video memory limit and ensure the reliability of gradient estimation.

[0140] The number of training epochs is set to 300, and the early stopping mechanism monitors the loss on the validation set. If there is no improvement for 20 consecutive epochs, the training is terminated.

[0141] Through the above loss function design and training strategy, this embodiment can effectively drive the network to learn the coarse-to-fine disparity estimation ability, while taking into account the large disparity range characteristics of satellite images and the generalization requirements of complex scenes.

[0142] As Figure 5 and Figure 6 shown, the disparity maps predicted by the trained context-aware cascaded stereo matching network of this embodiment on the WHU-Stereo and US3D datasets are respectively presented.

[0143] Generally speaking, the present invention constructs a heterogeneous feature extraction network (the main feature network uses lightweight MobileNetV2 to extract multi-scale detailed features, and the context encoding network captures global semantic information based on residual blocks), fuses multi-scale main features and context features to construct a population-related cost volume, strengthens the matching response in key regions by combining channel excitation weights, and uses a multi-scale attention aggregation strategy to optimize the cost volume to generate an initial disparity map; further introduces a gradient-guided double-branch refinement module, extracts the image edge features through the Sobel operator and cascades them with the context features, and restores high-resolution disparity details through a multi-level residual decoder, finally realizing high-precision 3D reconstruction of satellite images in low-texture, repetitive structure, and disparity mutation regions.

[0144] The satellite image cascaded matching system based on double-branch context awareness described below can be correspondingly referred to the method of satellite image cascaded matching based on double-branch context awareness described above.

[0145] Please refer to the attached Figure 7 , the present invention also provides a satellite image cascaded matching system based on double-branch context awareness, including:

[0146] The dual-branch feature extraction module 10 is used to obtain the main feature and the context feature;

[0147] The cost volume construction module 20 implements the calculation of population correlation and channel excitation;

[0148] The multi-scale aggregation module 30 includes a cross-scale information transmission and attention fusion unit;

[0149] The disparity optimization module 40 integrates gradient calculation and a dual-branch encoder.

[0150] The system of this embodiment can be used to execute the above method embodiment, and its principle and technical effect are similar, which will not be elaborated here.

[0151] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A satellite image cascade matching method based on dual-branch context awareness, characterized in that It includes the following steps: Extract the multi-scale main features and context features of the satellite image through a dual-branch network respectively; Fuse the main features and context features to construct a context-aware cost volume; Adopt a multi-scale aggregation strategy to optimize the cost volume and calculate the initial disparity map; Refine the disparity map based on gradient information and context features.

2. The satellite image cascade matching method based on double-branch context awareness according to claim 1, wherein In the step of extracting the multi-scale main features and context features of the satellite image through the dual-branch network respectively: The main features are extracted through a lightweight convolutional neural network, and the network structure includes an inverted residual module; The context features are extracted through an encoder structure, including residual blocks and downsampling layers.

3. The satellite image cascade matching method based on dual-branch context awareness according to claim 2, wherein, The lightweight convolutional neural network is MobileNetV2 and is pre-trained on the ImageNet dataset.

4. The satellite image cascade matching method based on dual-branch context awareness according to claim 1, wherein, The step of fusing the main features and context features to construct a context-aware cost volume includes: Concatenate the main features and context features in the channel dimension, and calculate the population correlation cost volume based on the concatenated features; Generate channel excitation weights using the left image features, and enhance the spatial attention of the population correlation cost volume through the weights; Perform regularization processing on the enhanced cost volume to obtain the optimized context-aware cost volume.

5. The satellite image cascade matching method based on dual-branch context awareness according to claim 4, wherein The population correlation cost volume is calculated by the following formula: Among them, C gwc (g, d, h, w) represents the correlation value at the g-th group, candidate disparity d, and spatial position (h, w) in the group-related cost volume; N c represents the total number of channels of the spliced feature map ; N g represents the number of groups for group-related calculation; represents the spliced feature value of the left image at the position (h, w); represents the spliced feature value of the right image at the position (h, w - d).

6. The satellite image cascade matching method based on dual-branch context awareness according to claim 1, wherein The step of adopting a multi-scale aggregation strategy to optimize the cost volume and calculate the initial disparity map includes: Gradually transfer the context-aware cost volumes of different scales from low resolution to high resolution, and superimpose the semantic information of the low-resolution cost volume on the high-resolution cost volume through upsampling operations; Introduce an attention mechanism in the disparity dimension, calculate the global attention vector and local attention vector respectively, and dynamically adjust the matching weights of different positions and disparity ranges through weighted fusion.

7. The satellite image cascade matching method based on dual-branch context awareness according to claim 1, wherein The step of refining the disparity map based on gradient information and context features includes: Calculate the gradient map of the left image through a gradient operator, and concatenate the initial disparity map and the gradient map in the channel dimension to obtain edge features; Construct a dual-branch encoding structure, where the first branch processes the edge features and the second branch reuses the extracted context features; Fuse the dual-branch features through a multi-level residual decoder, and gradually upsample to generate a refined disparity map.

8. The satellite image cascade matching method based on dual-branch context awareness according to claim 7, wherein The gradient operator is the Sobel operator, which is used to extract the horizontal and vertical gradients of the image, and the gradient map and the initial disparity map are concatenated and input into the first branch.

9. The satellite image cascade matching method based on dual-branch context awareness according to claim 1, wherein The method further includes: Adopt a weighted smooth L1 loss function to supervise and train the multi-scale initial disparity map and the refined disparity map, and the loss function satisfies: where n is the total number of output levels, λ i is the weight coefficient of the i-th level, d gt is the true value of the disparity, d i is the predicted disparity map of the i-th level, The function is defined as:

10. A satellite image cascade matching system based on dual-branch context awareness for performing the method according to any one of claims 1-9, characterized in that, It includes: A dual-branch feature extraction module for obtaining main features and context features; A cost volume construction module for realizing population correlation and channel excitation calculation; A multi-scale aggregation module including a cross-scale information transfer and attention fusion unit; A disparity optimization module integrating gradient calculation and a dual-branch encoder.

Citation Information

Cited By

  • Satellite image ground object height calculation method and system based on monocular parallax estimation

    CN120997275A

  • Satellite image ground object height calculation method and system based on monocular disparity estimation

    CN120997275B

  • Satisfaction evaluation method and system based on multi-dimensional data fusion

    CN121073271A

  • Binocular stereo matching method and system based on cross-scale cost fusion and edge perception enhancement

    CN122617947A