Semantic segmentation and matching method of image

By improving the semantic segmentation and feature matching module of the BGA-Net network, combining multi-scale feature pyramids and bidirectional guidance of attention, the problem of generating three-dimensional models with large differences in land and objects in remote sensing images is solved, and more accurate semantic segmentation and intensive matching effects are achieved.

CN120259663AActive Publication Date: 2025-07-04SHANDONG UNIV OF SCI & TECH

Patent Information

Application Number
CN202510350531.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-04
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

When the existing semantic segmentation and intensive matching methods deal with remote sensing images with large differences in land objects caused by seasonal changes and other reasons, it is difficult to generate high-precision three-dimensional models, and the semantic segmentation module ignores local context features, resulting in inaccurate image details segmentation.

Method used

The semantic segmentation module of the BGA-Net network is improved, and features of different scales are weighted through the attention mechanism of multiple convolution branches, combined with the feature pyramid module and the bidirectional guidance attention module to enhance the complementarity between global semantics and local details, and the feature matching module adopts a stacked hourglass network to improve efficiency.

Benefits of technology

It realizes more accurate image semantic segmentation results, generates a higher precision three-dimensional model, and improves the accuracy and efficiency of semantic segmentation and intensive matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259663A_ABST
    Figure CN120259663A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic segmentation and matching method of an image, and relates to the technical field of image processing. According to the method, a BGA-Net semantic segmentation module is improved, and a more accurate image semantic segmentation result is provided for generating a three-dimensional model: pixel information of an image is gradually reduced along a feature scale reduction direction, and semantic information is gradually increased; therefore, first initial features of different scales are weighted through an attention mechanism of a plurality of convolution branches, and pixel information in large-scale features and semantic information in small-scale features are focused at the same time; therefore, the output of the previous-scale decoder is added to the input of each convolution branch decoder, pixel information in the large-scale features can be transmitted to the decoding process of the small-scale features, complementation of global semantics and local details is achieved, and a more accurate semantic segmentation result is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to a method for semantic segmentation and matching of images. Background Art

[0002] Semantic segmentation and dense matching are hot research fields. As classic visual single-task learning problems, semantic segmentation and dense matching have been studied for many years. With the wave of deep learning in artificial intelligence, many researchers have designed simple and advanced networks to improve performance and made good progress in these two tasks. In the process of generating a 3D model based on remote sensing images, semantic segmentation helps to identify land types (such as forests, urban areas, water bodies, etc.) in remote sensing images, providing basic structural information for subsequent 3D modeling; dense matching is to establish corresponding relationships between the segmented regions, estimate their spatial positions and directions, and reconstruct surface details by comparing and analyzing the similarities between images. However, most of the existing methods separately process these two tasks or use a combination of multiple models to solve these two tasks. Due to the lack of sufficient information sharing and fusion, it is difficult to process remote sensing images with large differences in ground objects caused by seasonal changes, etc., resulting in difficulties in generating high-precision 3D models.

[0003] In recent years, multi-task fusion networks for semantic segmentation and dense matching, such as BGA-Net (Bidirectional Guided Attention Network), have been proposed. In order to bridge the gap between the two image semantic segmentation tasks and the dense matching task, BGA-Net first applies a shared unified backbone module to extract primary features from two images and generate two corresponding feature maps; then designs a bidirectional guided attention module as the bridge between the semantic segmentation task and the dense matching task. The goal of this bidirectional guided attention module is to extract the global information of the image from one of the two feature maps and generate an attention map; for the semantic segmentation task, this attention map will assign more weights to the objects (such as buildings, roads, and trees, etc.) in the image and less weights to the background. For the dense matching task, since each object in the image is regarded as a rigid surface and the same object has the same or similar disparity, this attention map will assign more weights to the disparity of the same object; under the guidance of the attention map, the semantic segmentation module and the feature matching module in BGA-Net respectively perform the semantic segmentation task and the dense matching task.

[0004] However, the semantic segmentation module pays more attention to the global context features of the image when performing the semantic segmentation task, while ignoring the image spatial detail information contained in the local context features, resulting in inaccurate segmentation of image details. Summary of the Invention

[0005] Based on this, in order to solve the technical problems in the prior art, the present invention provides a method for semantic segmentation and matching of images.

[0006] The present invention provides a method for semantic segmentation and matching of images, including: Replacing the semantic segmentation module in the original BGA-Net network with an improved semantic segmentation module, the improved semantic segmentation module includes a plurality of convolutional branches, and each convolutional branch includes an attention mechanism module and a decoder module connected in sequence; the output end of the decoder in each convolutional branch is connected to the input end of the decoder module in the next convolutional branch along the direction of decreasing scale; the attention mechanism module includes a parallel dot product attention branch and a Conv-BN-ReLU branch; replacing the feature matching module in the original BGA-Net network with an improved feature matching module, the improved feature matching module includes a convolutional module, a first encoding-decoding module, a second encoding-decoding module, a third encoding-decoding module, a bilinear interpolation layer, and a regression layer connected in sequence; adding a feature pyramid module between the original unified backbone module and the improved semantic segmentation module; constructing an improved BGA-Net network including the unified backbone module, the feature pyramid module, the bidirectional guiding attention module, the improved semantic segmentation module, the improved feature matching module, and the bidirectional attention module in the original BGA-Net network.

[0007] Collecting image data to construct a data set, and using the data set to train the improved BGA-Net network to obtain a semantic segmentation-dense matching model.

[0008] Input the first image and the second image to be processed into the semantic segmentation-dense matching model. Through the unified backbone module, perform preliminary feature extraction on the first image and the second image to obtain the first initial features at different scales and the second initial features at different scales; through the bidirectional guiding attention module, perform context feature extraction and weighting on the first initial features at different scales to obtain the bidirectional attention map; through the feature pyramid module, perform convolutional operations and transposed convolutional operations at different scales to enhance the semantic information in the first initial features at different scales, so as to input the enhanced first initial features into the improved semantic segmentation module; through the attention mechanism modules of several convolutional branches in the improved semantic segmentation module, perform attention weighting operations on the first initial features at different scales to obtain semantic enhanced feature maps at different scales; through the decoder module, decode the semantic enhanced feature maps at different scales to obtain decoded feature maps at different scales; wherein, the input of each decoder module also includes the output from the decoder module of the convolutional branch at the previous scale; by adding the bidirectional attention map to the decoded feature map at the smallest scale, obtain the initial segmentation map; through the feature matching module, integrate the second initial features at different scales into the cost volume, and perform dense matching on the cost volume and the bidirectional attention map to obtain the initial disparity map; through the bidirectional attention module, perform cross-fusion on the initial segmentation map and the initial disparity map to obtain the final segmentation map and the final disparity map.

[0009] Further, the specific steps for the bidirectional guiding attention module to perform context feature extraction and weighting include: After transforming the feature map with the shape of H×W×F input to the bidirectional guiding attention module, apply the softmax function to obtain the weight distribution. Transform the feature map with the shape of H×W×F into F×H×W, weight the feature map transformed into F×H×W with the weight distribution output by the softmax function; and perform a convolutional operation on the weighted feature map. Fuse the result of the convolutional operation with the feature map input to the bidirectional guiding attention module, and output the bidirectional attention map.

[0010] Further, the steps for the attention mechanism module to perform attention weighting operations specifically include: Through the dot-product attention branch, perform feature transformation on the input of the attention mechanism module to obtain matrices Q, K, and V. Use the softmax function to calculate the normalized similarity weights of each element in the input of the attention mechanism module relative to all other elements, and multiply the similarity weights by matrix V to generate the weight matrix of self-attention: Among them, is the input of the attention mechanism module, C is the number of channels, is the weight matrix of self-attention, and the superscript T represents the transpose operation; Perform convolution, batch normalization, and ReLU function activation operations on the first initial feature in sequence through the Conv-BN-ReLU branch: Fuse the output of the dot product attention branch and the Conv-BN-ReLU branch to obtain the output of the attention mechanism module.

[0011] Furthermore, the specific steps for obtaining the initial disparity map include: Perform three-dimensional convolution operation on the cost volume through the convolution module; obtain the first disparity map; Perform encoding and decoding operations on the first disparity map step by step through the first encoder-decoder module, the second encoder-decoder module, and the third encoder-decoder module. During the decoding process, achieve dense matching by introducing the bidirectional attention map output by the bidirectional guidance attention module to obtain the dense matching feature map; Interpolate the dense matching feature map through the bilinear interpolation layer, and perform regression on the interpolation result through the regression layer to output the initial disparity map.

[0012] Furthermore, the realization of dense matching by introducing the bidirectional attention map output by the bidirectional guidance attention module is achieved by fusing the bidirectional attention map with the output of the encoding layer in the encoder-decoder module and then inputting it into the decoding layer in the encoder-decoder module for decoding.

[0013] The above at least one technical solution adopted by the present invention can achieve the following beneficial effects: In the method for semantic segmentation and matching of images provided by the present invention, by improving the semantic segmentation module of BGA-Net, richer local detail features are obtained, providing a more accurate image semantic segmentation result for generating a three-dimensional model. Specifically: Since along the direction of decreasing feature scale of the first initial features of different scales, the pixel information of the image gradually decreases and the semantic information gradually increases, by weighting the first initial features of different scales through the attention mechanism of multiple convolution branches, it is possible to simultaneously focus on the pixel information in the large-scale features and the semantic information in the small-scale features; therefore, by adding the output of the decoder of the previous scale to the input of the decoder of each convolution branch, the pixel information in the larger-scale features can be effectively transmitted to the decoding process of the smaller-scale features, realizing the complementarity of global semantics and local details, and thus obtaining a more accurate image semantic segmentation result. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0015] Figure 1 It is a schematic flowchart of the method for semantic segmentation and matching of the image provided by the present invention; Figure 2 It is a schematic diagram of the original BGA-Net network structure provided by the present invention; Figure 3 It is a schematic diagram of the improved BGA-Net network structure provided by the present invention; Figure 4 It is a schematic diagram of the structure of the feature extraction module provided by the present invention; Figure 5 It is a schematic diagram of the GC block structure provided by the present invention; Figure 6 It is a schematic diagram of the structure of the improved semantic segmentation module provided by the present invention; Figure 7 It is a schematic diagram of the structure of the improved feature matching module provided by the present invention; Figure 8 It is a schematic diagram of the structure of the bidirectional attention module provided by the present invention. Detailed implementation manners

[0016] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0017] Rao et al. proposed a bidirectional guided attention network for 3D semantic detection of remote sensing images: BGA-Net. The network structure is as Figure 2 shown, including five modules: UBM (Unified Backbone Module), BGAM (Bidirectional Guided Attention Module), SSM (Semantic Segmentation Module), FMM (Feature Matching Module), and BFM (Bidirectional Fusion Module). The task of UBM is: First, the left and right images respectively pass through the same backbone network for feature extraction to obtain the feature map Xl and X r , these feature maps contain the low-level feature information of the image. Then, perform multi-scale feature fusion on X l and X r respectively to extract richer semantic and context information. After UBM processing, the unified feature maps X l and X r of the left and right images are obtained.

[0018] The task of SSM is: After transforming the features of X l and X r , extract multi-scale features and fuse information at different scales to obtain the semantic feature map S; then use the semantic segmentation attention map M' generated by BGAM to weight S to highlight the target object and obtain the activated feature map S'. Finally, upsample the size of S' to the size of the input image and generate the initial segmentation map through the convolutional layer. The task of FMM is: After transforming the features of X l and X r , calculate the similarity of the left and right feature maps at different disparities to construct the cost volume V, then use 3D convolution and NL block (Nonlocal block) to regularize the cost volume, and guide the matching process through the attention mechanism to obtain the regularized cost volume V'; finally, perform softmax operation on V' to calculate the probability of each disparity and obtain the initial disparity map through weighted summation. The task of BFM is: Concatenate the initial segmentation map, the input image, and the initial disparity map and fuse them through the residual learning network to obtain the refined segmentation map; fuse the initial disparity map to obtain the refined disparity map. Through the above modules, BGA-Net can simultaneously obtain the semantic segmentation map and the disparity map of the image, and use the bidirectional attention mechanism and the fusion module to improve the quality and robustness of the results.

[0019] SSM consists of three parts: a feature transformation part, a spatial feature learning part, and an upsampling part. The feature transformation part uses a two-dimensional convolutional layer with 256 channels and a 128-channel one to reduce the feature dimension, and then uses a 128-channel residual block and 3×3 filters to normalize the feature map. The spatial feature learning part first applies 16 bottleneck blocks with 128 channels to process the output of the feature transformation part, and then uses SPP (Spatial Pyramid Pooling) to divide it into different sub-regions; next, hierarchical context information and deep semantic feature maps are connected, and after fusing the semantic feature maps obtained through a 128-channel, a 64-channel 2D convolutional operation and 3×3 filters, the semantic feature maps are activated through an attention map to obtain activated semantic feature maps. Although SPP can capture features at different scales, the pooling operation will reduce the spatial resolution of the feature map, which may lead to the loss of detailed information, and the step of connecting hierarchical context information and deep semantic feature maps focuses more on combining global features and features at different scales rather than fine-grained local information. Therefore, SSM mainly extracts global features. Although this operation helps to understand the overall structure, it may not be sufficient to retain the fine details of the image. In summary, the semantic segmentation module of BGANet focuses on global context information, which is crucial for semantic information in complex scenes, but local context information also plays a key role in preserving rich spatial details. And in the feature matching module, NL-block (Nonlocal Block) and GC block (Global Context block) are mainly used to address the problem of long-distance information transmission and improve long-range dependencies, but they have a large computational cost, high memory consumption, and are prone to information redundancy.

[0020] Based on this, the present invention improves BGA-Net: using an attention mechanism to increase the connection between semantic segmentation and dense matching tasks, while optimizing the semantic information extraction module to comprehensively utilize global and local context information for semantic information extraction; optimizing the efficiency of the feature matching module, and using a simple and efficient stacked hourglass network for cost aggregation to improve the accuracy and efficiency of the multi-task learning network.

[0021] Embodiment 1 Figure 1 Shows the process flow of semantic segmentation and matching of the image in this embodiment. The following specifically combines Figure 1 to elaborate on this method in detail, which specifically includes the following steps:

[0022] S1: Replace the semantic segmentation module in the original BGA-Net network with an improved semantic segmentation module. The improved semantic segmentation module includes several convolutional branches, and each convolutional branch includes an attention mechanism module and a decoder module connected in sequence. The output end of the decoder in each convolutional branch is connected to the input end of the decoder module in the next convolutional branch in the direction of decreasing scale. Construct an improved BGA-Net network including the unified backbone module, bidirectional guiding attention module, improved semantic segmentation module, feature matching module, and bidirectional attention module in the original BGA-Net network.

[0023] Figure 3 The constructed improved BGA-Net network structure is shown, which is mainly composed of a feature extraction module, a bidirectional guiding attention module, an improved semantic segmentation module, an improved feature matching module, and a bidirectional attention module. First, use the feature extraction module to extract features from the left image and the right image respectively, and then use the left image features to apply to the bidirectional attention module to share global features and guide the semantic segmentation module and the feature matching module to learn, outputting an initial semantic map and an initial disparity map. Finally, use the bidirectional attention module to calculate the residuals of the initial segmentation map and the disparity map to refine the segmentation map and the disparity map to further improve the accuracy. More specifically:

[0024] Feature extraction module.

[0025] To increase the connection between the dense matching and semantic segmentation modules, the present invention performs unified feature extraction on the left and right images, which mainly consists of two parts, a unary feature extraction part and a multi-scale feature fusion part, as Figure 4As shown, a convolution operation is performed on the first image through the first convolution kernel. The size of the first convolution kernel is 3x3, the number of channels is 16, and the stride is 2. A convolution operation is performed on the output of the first convolution kernel through the first convolution block. The first convolution block includes three 3×3 convolution kernels, and the numbers of channels of the three 3×3 convolution kernels are 16, 16, and 32 respectively. A convolution operation is performed on the output of the first convolution block through the second convolution block. The second convolution block includes three 3×3 convolution kernels, and the numbers of channels of the three 3×3 convolution kernels are all 32. A convolution operation is performed on the output of the second convolution block through the third convolution block. The third convolution block includes sixteen 3×3 convolution kernels with 32 channels, sixteen 3×3 convolution kernels with 32 channels, sixteen 3×3 convolution kernels with 32 channels, and one 3×3 convolution kernel with 64 channels. A convolution operation is performed on the output of the third convolution block through the fourth convolution block. The fourth convolution block includes three 3×3 convolution kernels, and the numbers of channels of the three 3×3 convolution kernels are all 92, and the strides are 2, 2, and 1 respectively. A convolution operation is performed on the output of the fourth convolution block through the fifth convolution block. The fifth convolution block includes three 3×3 convolution kernels, and the numbers of channels of the three 3×3 convolution kernels are all 128, and the strides are 4, 4, and 1 respectively. Convolution operations are respectively performed on the output of the fourth convolution block through the second convolution kernel, the third convolution kernel, the fourth convolution kernel, and the fifth convolution kernel to obtain outputs of four scales. Among them, the sizes, numbers of channels, and strides of the second convolution kernel, the third convolution kernel, the fourth convolution kernel, and the fifth convolution kernel are {1×1, 16, 1}, {3×3, 16, 6}, {3×3, 16, 12}, and {3×3, 16, 18} respectively. After merging the outputs of the third convolution block, the fourth convolution block, the fifth convolution block, the second convolution kernel, the third convolution kernel, the fourth convolution kernel, and the fifth convolution kernel, convolution operations are performed using two convolution kernels with 256 and 384 channels respectively to obtain the output. The squares are respectively the convolution kernel size, the output dimension, and the stride. If the stride is not shown, it is defaulted to 1.

[0026] After feature extraction, feature transformation needs to be further performed, which is respectively applied to the bidirectional attention module, the feature matching module, and the semantic segmentation module. For the dense matching and bidirectional attention modules, a series of two-dimensional convolution operations are used to reduce the feature dimension. For the semantic segmentation module, a multi-scale feature pyramid is constructed.

[0027] Bidirectional guiding attention module.

[0028] The bidirectional guiding attention module includes three parts: the feature transformation part, the global context extraction part, and the bidirectional attention part. The first step of the feature transformation part is to reduce the feature dimension through a series of 2-D convolution operations, and a residual block with 64 channels is used to regularize the feature map. The global context extraction part is to capture the global context information through a group of global context information blocks, and the computational amount in object detection is significantly reduced. The process is as Figure 5As shown, first, based on the input feature map with the shape of H×W×F, where H is the height, W is the width, and F is the number of channels, after transforming the feature map with the shape of H×W×F of the input bidirectional guiding attention module, the softmax function is applied to obtain the weight distribution, so as to form context-related features by calculating the global attention map and sharing it for all positions. Secondly, the feature map with the shape of H×W×F is transformed into F×H×W, and the feature map transformed into F×H×W is weighted by the weight distribution output by the softmax function; and a convolution operation is performed on the weighted feature map for feature transformation to capture the dependencies in the channel manner. Finally, the result of the convolution operation is fused with the feature map of the input bidirectional guiding attention module, and the global context features are fused into the features at all positions, and the output shape is still H×W×F. The last step is to generate a bidirectional attention map for semantic segmentation and dense matching through the sigmoid function and two-dimensional convolution operation. For semantic segmentation, the semantic attention map will assign more weights to the ground objects (such as buildings, roads, and trees). For dense matching, each object is usually regarded as a rigid surface; therefore, the same object has the same or similar disparity. The disparity attention map will assign different weights within the disparity range.

[0029] The role of the bidirectional attention map in the semantic segmentation module is as follows: The semantic weights in the bidirectional attention map (such as buildings and roads) directly enhance the features related to the ground objects, making the segmentation module pay more attention to the category saliency regions and suppressing background interference at the same time. The role in the dense matching module is as follows: The disparity attention map assigns similar weights to the same rigid object (such as the same disparity on the vehicle surface), reduces the matching ambiguity, and improves the continuity of disparity prediction. By complementary bidirectional information, it provides scene understanding (global) for the semantic attention map and geometric constraints (local) for the disparity attention map. The two further cooperate through the cross-fusion module to ensure the consistency of the segmentation and disparity results in terms of space and semantics.

[0030] Semantic segmentation module.

[0031] The semantic segmentation module applies a series of two-dimensional convolutional operations to generate an initial segmentation map. Multi-scale information has been proven to be helpful for semantic segmentation in scene understanding. Therefore, the present invention uses constructing a feature pyramid to extract multi-scale information. The multi-scale feature extraction performs convolutional operations on the input image at different scales to generate feature maps with different resolutions. Top-down path: Starting from the high-resolution feature map, the size of the feature map is gradually increased through upsampling operations, and at the same time, element-wise addition or concatenation is performed with the high-resolution feature map from the lower layer to restore fine-grained information. Lateral connection: In the top-down path, the upsampling output of each layer is fused with the original-resolution feature map of the same layer (i.e., the input of the lateral connection), which is usually achieved through 1x1 convolution to reduce the number of channels of the feature map to be the same as that of the upsampled feature map, and then element-wise addition is performed. Convolutional fusion: After the lateral connection is completed, a 3x3 convolutional operation is performed on the fused feature map to further refine the features and enhance the representation ability of the features.

[0032] The feature map is refined through a multi-scale feature learning module, which can also be called an attention mechanism module, where the feature map is combined with the upsampled high-level features. After that, the fused features are upsampled accordingly and contribute to the refinement of the low-level features. Finally, a transposed convolutional layer is applied to upsample the output of the final decoder to match the spatial resolution of the input image, and then the attention mechanism is introduced, and then it is input into the final convolutional layer to extract semantic information, as Figure 4 shown, the feature 16x with the largest scale output by the pyramid module is feature-enhanced through the first attention mechanism module, and the enhanced result is input into the first decoder for decoding to obtain the first decoded feature; the feature 8x output by the pyramid module is feature-enhanced through the second attention mechanism module, and the enhanced result and the first decoded feature are merged and then input into the second decoder for decoding to obtain the second decoded feature; the feature 4x output by the pyramid module is feature-enhanced through the third attention mechanism module, and the enhanced result and the second decoded feature are merged and then input into the third decoder for decoding to obtain the third decoded feature; the feature 2x output by the pyramid module is feature-enhanced through the fourth attention mechanism module, and the enhanced result and the third decoded feature are merged and then input into the fourth decoder for decoding to obtain the fourth decoded feature. By adding a bidirectional attention map to the fourth decoded feature and fusing it through a 1×1 convolutional layer, the output of the semantic segmentation module is obtained.

[0033] Global context information is crucial for semantic information in complex scenarios, and local context information also plays a key role in preserving rich spatial details. The multi-scale feature learning module focuses on both global and local context information. One branch of the attention mechanism module uses the dot-product attention mechanism. First, the feature X is convolved using a one-dimensional convolution to obtain the matrix Q. Each column of the transformed result is a one-dimensional sequence of feature channels, and each row is the values of different channels at the same position in the feature map. Additionally, the K and V matrices can be obtained in the same way, as shown in the following formula. Calculate the similarity between each row vector and the corresponding elements of other row vectors. By applying the softmax function to each row of the similarity matrix, the normalized similarity weights of each element in the feature map relative to all other elements can be generated. The similarity weights are multiplied by the matrix V to generate the weight matrix of self-attention:

[0034] ; ; ; ; Among them, is the input of the attention mechanism module, C is the number of channels, is the weight matrix of self-attention, and the superscript T represents the transpose operation.

[0035] The first initial feature is successively subjected to convolution, batch normalization, and ReLU function activation operations through the Conv-BN-ReLU branch: The output of the dot-product attention branch and the Conv-BN-ReLU branch are fused to obtain the output of the attention mechanism module.

[0036] Feature matching module.

[0037] After integrating and transforming the second initial features of different scales, a four-dimensional cost volume (height × width × disparity range × feature size) is formed. When constructing the matching cost cube, the minimum disparity is introduced to adapt to different terrain regions and improve the accuracy and efficiency of deep learning dense matching. To obtain more context information, the present invention uses a stacked hourglass (encoding-decoding) structure, which consists of multiple repeated top-down / bottom-up processes with intermediate layer supervision. The network structure includes 4 3×3×3 convolutional layers and three stacked hourglass networks, which are composed of multiple repeated top-down / bottom-up processes with intermediate layer supervision. Attention information is introduced into each stacked hourglass network, such as Figure 7As shown, a three-dimensional convolution operation is performed on the cost volume through a convolution module; a first disparity map is obtained; the first disparity map is gradually encoded and decoded through a first encoder-decoder module, a second encoder-decoder module, and a third encoder-decoder module. During the decoding process, dense matching is achieved by introducing the bidirectional attention map output by the bidirectional guidance attention module, and a dense matching feature map is obtained; the dense matching feature map is interpolated through a bilinear interpolation layer, and the interpolation result is regressed through a regression layer to output an initial disparity map.

[0038] Bidirectional attention module.

[0039] The bidirectional attention module is divided into a semantic fusion part and a disparity fusion part. The input information is the left image, the initial disparity map, and the initial segmentation map. Except for the last convolutional layer, the other parts of the two parts are the same, as Figure 8 shown. The output dimension of the last convolutional layer of the semantic fusion part is class, where class represents the ground object category of the dataset, and the semantic residual result of the fused information is obtained. Then, it is added to the initial segmentation map to obtain a refined segmentation map. The disparity fusion part is consistent with the semantic fusion part. The refined score map is calculated by adding the fusion result and the initial score map, and then the refined segmentation map is obtained through the Argmax operation. To reduce the computational load, the output dimension of the last convolutional layer of the disparity fusion part is 1.

[0040] S2: Image data is collected to construct a dataset, and the improved BGA-Net network is trained using the dataset to obtain a semantic segmentation-dense matching model.

[0041] The present invention uses the TensorFlow framework and uses the ADAM optimizer (β1 = 0.9, β2 = 0.999) to train the model in an end-to-end manner. Before training, the input images are normalized, and the pixel intensity level is from -1 to 1. The batch size is set to 1, and the disparity is 96. The total number of epochs is set to 40. The initial learning rate is set to 0.001, and as the training progresses, every 10 epochs, the learning rate drops to half. Note that this training process is executed on a Tesla V100s GPU. The present invention adopts the open-source US3D dataset, which has a total of 4292 pairs of images. These images are WorldView-3 satellite images. 4000 pairs of images from the Jacksonville and Omaha regions are used for validation and training, and the rest are used for testing.

[0042] For the US3D dataset, under 40 epochs of training, for the dense matching accuracy metrics, the EPE is 1.38 and the D1 is 8.23%, which are better than 1.62 and 11.29% of BGANet. The semantic segmentation accuracy metric is that the F1 score is 80%.

[0043] The present invention can achieve high-precision metrics for tasks with fewer training epochs (compared to other networks), improving task efficiency. The semantic segmentation accuracy metric needs to be optimized for satellite images, and the attention mechanism used needs to be further optimized to improve task efficiency.

[0044] S3: Input the first image and the second image to be processed into the semantic segmentation-dense matching model. Perform preliminary feature extraction on the first image and the second image through the unified backbone module to obtain the first initial features of different scales and the second initial features of different scales. Perform context feature extraction and weighting on the first initial features of different scales through the bidirectional guiding attention module to obtain the bidirectional attention map. Perform attention weighting operations on the first initial features of different scales through the attention mechanism modules of several convolutional branches in the improved semantic segmentation module to obtain the semantic enhancement feature maps of different scales. Decode the semantic enhancement feature maps of different scales through the decoder module to obtain the decoded feature maps of different scales. Wherein, the input of each decoder module further includes the output from the decoder module of the convolutional branch of the previous scale. Add the bidirectional attention map to the decoded feature map of the smallest scale to obtain the initial segmentation map. Integrate the second initial features of different scales into the cost volume through the feature matching module, and perform dense matching on the cost volume and the bidirectional attention map to obtain the initial disparity map. Perform cross-fusion on the initial segmentation map and the initial disparity map through the bidirectional attention module to obtain the final segmentation map and the final disparity map.

[0045] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to the memory, storage, database, or other media used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0046] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present invention.

Claims

1. A method for semantic segmentation and matching of images, characterized in that, Including: Replacing the semantic segmentation module in the original BGA-Net network with an improved semantic segmentation module, where the improved semantic segmentation module includes a number of convolutional branches, and each convolutional branch includes an attention mechanism module and a decoder module connected in sequence; the output end of the decoder in each convolutional branch is connected to the input end of the decoder module in the next convolutional branch in the direction of decreasing scale; constructing an improved BGA-Net network including the unified backbone module, bidirectional guiding attention module, improved semantic segmentation module, feature matching module, and bidirectional attention module in the original BGA-Net network; Collecting image data to construct a dataset, and using the dataset to train the improved BGA-Net network to obtain a semantic segmentation-dense matching model; Inputting the first image and the second image to be processed into the semantic segmentation-dense matching model, and performing preliminary feature extraction on the first image and the second image through the unified backbone module to obtain the first initial features of different scales and the second initial features of different scales; Performing context feature extraction and weighting on the first initial features of different scales through the bidirectional guiding attention module to obtain a bidirectional attention map; performing attention weighting operations on the first initial features of different scales through the attention mechanism modules of a number of convolutional branches in the improved semantic segmentation module to obtain semantic enhanced feature maps of different scales; decoding the semantic enhanced feature maps of different scales through the decoder module to obtain decoded feature maps of different scales; where the input of each decoder module also includes the output from the decoder module of the convolutional branch at the previous scale; obtaining an initial segmentation map by adding the bidirectional attention map to the decoded feature map at the smallest scale; integrating the second initial features of different scales into a cost volume through the feature matching module, and performing dense matching on the cost volume and the bidirectional attention map to obtain an initial disparity map; Performing cross-fusion on the initial segmentation map and the initial disparity map through the bidirectional attention module to obtain a final segmentation map and a final disparity map.

2. The semantic segmentation and matching method of an image according to claim 1, wherein The attention mechanism module includes a parallel dot product attention branch and a Conv-BN-ReLU branch; The steps of the attention mechanism module performing attention weighting operations specifically include: Performing feature transformation on the input of the attention mechanism module through the dot product attention branch to obtain matrices Q, K, and V, calculating the normalized similarity weights of each element in the input of the attention mechanism module relative to all other elements using the softmax function, and multiplying the similarity weights by matrix V to generate a weight matrix for self-attention: Among them, is the input of the attention mechanism module, C is the number of channels, is the weight matrix of self-attention, and the superscript T represents the transpose operation; Performing convolution, batch normalization, and ReLU function activation operations on the first initial features in sequence through the Conv-BN-ReLU branch: Fusing the output of the dot product attention branch and the Conv-BN-ReLU branch to obtain the output of the attention mechanism module.

3. The semantic segmentation and matching method of an image according to claim 1, wherein A feature pyramid module is provided between the unified backbone module and the improved semantic segmentation module, and the feature pyramid module enhances the semantic information in the first initial features through convolutional operations of different scales and deconvolutional operations of different scales.

4. The semantic segmentation and matching method of an image according to claim 1, wherein The feature matching module is an improved feature matching module, which includes a convolutional module, a first encoding-decoding module, a second encoding-decoding module, a third encoding-decoding module, a bilinear interpolation layer, and a regression layer connected in sequence; The specific steps for the improved feature matching module to output the initial disparity map include: Performing a three-dimensional convolution operation on the cost volume through the convolutional module to obtain a first disparity map; Performing encoding and decoding operations on the first disparity map step by step through the first encoding-decoding module, the second encoding-decoding module, and the third encoding-decoding module. During the decoding process, dense matching is achieved by introducing the bidirectional attention map output by the bidirectional guiding attention module to obtain a dense matching feature map; Interpolating the dense matching feature map through the bilinear interpolation layer and regressing the interpolation result through the regression layer to output the initial disparity map.

5. The semantic segmentation and matching method of an image according to claim 4, characterized in that, The dense matching achieved by introducing the bidirectional attention map output by the bidirectional guiding attention module is realized by fusing the bidirectional attention map with the output of the encoding layer in the encoding-decoding module and then inputting it into the decoding layer in the encoding-decoding module for decoding.

6. The semantic segmentation and matching method of an image according to claim 1, characterized in that, The specific steps for the bidirectional guiding attention module to perform context feature extraction and weighting include: After transforming the feature map with the shape of H×W×F input to the bidirectional guiding attention module, applying the softmax function to obtain the weight distribution; Transforming the feature map with the shape of H×W×F into F×H×W, weighting the feature map transformed into F×H×W with the weight distribution output by the softmax function; and performing a convolution operation on the weighted feature map; Fusing the result of the convolution operation with the feature map input to the bidirectional guiding attention module to output the bidirectional attention map.

Citation Information

Patent Citations

  • Lightweight multi-scale feature fusion real-time image semantic segmentation method and system

    CN114445430A

  • Power transmission line segmentation method based on binocular image

    CN115953698A

  • Remote sensing image semantic segmentation method of asymmetric double-branch coding network based on clustering mutual contrast loss

    CN118864857A

  • Point cloud semantic segmentation method and model based on density perception and feature enhancement

    CN119478409A

  • Semantic segmentation-based unmanned aerial vehicle image georeferencing method, and related device

    WO2024221946A1

Cited By

  • Deep learning-based pathological image full-automatic segmentation and classification method and system

    CN120635897A