Optical remote sensing image salient target detection method based on double-branch progressive interaction and multi-direction feature enhancement
By using a dual-branch progressive interactive encoder and a multi-directional feature enhancement module in the significant object detection of optical remote sensing images, the problem of insufficient feature extraction and fusion in the prior art is solved, and higher detection accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202510209705.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-25
AI Technical Summary
The existing optical remote sensing image significant object detection method is inadequate in feature extraction and fusion when processing complex backgrounds and multi-scale objects, resulting in a decrease in detection accuracy.
A method for detecting significant object of optical remote sensing images based on dual-branch progressive interaction and multi-direction feature enhancement is proposed. Through a dual-branch progressive interaction encoder combined with CNN and Transformer branches, cross-branch interaction and dynamic weight adjustment are performed; and through semantic correlation enhancement module and multi-scale adaptive fusion module, the semantic correlation and multi-scale fusion capability of features are improved.
It significantly improves the feature extraction ability of various targets in complex optical remote sensing images, improves detection accuracy and adaptability, and can more effectively deal with complex backgrounds and multi-scale targets.
Smart Images

Figure CN119992068A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method for detecting salient targets in optical remote sensing images based on dual-branch progressive interaction and multi-directional feature enhancement, and belongs to the technical field of optical remote sensing image processing and analysis. Background Art
[0002] Optical remote sensing image salient object detection, hereinafter referred to as ORSI-SOD, aims to accurately detect the most salient objects or areas in optical remote sensing images, and can provide an important preprocessing method for other optical remote sensing image tasks. However, since optical remote sensing images are taken by remote sensing satellites or cameras on aircraft at a bird's-eye view, and the imaging is affected by the climate, ORSI-SOD faces challenges such as blurred target boundaries, complex backgrounds, and large differences in target scales. Therefore, the detected salient objects are often not accurate or complete enough.
[0003] In 2019, Li et al. published the first paper on remote sensing saliency detection using deep learning, constructed the first open source dataset ORSSD, and proposed LVNet. In this model, the L-shaped two-stream pyramid module is responsible for multi-scale feature extraction, and the V-shaped nested encoder-decoder module highlights salient objects and suppresses the background. Zhang et al. proposed an attention-intensive transfer for more efficient transfer, expanded the first dataset and named it EORSSD, introduced edge supervision, and proposed DAFNet. This model implements shallow-to-deep attention information transfer, simultaneously infers saliency maps and edge maps, and enhances edge information. Tu et al. published a more comprehensive RSI dataset, ORSI-4199. They creatively designed a multi-scale joint boundary and region model. Due to the development of deep learning and the accessibility of a wide range of datasets such as ORSSD, EORSSD, and ORSI-4199, significant progress has been made in the field of ORSI-SOD. Cong et al. constructed a relational reasoning model for ORSI-SOD using a parallel multi-scale attention mechanism. Zhou et al. used edge information to generate target location clues at each decoding step. In addition, some networks draw inspiration from Transformer and integrate long-range self-attention mechanisms into ORSI-SOD, thereby enhancing the representation of global information. Gao et al. designed an adaptive spatial tokenization Transformer encoder to extract global-local features. Bao et al. took a different route and proposed a dual-branch architecture combining CNN and Transformer, which facilitated interactive guided aggregation of multi-level global and local features. In addition, some studies aim to build lightweight models such as FSMINet, CorrNet, and SeaNet.
[0004] Although deep learning technology has made significant progress in the field of salient object detection in remote sensing images in recent years, existing methods still have the following problems when dealing with complex backgrounds and multi-scale objects:
[0005] Insufficient feature extraction and fusion: Traditional CNNs have difficulty capturing the global features of the target, while Transformers are weak in extracting local details. In addition, existing methods often fail to fully utilize global context information when extracting features and are not adaptable enough to complex backgrounds. At the same time, existing methods fail to effectively fuse multi-scale features, resulting in reduced detection accuracy.
[0006] Multi-scale target detection problem: The targets in optical remote sensing images may be at different scales. Traditional convolutional neural networks have limitations in capturing multi-scale information and are difficult to effectively process targets of different scales. Summary of the invention
[0007] The present invention aims to solve the above problems and proposes a method for detecting salient objects in optical remote sensing images based on dual-branch progressive interaction and multi-directional feature enhancement, which specifically comprises the following steps:
[0008] S1: The original dataset of salient object detection in optical remote sensing images is cut into non-overlapping sub-images according to the specified size, and divided into training set, validation set, and test set; data enhancement is performed before input into the network; the specific process is as follows:
[0009] S1.1: Crop the original images of the dataset into 224×224 images to adapt to the network input size;
[0010] S1.2: Use stratified sampling method to divide into training set, validation set and test set according to different proportions;
[0011] S1.3: Perform random horizontal flipping, random vertical flipping, random rotation and random cropping operations on the training set images. Bilinear interpolation is used for resampling during random rotation. During random cropping, the overlap rate between the center area of the cropped image and the original target area is ensured to be no less than 50%;
[0012] S1.4: Normalize the image and map the pixel values to the interval [0, 1];
[0013] S2: Input the input image into the dual-branch progressive interaction encoder for feature extraction; insert large windmill convolution into the backbone network, dynamically adjust the convolution kernel, and increase the receptive field; use the progressive interaction mechanism to perform cross-branch interaction and dynamic weight adjustment on the features of the dual branches;
[0014] S3: The features extracted by the dual-branch progressive interaction encoder are input into the semantic association enhancement module to calculate the similarity between local and global features, generate semantic association weights, and enhance the semantic consistency of salient targets; add directional convolution units to enhance the perception of the direction of salient targets; introduce context-aware attention mechanism to highlight the features of salient targets;
[0015] S4: Input the features output by the semantic association enhancement module into the multi-scale adaptive fusion module to dynamically adjust the fusion ratio of features at different scales; introduce deformable pooling operations to adaptively adjust the pooling kernel to avoid the loss of key features; add directional convolution units to refine the target boundary;
[0016] S5: Input the features output by the multi-scale adaptive fusion module into the decoder, and fuse the features of the encoder and decoder through cross-layer connections;
[0017] S6: Train the model on the training set and save the best model parameters; input the test set images into the model to obtain the salient object detection results.
[0018] Furthermore, the dual-branch progressive interaction encoder is specifically as follows: inserting a windmill convolution into the CNN branch backbone network, and using the convolution layer to extract the shallow feature stage; the windmill convolution adopts an asymmetric filling strategy to capture the distribution characteristics of the target; by dynamically adjusting the shape and size of the convolution kernel, the receptive field is significantly increased; the Transformer branch is used to extract global context features; cross-branch interaction nodes are set in the CNN branch and the Transformer branch for feature interaction; the CNN branch contains five layers, and the Transformer branch contains three layers. The feature interaction starts from the third layer of the CNN branch and interacts layer by layer. First, the progressive interaction mechanism is used to perform cross-branch interaction and dynamic weight adjustment on the features of the two branches; secondly, the fourth layer of the CNN branch interacts with the second layer of the Transformer branch, and the first operation needs to be repeated at this time; finally, the fifth layer of the CNN branch interacts with the third layer of the Transformer branch, and no dynamic weight adjustment is required at this time, and the features are directly input into the semantic association enhancement module.
[0019] Furthermore, the progressive interaction mechanism is specifically as follows: a cross-attention mechanism is introduced in the cross-branch interaction node, attention weights are generated by calculating the similarity of the features of the CNN branch and the Transformer branch, and the feature fusion ratio is dynamically adjusted; a dynamic branch weight adjustment mechanism is introduced, the global features of the image are analyzed through the global feature evaluation module, a weight adjustment signal is generated, and the weights of the CNN branch and the Transformer branch are dynamically adjusted.
[0020] Furthermore, the semantic association enhancement module is specifically as follows: adding a directional convolution unit, extracting directional features through its multi-directional parallel convolution operation, and enhancing feature expression capabilities; enhancing the perception of target direction by aggregating convolution features in different directions; introducing a context-aware attention mechanism, incorporating contextual semantic information while considering the spatial relationship between pixels, taking 5×5 neighborhood pixels centered on the current pixel, calculating the sum of the absolute values of the differences, and obtaining semantic association weights through dot product operations and the Softmax function, focusing on significant target features and suppressing background interference.
[0021] Furthermore, the multi-scale adaptive fusion module is specifically as follows: according to the feature map resolution and the target average activation strength, the fusion ratio of features of different scales is dynamically adjusted, and for images containing targets of multiple scales, features of different scales are combined through a weighted fusion mechanism; in the feature fusion process, a deformable pooling operation is introduced, and the shape and position of the pooling area are adaptively adjusted by learning the offset to avoid the loss of key features and provide a more accurate feature representation; a directional convolution unit is added to extract target boundary features through convolution kernels in different directions, refine the target boundary, and provide the decoder with features with refined boundaries and high discrimination.
[0022] Compared with the prior art, the present invention has the following beneficial results:
[0023] 1. The present invention proposes a dual-branch progressive interaction encoder structure based on dual-branch progressive interaction and multi-directional feature enhancement, combining the windmill convolution of the CNN branch and the self-attention mechanism of the Transformer branch, which not only utilizes the powerful capture ability of CNN for local detail features, but also gives play to the Transformer's effective perception advantage of global context information. Through cross-branch interaction and dynamic weight adjustment, it can adaptively balance the fusion of local and global features in different scenarios, realize multi-scale feature capture from tiny target details to large-area scene patterns, and significantly improve the feature extraction capability of various targets in complex optical remote sensing images.
[0024] 2. The present invention constructs a semantic association enhancement module to effectively improve the semantic association between local and global features. This module calculates the similarity between local and global features through attention mechanism and contrastive learning, generates semantic association weights, focuses on salient target features, and suppresses background interference. At the same time, a context-aware attention mechanism is added, and the sum of the absolute values of the difference between neighboring pixels is calculated with the current pixel as the center. The semantic association weights are obtained through dot product operation and Softmax function to further highlight the salient target features. This series of operations ensures the semantic consistency between features of different scales and levels, reduces the semantic gap between features, and improves the model's understanding and detection accuracy of salient targets.
[0025] 3. The multi-scale adaptive fusion module designed in the present invention realizes the dynamic fusion and processing of features of different scales. The module uses the feature map resolution and the average activation intensity of the target to judge the target scale, dynamically adjusts the fusion ratio of features of different scales, and combines features of different scales through a weighted fusion mechanism for images containing targets of multiple scales, giving full play to the structural representation ability of large-scale features for large target areas and the ability of small-scale features to depict the details of small targets. In addition, the integrated deformable pooling operation adaptively adjusts the shape and position of the pooling area by learning the offset, avoids the loss of key features, provides more accurate feature representation, enhances the model's detection ability for multi-scale and irregularly shaped targets, and improves adaptability and detection accuracy in complex scenes.
[0026] 4. The present invention adopts a combination of multiple training strategies to effectively improve the performance and generalization ability of the model. End-to-end training combines cross entropy loss, Dice loss and IoU loss to construct a total loss function, comprehensively considers multiple aspects of target detection, and optimizes model training. The AdamW optimizer is selected and combined with the cosine annealing learning rate strategy to achieve a balance between rapid convergence and fine optimization during the training process. At the same time, adversarial training enhances the robustness of the model, multi-task collaborative training improves the accuracy and completeness of detection, and self-supervised learning uses a large amount of unlabeled data to learn general image features in the pre-training stage, so that the model can better adapt to different optical remote sensing image scenes, enhance the generalization ability of the model, and reduce dependence on specific data sets. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:
[0028] Figure 1 A flowchart of the steps of a method for detecting salient objects in optical remote sensing images based on dual-branch progressive interaction and multi-directional feature enhancement provided by the present invention;
[0029] Figure 2 A network structure diagram of a method for detecting salient objects in optical remote sensing images based on dual-branch progressive interaction and multi-directional feature enhancement provided by the present invention;
[0030] Figure 3 This is the structural diagram of the semantic association enhancement module;
[0031] Figure 4 This is the structural diagram of the multi-scale adaptive fusion module. DETAILED DESCRIPTION
[0032] In order to make those skilled in the art better understand the present invention, the technical solution of the present invention is described in detail below in conjunction with specific embodiments. The specific embodiments described here are only used to explain the present invention and are not used to limit the present invention.
[0033] like Figure 1 As shown, it is a flowchart of a method for detecting salient targets in optical remote sensing images based on dual-branch progressive interaction and multi-directional feature enhancement according to an embodiment of the present invention, including steps S1-S6:
[0034] S1, the original dataset of salient object detection in optical remote sensing images is cut into non-overlapping sub-images according to the specified size, and divided into training set, validation set, and test set. Data enhancement is performed before input into the network.
[0035] In the example of the present invention, step S1 specifically includes steps S11-S13:
[0036] S11, the original images of the ORSSD, EORSSD and ORSI-4199 datasets are cropped into non-overlapping sub-images of the specified size, with a size of 224×224 pixels to adapt to the network input size;
[0037] S12: Divide the cropped images into training sets, validation sets, and test sets according to different ratios. For the ORSSD dataset, the training sets, validation sets, and test sets are divided according to a ratio of 600 / 100 / 100; for the EORSSD dataset, the training sets, validation sets, and test sets are divided according to a ratio of 1500 / 250 / 250; for the ORSI-4199 dataset, the training sets, validation sets, and test sets are divided according to a ratio of 3000 / 600 / 599;
[0038] S13, before inputting into the network, these data sets are randomly flipped horizontally, randomly flipped vertically, and randomly rotated, that is, the angle range is [-30°, 30°] and randomly cropped, and the overlap rate between the center area of the cropped image and the original target area is guaranteed to be no less than 50% during random cropping. Finally, the image is normalized and the pixel value is mapped to the interval [0, 1];
[0039] S2, input the preprocessed image into the dual-branch progressive interaction and multi-directional feature enhancement network, and use the progressive interaction mechanism to perform cross-branch interaction and dynamic weight adjustment on the features of the two branches;
[0040] like Figure 2 As shown, a dual-branch progressive interaction and multi-directional feature enhancement network structure diagram of an embodiment of the present invention is provided. The backbone network is divided into a CNN branch and a Transformer branch, and the specific steps of step S2 include S21-S24:
[0041] S21, CNN branch feature extraction: The first layer is at the beginning of ResNet50 (after conv1 layer), and a specially designed windmill convolution is performed. First, two consecutive windmill convolution operations are performed. The convolution kernel size of each windmill convolution layer is 5×5, the step size is 1, the padding is 2, and the dilation rate is 1. The windmill convolution operation can be expressed as:
[0042]
[0043] Among them, w m,n is the original convolution kernel weight, is the weight of the convolution kernel after rotating 45 degrees. In this way, the process of extracting features from multiple directions by simulating the windmill convolution is simulated. After two windmill convolutions, the image resolution is reduced from 224×224 to 112×112, and the number of channels is increased from 3 to 64. Windmill convolution can better capture complex features such as edges and textures, and can more effectively extract unique information of the target than ordinary convolution. The second layer is processed by two ordinary convolutions with a convolution kernel size of 3. These two convolution layers act in sequence to reduce the feature map from 112×112 to 56×56, and the number of channels is increased from 64 to 128. The features of the image are further extracted, and the features extracted by the previous windmill convolution are refined and integrated to provide more representative feature maps for subsequent layers. The third layer is in the middle layer of ResNet50, that is, conv3_x position, and two windmill convolution operations are performed. The convolution kernel size, stride, padding and dilation rate are the same as those of the windmill convolution in the first layer. The feature map is reduced from 56×56 to 28×28, and the number of channels is increased from 128 to 256. For the feature map processed by the first two layers of convolution, the windmill convolution further captures the characteristics of complex-shaped targets and enhances the model's perception of target details. The fourth layer uses two ordinary convolutions with a convolution kernel size of 3 to reduce the feature map from 28×28 to 14×14, and the number of channels is increased from 256 to 512, and further integrates the features to deeply integrate the various features extracted previously. The fifth layer uses two ordinary convolutions with a convolution kernel size of 3 to output a high-dimensional, high-resolution feature map, providing rich information for subsequent cross-branch interactions and feature fusion;
[0044] S22, Transformer branch feature extraction: Swin Transformer is used as the backbone network of the Transformer branch, and the input 224×224 image is processed as follows: The first layer uses a convolution operation with a convolution kernel size of 7, a padding of 3, and a stride of 4 to increase the dimension of the feature map. Assuming the convolution kernel is Kconv1, the convolution operation can be expressed as:
[0045] F conv1 =Kconv1 ×I
[0046] After this convolution, the feature map dimension size becomes 64, and the feature map scale is 56×56. After that, the feature map is layer normalized to stabilize the distribution of features. The layer normalization formula is:
[0047]
[0048] Among them, μ and σ 2 are the mean and variance of the feature map in the channel dimension, ε is a small constant to prevent the denominator from being 0, and γ and β are learnable parameters.
[0049] Perform multi-head self-attention operations. In the multi-head self-attention mechanism, the formula for calculating the attention score is:
[0050]
[0051] Where Q = W Q F LN , K=W K F LN , V=W V F LN , W Q , W K , WV is the linear transformation matrix, F LN is the feature map after layer normalization, d k is the dimension of the key. Calculate the correlation between the feature at each position and the features at all other positions to obtain global context information. Add the feature map obtained after the operation to the feature map without multi-head self-attention. This residual connection method helps to retain important information in the original features and avoid losing key details during the self-attention operation. That is:
[0052] F add =Attention(Q, K, V)+F LN
[0053] Then perform layer normalization and multi-layer perceptron operations in sequence. The multi-layer perceptron can be expressed as:
[0054] MLP(F add )=W 2 ×ReLU(W 1 ×F add +b 1 )+b 2 Among them, W 1 , W 2 is the weight matrix, b 1 , b 2 It is a bias term, which further transforms and integrates the features to enhance the expressiveness of the features.
[0055] The second layer uses convolution with a kernel size of 3 and a padding of 1 to change the feature map from 56×56 to 28×28, and the number of channels increases from 64 to 128. In this process, the convolution operation not only changes the size and number of channels of the feature map, but also further extracts feature information at different scales.
[0056] The third layer uses convolution with a kernel size of 3 and padding of 1 to convert the feature map from 28×28 to 14×14, and the number of channels increases from 128 to 256. The feature map is further processed to obtain a more refined feature representation. After that, a multi-head self-attention operation is performed, and the self-attention mechanism is used again to capture global information, strengthen the correlation between features, and improve the model's ability to understand complex scenes;
[0057] S23, cross-branch interaction node: set cross-branch interaction nodes in the convolutional layers of the CNN branch, i.e. conv3_x, conv4_x, and the first and second layers of the Transformer branch, and perform feature interaction in each layer. Introduce the cross-attention mechanism in the cross-branch interaction node, calculate the similarity of the features of the CNN branch and the Transformer branch, generate attention weights, and dynamically adjust the feature fusion ratio. The specific calculation process is: linearly transform the feature maps from the two branches to obtain Q, K, and V vectors, calculate the dot product of Q and K and perform a Softmax operation to obtain the attention weight, and then perform a weighted summation of the attention weight and V to achieve dynamic adjustment of the feature fusion ratio. In this way, the features of the two branches complement each other and give full play to their respective advantages;
[0058] S24, dynamic weight adjustment: introduce a dynamic branch weight adjustment mechanism, analyze the global features of the image through the global feature evaluation module, generate weight adjustment signals, and dynamically adjust the weights of the CNN branch and the Transformer branch. The global feature evaluation module consists of 3 convolutional layers, ReLU activation and 2 fully connected layers, which are used to extract scene complexity indicators and target scale information. When the grayscale co-occurrence matrix calculates the texture complexity, for the grayscale image I, the grayscale co-occurrence matrix G(i, j, d, θ) represents the frequency of occurrence of pixel pairs with grayscale values i and j when the distance is d and the direction is θ. The texture complexity T can be calculated by the following formula:
[0059]
[0060] Where L is the number of gray levels. A weight adjustment signal is generated based on the evaluation results and transmitted to the weight control nodes of the CNN branch and the Transformer branch respectively to dynamically adjust the branch weights. When detecting images containing objects of various scales, the weight of the CNN branch that is more favorable for small target detection is increased according to the target scale information, or the weight of the Transformer branch that is more effective for large targets and global scene perception is enhanced, so that the model can more efficiently utilize the features of the two branches in different scenarios.
[0061] S3, the features extracted by the dual-branch progressive interaction encoder are input into the semantic association enhancement module.
[0062] like Figure 3 As shown, a structure diagram of a semantic association enhancement module according to an embodiment of the present invention is provided, and the specific steps of step S3 include S31-S34.
[0063] S31, directional convolution to enhance feature expression: add directional convolution units to extract and enhance directional features of dual-branch salient target features. For directional convolution, let the input feature map be F, and the convolution kernel in the kth direction be K k , then the output of the directional convolution for:
[0064]
[0065] The convolution operation is performed on the feature map through these convolution kernels in different directions to capture the direction-sensitive features of the salient target and enhance the feature expression ability; the 8-directional feature maps of the dual branches are spliced and the number of channels is compressed through 1×1 convolution;
[0066] S32, generate Q, K, V vectors, and generate attention vectors based on the fused features;
[0067] S33, context-aware attention mechanism: Introduce the context-aware attention mechanism, take 5×5 neighboring pixels with the current pixel as the center, calculate the sum of the absolute values of the differences, and obtain the semantic association weight through dot product operation and Softmax function. Let the current pixel value be, the neighboring pixel value be p, and the neighboring pixel value be p ij , i, j∈[-2, 2], the sum of the absolute values of the differences D is:
[0068]
[0069] Assume that the global feature vector is G, and the semantic association weight β is calculated as follows:
[0070] β=Softmax(D×G)
[0071] In this way, salient target features are focused, background interference is suppressed, and the semantic correlation between local and global features is enhanced.
[0072] S4, inputs the features output by the semantic association enhancement module into the multi-scale adaptive fusion module for multi-scale feature fusion.
[0073] like Figure 4 As shown, a multi-scale adaptive fusion module structure diagram of an embodiment of the present invention is provided, and the specific steps of step S4 include S41-S45.
[0074] S41, target scale judgment: using feature map resolution: width and height can be directly obtained; target average activation intensity: calculate the average pixel value of the target area, and the target area is obtained by segmentation with a threshold of 0.5 to judge the target scale. Assume that the feature map resolution is (H, W), the average activation intensity of the target area is A, and define the scale judgment function S(H, W, A):
[0075]
[0076] Among them, T 1 , T 2 , T 3 , T 4 It is a threshold set according to the experiment. By analyzing the resolution of the feature map, the approximate scale range of the target can be preliminarily determined; combined with the average activation intensity of the target, the judgment of the target scale can be further refined. For areas with low resolution but high average activation intensity, they may correspond to small-scale but significant targets; while areas with high resolution and moderate average activation intensity may contain large-scale targets. In this way, it provides a basis for subsequent scale feature fusion;
[0077] S42, scale feature fusion: dynamically adjust the fusion ratio of different scale features. For images containing objects of multiple scales, combine different scale features through a weighted fusion mechanism. Let the large-scale feature map be F and the small-scale feature map be F large , according to the target scale judgment result F samll , the fused feature map F fusion for:
[0078]
[0079] Among them, ω 1 and ω 2 is adaptively generated through the attention mechanism, and ω 1 +ω 2=1. Different weights are assigned to feature maps of different scales according to the target scale judgment results. For large-scale targets, the weight of large-scale feature maps is increased to highlight their overall structural information; for small-scale targets, the weight of small-scale feature maps is increased to retain their detailed features. Through this weighted fusion method, the advantages of features of different scales are fully utilized to improve the detection accuracy of multi-scale targets;
[0080] S43, deformable pooling operation: In the feature fusion process, the deformable pooling operation is integrated. Deformable pooling adaptively adjusts the shape and position of the pooling area by learning the offset to avoid the loss of key features and provide more accurate feature representation. The sampling point offset of the deformable pooling module is predicted by a small convolutional network containing 2 convolutional layers and ReLU activation. Let the predicted offsets be Δx and Δy. For the input feature map F, the output F of the deformable pooling is def for:
[0081] F def (i,j)=F(i+Δx(i,j),j+Δy(i,j))
[0082] According to the predicted offset, when performing pooling operations on feature maps of different scales, the pooling area can be adaptively adjusted to better adapt to the shape and size changes of the target and effectively retain the key features of the target. Deformable pooling can dynamically adjust the pooling area according to the boundary shape of the image to avoid losing edge details;
[0083] S44, directional convolution to refine boundaries: add directional convolution units, and use these directional convolution kernels to perform convolution operations on the fused feature maps to further highlight the target boundary features and refine the target boundaries. The directional convolution kernel detects and strengthens the target boundaries from multiple directions, providing the decoder with features with refined boundaries and high discrimination, which helps to improve the accuracy and clarity of the target boundaries; add residual connections after directional convolution to avoid feature information loss;
[0084] S45, output fused feature map: output the feature map after multi-scale adaptive fusion, deformable pooling and directional convolution for subsequent decoder and output layer processing. These fused feature maps contain multi-scale, directional and boundary information, providing strong support for the decoder to accurately restore the target image.
[0085] S5, input the features output by the multi-scale adaptive fusion module into the decoder, fuse the features of the encoder and the decoder through cross-layer connection, and generate a salient object detection result. Step S5 specifically includes steps S51-S54.
[0086] S51, decoder upsampling and convolution operation: The decoder is composed of 4 layers of upsampling and deconvolution operations to gradually restore the image resolution. After each layer of upsampling, a directional convolution unit is added to gradually restore the low-resolution feature map to a high-resolution image through upsampling and deconvolution operations; the directional convolution unit further refines the target boundaries and features, so that the restored image can more accurately present the details of the target;
[0087] S52, cross-layer connection feature fusion: set up cross-layer connections in the decoder to fuse the features of different levels of the encoder with the features of the corresponding levels of the decoder. Through cross-layer connections, the features extracted from different levels in the encoder and the features of the corresponding levels of the decoder are concatenated in the channel dimension and fused through a 1×1 convolutional layer. This cross-layer connection method can make full use of information at different levels, so that the decoder can better reconstruct the shape and structure of the target in the process of restoring the image resolution, thereby improving the accuracy and completeness of target detection;
[0088] S53, multi-task output: The output layer uses multi-task output, and uses the convolution layer to output the salient target detection results and target category information. By setting different convolution kernels and classifiers in the output layer, the fused feature map is processed, and the location, contour information and category of the target are output at the same time;
[0089] S54, output and save results: Visualize and save the salient target detection results obtained by the output layer. Use the image processing library to display the detection results in the form of images to intuitively present the detection effect of the model; at the same time, save the detection results as files for subsequent analysis and evaluation;
[0090] S6, train the model on the training set and save the best model parameters. Input the test set images into the model to obtain the salient object detection results.
[0091] In the example of the present invention, an end-to-end training method is adopted, and the total loss function is constructed by combining cross entropy loss, Dice loss and IoU loss:
[0092] L total =L CE +0.5L Dice +0.3L IoU
[0093] Select AdamW optimizer and set the initial learning rate to 7×10 -5, the weight decay coefficient is 0.0001. The cosine annealing learning rate strategy is used to dynamically adjust the learning rate during the training process. Construct an adversarial network. The generator consists of 5 transposed convolutional layers, i.e., convolution kernel 4×4, step size 2, padding 1, ReLU activation and the last layer Tanh activation. The discriminator consists of 5 convolutional layers, LeakyReLU activation and the last layer Sigmoid activation. The Wasserstein GAN loss is used for adversarial training to enhance the robustness of the model. Add a target classification task head, i.e., it contains 2 fully connected layers, the number of hidden layer nodes is 256 and 128, and Softmax activation; the edge detection task head, i.e., it contains 3 convolutional layers and 1 deconvolution layer, ReLU activation and the last layer Sigmoid activation. The target classification loss and edge detection loss use cross entropy loss and Dice loss respectively to achieve multi-task collaborative training and improve the accuracy and completeness of detection. Introduce self-supervised learning tasks in the pre-training stage, and use a large amount of unlabeled data to improve the feature representation ability of the model. During the training process, the validation set is evaluated every 5 epochs, the performance indicators of the model are recorded, and the model with the best performance is selected for saving.
[0094] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various modifications or variations within the scope of the claims, which does not affect the essence of the present invention. The above preferred features can be used in any combination without conflicting with each other.
Claims
1. A method for detecting salient objects in optical remote sensing images based on dual-branch progressive interaction and multi-directional feature enhancement, characterized in that: The steps include: S1: The original dataset of salient object detection in optical remote sensing images is cut into non-overlapping sub-images according to the specified size, and divided into training set, validation set, and test set; data enhancement is performed before input into the network; the specific process is as follows: S1.1: Crop the original images of the dataset into 224×224 images to adapt to the network input size; S1.2: Use stratified sampling method to divide into training set, validation set and test set according to different proportions; S1.3: Perform random horizontal flipping, random vertical flipping, random rotation and random cropping operations on the training set images. Bilinear interpolation is used for resampling during random rotation. During random cropping, the overlap rate between the center area of the cropped image and the original target area is ensured to be no less than 50%; S1.4: Normalize the image and map the pixel values to the interval [0,1]; S2: Input the input image into the dual-branch progressive interaction encoder for feature extraction; insert large windmill convolution into the backbone network, dynamically adjust the convolution kernel, and increase the receptive field; use the progressive interaction mechanism to perform cross-branch interaction and dynamic weight adjustment on the features of the dual branches; S3: The features extracted by the dual-branch progressive interaction encoder are input into the semantic association enhancement module to calculate the similarity between local and global features, generate semantic association weights, and enhance the semantic consistency of salient targets; add directional convolution units to enhance the perception of the direction of salient targets; introduce context-aware attention mechanism to highlight the features of salient targets; S4: Input the features output by the semantic association enhancement module into the multi-scale adaptive fusion module to dynamically adjust the fusion ratio of features at different scales; introduce deformable pooling operations to adaptively adjust the pooling kernel to avoid the loss of key features; add directional convolution units to refine the target boundary; S5: Input the features output by the multi-scale adaptive fusion module into the decoder, and fuse the features of the encoder and decoder through cross-layer connections; S6: Train the model on the training set and save the best model parameters; input the test set images into the model to obtain the salient object detection results.
2. The optical remote sensing image salient target detection method based on dual-branch progressive interaction and multi-directional feature enhancement according to claim 1, characterized in that: Constructing the dual-branch progressive interactive encoder includes: Insert a pinwheel convolution into the CNN branch backbone network, and use the convolution layer to extract shallow feature stages; the pinwheel convolution adopts an asymmetric padding strategy to capture the distribution characteristics of the target; by dynamically adjusting the shape and size of the convolution kernel, the receptive field is significantly increased; the Transformer branch is used to extract global context features; Cross-branch interaction nodes are set in the CNN branch and the Transformer branch for feature interaction. The CNN branch contains five layers, and the Transformer branch contains three layers. The feature interaction starts from the third layer of the CNN branch and interacts layer by layer. First, the progressive interaction mechanism is used to perform cross-branch interaction and dynamic weight adjustment on the features of the two branches. Secondly, the fourth layer of the CNN branch interacts with the second layer of the Transformer branch, and the first operation needs to be repeated at this time. Finally, the fifth layer of the CNN branch interacts with the third layer of the Transformer branch. At this time, there is no need to perform dynamic weight adjustment, and the features are directly input into the semantic association enhancement module.
3. The optical remote sensing image salient object detection method based on dual-branch progressive interaction and multi-directional feature enhancement according to claim 1, characterized in that: The progressive interaction mechanism specifically includes cross-branch interaction and dynamic weight adjustment of dual-branch features, including: Introducing the cross-attention mechanism in the cross-branch interaction nodes, by calculating the similarity of the features of the CNN branch and the Transformer branch, generating attention weights and dynamically adjusting the feature fusion ratio; A dynamic branch weight adjustment mechanism is introduced. The global feature evaluation module analyzes the global features of the image, generates weight adjustment signals, and dynamically adjusts the weights of the CNN branch and the Transformer branch.
4. The method for detecting salient targets in optical remote sensing images based on dual-branch progressive interaction and multi-directional feature enhancement according to claim 1, characterized in that: Constructing the semantic association enhancement module includes: Add a directional convolution unit to extract directional features and enhance feature expression capabilities through its multi-directional parallel convolution operation; enhance the perception of target direction by aggregating convolution features in different directions; A context-aware attention mechanism is introduced to integrate contextual semantic information while considering the spatial relationship between pixels. A 5×5 neighborhood pixel is taken with the current pixel as the center, the sum of the absolute values of the difference is calculated, and the semantic association weight is obtained through dot product operation and Softmax function to focus on salient target features and suppress background interference.
5. The optical remote sensing image salient object detection method based on dual-branch progressive interaction and multi-directional feature enhancement according to claim 1, characterized in that: Constructing the multi-scale adaptive fusion module includes: According to the feature map resolution and the average activation strength of the target, the fusion ratio of features of different scales is dynamically adjusted. For images containing targets of multiple scales, features of different scales are combined through a weighted fusion mechanism. In the feature fusion process, a deformable pooling operation is introduced to adaptively adjust the shape and position of the pooling area by learning the offset to avoid the loss of key features and provide more accurate feature representation; Add a directional convolution unit to extract target boundary features through convolution kernels in different directions, refine the target boundary, and provide the decoder with features with refined boundaries and high discrimination.
Citation Information
Patent Citations
Remote sensing image semantic segmentation method based on double-branch feature fusion
CN115797931A
Image processing method and system based on double-branch multi-scale semantic segmentation network
CN116580241A
Monocular depth prediction method based on multi-path feature extraction and multi-scale feature fusion
CN116758130A
Remote sensing image segmentation method based on dual-branch multi-scale feature fusion
CN118314353A
Remote sensing image segmentation method based on channel enhancement and cross-level multi-input features
CN119380018A
Cited By
Texture multi-scale feature fusion detection method based on light convolution and deformation attention
CN120198763A
CTSP solving method and system based on convolution enhanced proxy attention mechanism
CN120354086A
Target detection network for small target and training method and detection method thereof
CN120599386A
Optical remote sensing image salient target detection method based on progressive attention enhancement
CN120894536A
Image feature enhancement method and device, electronic equipment and product
CN121147045A