Salient object detection method in optical remote sensing images based on Transformer boundary perception
By adopting a Transformer-based boundary perception method in the detection of significant targets of optical remote sensing images, problems such as incomplete detection of large targets and missed detection of small targets in the prior art are solved, and the complete detection of significant targets in optical remote sensing images are achieved and the boundary clarity is improved.
Patent Information
- Application Number
- CN202210458900.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-27
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-04-27
AI Technical Summary
Optical remote sensing images have significant target detection faces challenges such as incomplete large target detection, small target miss detection, large scale changes, many noise interferences, and complex imaging conditions. Existing convolutional neural networks are difficult to effectively utilize global clues.
Using a boundary perception method based on Transformer, multi-level features are extracted through feature encoder, combined with spatial attention mechanism and global context module, the interaction between shallow and high-level features is optimized, and boundary information is further enhanced in the boundary perception decoder, and supervised training is performed through joint loss functions.
Complete detection of significant targets in optical remote sensing images is achieved, boundary clarity is improved, the problem of unbalanced foreground background categories is alleviated, and detection performance is significantly improved.
Smart Images

Figure CN115049921B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer vision processing technology, and in particular to a method for detecting salient objects in optical remote sensing images based on Transformer boundary perception. Background Art
[0002] Salient object detection is an important topic in the field of computer vision. This task aims to detect the most visually distinctive object from a given image. In recent years, salient object detection in natural scene images has become relatively mature, but salient object detection in optical remote sensing images remains to be explored. Salient object detection in optical remote sensing images has strong practical value. This task can be applied to various fields of computer vision to play a preprocessing role, such as object segmentation, visual tracking, image retrieval, cropping, image quality assessment, etc.
[0003] Optical remote sensing images are taken by satellites and aerial sensors, while natural scene images are usually taken by handheld cameras. Therefore, the huge difference in the acquired images makes the detection of salient targets in optical remote sensing images more challenging in terms of scale change, imaging conditions, noise interference, and imaging direction. Optical remote sensing images are taken outdoors at a very high angle, so the scales of the targets they contain vary greatly, including large targets such as buildings and islands, as well as small targets such as airplanes and ships. In addition, the salient targets exist in various directions due to the bird's-eye view, while natural scene images are all in vertical directions. At the same time, the background of optical remote sensing images is complex, the noise interference is diverse, and they are easily affected by imaging conditions such as light intensity, shooting time, and shooting height. Therefore, it is difficult to obtain satisfactory results when directly using the most advanced natural scene image salient target detection method to process optical remote sensing images.
[0004] In existing research, most of them use convolutional neural networks to make great progress in salient object detection in optical remote sensing images. However, convolutional neural networks have the inherent limitation of extracting features from neighborhood pixels, so previous methods have difficulty in utilizing key global cues. Recently, Swin Transformer was proposed, which realizes entity pair interaction within a local window through multi-head self-attention and establishes long-distance dependencies across windows through a sliding window scheme. The features extracted from Transformer have more global information than those extracted from convolutional neural networks, and global context information has been shown to be the key to saliency detection.
[0005] At present, there are still some challenges to be solved in optical remote sensing image salient object detection that affect its performance:
[0006] On the one hand, there will be incomplete detection of large targets and missed detection of small targets. There are some narrow and long targets or targets that occupy a large area of the entire image in optical remote sensing images, which can easily cause incomplete detection. However, convolutional neural networks cannot establish good long-distance semantic dependencies to solve the problem of long-distance feature inconsistency. At the same time, since there are many small targets in the data set and the foreground accounts for a small proportion, there is an imbalance in the foreground and background categories. This problem will lead to a decrease in detection accuracy and missed detection of small targets.
[0007] On the other hand, many deep learning methods for salient detection in visible light images can use the edge information of salient objects to more accurately predict the boundaries of salient maps. However, due to the complexity of the foreground and background of optical remote sensing images, simply and directly fusing the boundary information of multiple layers of salient objects is not ideal for salient detection in optical remote sensing images. Summary of the invention
[0008] Purpose of the invention: The purpose of the present invention is to solve the deficiencies in the prior art and to provide a method for detecting salient objects in optical remote sensing images based on Transformer boundary perception.
[0009] Technical solution: The present invention provides a method for detecting salient targets in optical remote sensing images based on Transformer boundary perception, which is characterized by comprising the following steps:
[0010] Step S1, using a feature encoder to encode and obtain multi-level features of a remote sensing image, and marking the remote sensing image features as f1-f4;
[0011] Step S2: Use the spatial attention mechanism to guide the shallow features f1 and f2 to interactively optimize, and fuse and extract the common area feature information to obtain new shallow features F1 and F2;
[0012] Step S3, high-level features f3 and f4 are processed by the global information based on the Transformer global context module to obtain features F3 and F4 containing sufficient global context information respectively;
[0013] Step S4: In the boundary-aware decoder, the boundary information in the features F1-F3 is further enhanced, and the boundary features F′1-F′3 are obtained through boundary truth supervision. Then, the boundary features of each layer and the features of the emphasized salient areas are gradually fused through three repeated boundary modules to obtain a complete salient map with clear boundaries.
[0014] Step S5: Combine the loss function L final Supervised training of the network yields the final prediction graph.
[0015] Furthermore, the feature encoder in step S1 adopts a Swin-Transformer partial structure, including 3 stages; each stage includes multiple layers of repeated transformer-like modules, each transformer module includes core window attention and shift window attention; and each stage can reduce the resolution of the input feature map; for example: the number of channels before the stage output is 128, and the number of channels after passing through three stages are 256, 512 and 1024 respectively.
[0016] Furthermore, in step S2, the shallow features f1 and f2 are firstly optimized for detail information by adaptively modulating spatial attention to highlight the significant target area;
[0017] First, we get the spatial attention map A of the shallow features. s :
[0018] A s =SpatialAttn(f i )
[0019] SpatialAttn(f i )=σ(Conv(Cat(AvgPool(f i ),Maxpool(f i ))))
[0020] Where i refers to the i-th layer, where i = 1, 2, σ(*) represents the sigmoid activation function, Conv represents the convolutional layer, Cat represents the concatenation of features by channel dimension, AvgPooling represents average pooling, and Maxpooling represents maximum pooling;
[0021] The spatial attention map A s And the corresponding feature map f i Multiply, and get the new feature map f′ i :
[0022] f′ i =A s ×f i
[0023] Then, the shallow features are multiplied and processed by 3*3 convolution to obtain the shallow fusion feature f fuse :
[0024] f fuse =Conv(f1×f2)
[0025] The shallow fusion feature f fuseThe common region features are extracted by fusing with f′1 and f′2 through element-wise addition, and then new shallow features F1 and F2 are obtained;
[0026] F i =f′ i +f fuse
[0027] Where F1=f′1+f fuse , F2=f′2+f fuse .
[0028] Furthermore, the specific method based on the Transformer global context module in step S3 is:
[0029] Step S3.1: Downsample the higher-level feature f3 to the same resolution as f4 and add them together to obtain the high-level fusion feature
[0030] Step S3.2: Fusion of high-level features Acquire global context information through a dual-branch structure;
[0031] Branch 1 obtains the global feature G1 with a global receptive field through the Transformer layer:
[0032]
[0033] Among them, T represents a group of Transformer layers, Conv1 represents a 1*1 convolution operation, BN represents a batch normalization layer, and ρ(*) represents a ReLU function;
[0034] Branch 2 is obtained by average pooling The global feature G2 supplements branch 1 and enhances the target area:
[0035]
[0036] Step S3.3, fuse the two branches by element-wise addition, and then use a Sigmoid activation function to generate a weight φ to act on the high-level features f3 and f4 to obtain features F3 and F4 containing sufficient global context information:
[0037] φ=Sigmoid(G1+G2)
[0038] F3=f3×φ
[0039] F4=f4×φ.
[0040] Furthermore, the specific process of step S4 is as follows:
[0041] S4.1. The new shallow features F1, F2 and F3 are further mined through channel attention and residual connection to obtain boundary information, and F′1, F′2, F′3 are obtained. The boundary features are obtained by boundary truth supervision:
[0042] A c = ChannelAtt n(F i )
[0043] ChannelAtt n(F i )=σ(Maxpool(F i ))
[0044] F′ i =F i +A c ×F i , i=1,2,3
[0045] For example, i refers to the i-th level, where
[0046] S4.2. In the boundary-aware decoder, the features F1 to F4 and boundary features F′1 to F′3 obtained by the encoder are input. The highest-level feature F4 is first upsampled to the same size as F3 and fused. Then, the boundary features F′1 to F′3 and F1 to F4 are interactively fused through three repeated boundary modules, and gradually refined to generate a more complete saliency map with clearer edges:
[0047] Here we take one of the boundary modules as an example to describe its processing content:
[0048] The boundary module in the decoder consists of two branches. One branch upsamples the boundary features twice to the same resolution as the next layer features. The two are multiplied pixel by pixel to obtain the edge area E1:
[0049] E1=Up(F′3)×F2
[0050] Subtract the boundary feature from 1 to get the region without the boundary
[0051]
[0052] The other branch upsamples feature F3 and The same size, then remove the border area Multiplying with feature F3 emphasizes the more distinct salient area E2:
[0053]
[0054] S4.3, the salient region E2 and the complementary edge region E1 are fused to obtain F″2, and the new feature map F″2 is fused with the next layer of features and then repeatedly sent to the boundary module:
[0055] F″2=E1+E2
[0056] F2=F″2+F2.
[0057] Furthermore, the joint loss function L in step S5 final for:
[0058] L final =L s +L e
[0059] Among them, L s is the significant loss function, L e is the boundary loss function;
[0060] The significant loss function is specifically:
[0061] L s =L bce +L dice +L smooth
[0062] Among them, L bce is the binary cross entropy BCE loss function, L dice is the Dice loss function, L smooth is the smoothness loss function;
[0063] Given a saliency map S = {S(j)|j = 1, ..., T} and a true value label Y = {Y(j)|j = 1, ..., T}, where j represents the jth pixel and T is the total number of pixels in the saliency map;
[0064] The binary cross entropy BCE loss function is:
[0065]
[0066] The Dice loss function can handle the problem of pixel imbalance between foreground and background areas caused by the presence of many small objects in the dataset:
[0067]
[0068] The smoothness loss function can make the boundary of the salient target smoother and produce a more refined segmentation map, specifically:
[0069]
[0070]
[0071] In the smoothness loss function here, i and j represent the position of each pixel on the horizontal and vertical axes respectively.
[0072] Given a boundary map E = {E(k)|k = 1, ..., K} and a boundary truth label G = {G(k)|k = 1, ..., K}, where k represents the kth pixel and K is the total number of pixels in the boundary map;
[0073] The feature map that focuses on the boundary Upsample to the resolution of the boundary truth map, and then supervised by the binary cross entropy BCE loss function, the boundary loss function L e Specifically:
[0074] L e =α*L1+β*L2+γ*L3
[0075]
[0076]
[0077]
[0078] Among them, α, β, and γ are weight parameters that control different losses, which are 0.5, 0.8, and 1 respectively.
[0079] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0080] (1) This paper introduces Transformer into the remote sensing saliency task for the first time, and designs a Transformer-based global context module to solve the problem of long-distance feature inconsistency, thereby detecting complete salient targets.
[0081] (2) The boundary-aware decoder of the present invention can effectively utilize boundary information and obtain clearer boundaries while paying attention to salient areas.
[0082] (3) The present invention adopts a new joint loss function to alleviate the problem of category imbalance between foreground and background, thereby achieving better detection performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] Figure 1 It is a schematic diagram of the overall process of the present invention;
[0084] Figure 2 It is a schematic diagram of the network model of the present invention;
[0085] Figure 3 FIG. 4 is a schematic diagram of a PR curve according to an embodiment of the present invention. DETAILED DESCRIPTION
[0086] The technical solution of the present invention is described in detail below, but the protection scope of the present invention is not limited to the embodiments.
[0087] like Figure 1 As shown, the optical remote sensing image salient object detection method based on Transformer boundary perception of this embodiment includes the following steps:
[0088] Step S1, use the feature encoder to encode and obtain the multi-level features of the remote sensing image, and mark the remote sensing image features as f1~f4; the feature encoder here adopts the Swin-Transformer partial structure, including 3 stages; each stage includes multiple layers of repeated transformer modules, each transformer module includes window attention and shift window attention; and each stage can reduce the resolution of the input feature map;
[0089] Step S2: Use the spatial attention mechanism to guide the shallow features f1 and f2 to interactively optimize, and fuse and extract the common area feature information to obtain new shallow features F1 and F2;
[0090] First, the shallow features f1 and f2 are optimized by adaptively modulating the spatial attention to highlight the salient target area.
[0091] First, we get the spatial attention map of shallow features:
[0092] A s =SpatialAtt n(f i )
[0093] SpatialAtt n(f i )=σ(Conv(Cat(AvgPool(f i ),Maxpool(f i ))))
[0094] Where i=1,2, σ(*) represents the sigmoid activation function, Conv represents the convolutional layer, Cat represents the concatenation of features by channel dimension, AvgPooling represents average pooling, and Maxpooling represents maximum pooling.
[0095] The spatial attention map is combined with the corresponding feature map f i Multiply, and get the new feature map f′ i :
[0096] f′ i =A s ×fi
[0097] Then, the shallow features are multiplied and processed by 3*3 convolution to obtain the shallow fusion feature f fuse :
[0098] f fuse =Conv(f1×f2)
[0099] The shallow fusion feature f fuse The common region features are extracted by fusing with f′1 and f′2 through element-wise addition, and then new shallow features F1 and F2 are obtained;
[0100] F i =f′ i +f fuse
[0101] Where F1=f′1+f fuse , F2=f′2+f fuse ;
[0102] Step S3, high-level features f3 and f4 are processed by the global information based on the Transformer global context module to obtain features F3 and F4 containing sufficient global context information respectively;
[0103] The specific method based on the Transformer global context module in step S3 is:
[0104] Step S3.1: Downsample the higher-level feature f3 to the same resolution as f4 and add them together to obtain
[0105] Step S3.2: Fusion of high-level features Acquire global context information through a dual-branch structure;
[0106] Branch 1 obtains the global feature G1 with a global receptive field through the Transformer layer:
[0107]
[0108] Among them, T represents a group of Transformer layers, Conv1 represents a 1*1 convolution operation, BN represents a batch normalization layer, and ρ(*) represents a ReLU function;
[0109] Branch 2 is obtained by average pooling The global feature G2 supplements branch 1 and enhances the target area:
[0110]
[0111] Step S3.3, fuse the two branches by element-wise addition, and then use a Sigmoid activation function to generate a weight φ to act on the high-level features f3 and f4 to obtain features F3 and F4 containing sufficient global context information:
[0112] φ=Sigmoid(G1+G2)
[0113] F3=f3×φ
[0114] F4=f4×φ;
[0115] Step S4: In the boundary-aware decoder, the boundary information in the features F1 to F3 is enhanced respectively, and the boundary features F′1 to F′3 are obtained through boundary truth supervision. The boundary features of each layer and the features of the emphasized salient areas are gradually fused through three boundary modules to obtain a complete salient map with clear boundaries.
[0116] S4.1. The new shallow features F1, F2 and F3 are further mined through channel attention and residual connection to obtain boundary information, and F′1, F′2, F′3 are obtained. The boundary features are obtained by boundary truth supervision:
[0117] A c = ChannelAttn(F i )
[0118] ChannelAttn(F i )=σ(Maxpool(F i ))
[0119] F′ i =F i +A c ×F i , i=1,2,3
[0120] S4.2, in the boundary-aware decoder, the features F1~F4 of each level and the boundary features F′1~F′3 obtained by the encoder are input, the highest level feature F4 is first upsampled to the same size as F3 and fused, and then the boundary features F′1~F′3 and F1~F4 are interactively fused through three repeated boundary modules, and gradually refined to generate a more complete saliency map with clearer edges;
[0121] Each boundary module in the decoder consists of two branches. One branch upsamples the boundary features twice to the same resolution as the next layer features. The two are multiplied pixel by pixel to obtain the edge area E1:
[0122] E1=Up(F′3)×F2
[0123] Subtracting the boundary feature from 1 gives the region without the boundary:
[0124]
[0125] The other branch upsamples feature F3 and The same size, then multiply the area without the border and feature F3 to emphasize the clearer salient area E2:
[0126]
[0127] S4.3, the salient region E2 and the complementary edge region E1 are fused to obtain F″2, and the new feature map F″2 is fused with the next layer of features and then repeatedly sent to the boundary module:
[0128] F″2=E1+E2;
[0129] F2=F″2+F2;
[0130] Step S5: Combine the loss function L final Supervise the training network to obtain the final prediction graph;
[0131] Joint loss function L final for:
[0132] L final =L s +L e
[0133] Among them, L s is the significant loss function, L e is the boundary loss function;
[0134] The significant loss function is specifically:
[0135] L s =L bce +L dice +L smooth
[0136] Among them, L bce is the binary cross entropy BCE loss function, L dice is the Dice loss function, L smooth is the smoothness loss function;
[0137] Given a saliency map S = {S(j)|j = 1, ..., T} and a true value label Y = {Y(j)|j = 1, ..., T}, where j represents the jth pixel and T is the total number of pixels in the saliency map;
[0138] The binary cross entropy BCE loss function is:
[0139]
[0140] Dice loss function L dice for:
[0141]
[0142] Smoothness loss function L s Specifically:
[0143]
[0144]
[0145] Given a boundary map E = {E(k)|k = 1, ..., K} and a boundary truth label G = {G(k)|k = 1, ..., K}, where k represents the kth pixel and K is the total number of pixels in the boundary map;
[0146] The feature map that focuses on the boundary Upsample to the resolution of the boundary truth map, and then supervised by the binary cross entropy BCE loss function, the boundary loss function L e Specifically:
[0147] L e =α*L1+β*L2+γ*L3
[0148]
[0149]
[0150]
[0151] Among them, α, β, and γ are weight parameters that control different losses, which are 0.5, 0.8, and 1 respectively.
[0152] Example:
[0153] This embodiment uses public datasets, such as ORSSD and EORSSD. The ORSSD dataset contains 600 images and their corresponding true value labels, including 400 training images and 200 test images. The EORSSD dataset is an extended version of the ORSSD dataset, including 1400 training images and 600 test images.
[0154] In this embodiment, data enhancement is first performed on the EORSSD training set, including random flipping, rotation, cropping, and affine transformation. In order to make the optical remote sensing image salient target detection network of the present invention converge, the target detection network of this embodiment is trained 55 times on the NVIDIATITAN Xp GPU with a batch size of 8. The backbone parameters of the network adopt the parameters of Swin-B, and the network parameters are optimized using the Adams optimizer, with a learning rate of 5e -5, the input image size is 384×384.
[0155] To facilitate quantitative evaluation, this embodiment adopts five widely used indicators.
[0156] (1) Mean absolute error (MAE), which represents the average value of the absolute error between the predicted value and the observed value. MAE is defined as:
[0157]
[0158] Where T is the total number of pixels and S is the predicted saliency map, and Y is the true value map.
[0159] (2) F-measure (Fm) refers to the weighted harmonic mean of precision and recall. The formula for F-measure is:
[0160]
[0161] where β 2 =0.3, which means more attention is paid to accuracy.
[0162] (3) S-measure (S m ). m The target-aware structural similarity (S0) and region-aware structural similarity (S0) between the predicted image and the true value label are calculated. r ). m As shown below:
[0163] S m =α·S0+(1-α)·S r
[0164] where α is set to 0.5.
[0165] (4) E-measure (Em) metric is an enhanced alignment metric that can simultaneously measure global pixel error and local pixel error.
[0166] (5) PR curve: Use the true value label and prediction graph to calculate the precision and recall rate, and draw the PR curve based on the precision and recall rate.
[0167] The closer the PR curve is to the coordinate (1,1), the better the network performance.
[0168] Compare the technical solution of the present invention with other prior arts.
[0169] This embodiment compares the network of the technical solution of the present invention with other 13 methods.
[0170] The comparison methods include not only natural scene image methods, but also optical remote sensing image methods, including DAFNet, MCCNet, EMFINet, LV-Net, MINet, RRNet, MJRBM, PA-KRN, GCPANet, GateNet, SUCA, ITSD, and EGNet. All results are generated by the code provided by the author.
[0171] Compare:
[0172] The specific comparative test results of this embodiment are shown in Table 1. This embodiment uses Em, Sm, Fm, MAE and wFm on three data sets to evaluate the corresponding saliency maps.
[0173] On the ORSSD dataset, compared with the suboptimal MCCNet method, the five evaluation indicators are improved by 0.6%, 0.3%, 0.9%, 0.3% and 1.2% respectively. On the EORSSD dataset, the invention improves by 2.1%, 0.7%, 4.6%, 0.3% and 1.7% on the five evaluation indicators compared with the suboptimal MCCNet method.
[0174] Table 1 Schematic diagram of comparison of measurement indicators
[0175]
[0176] like Figure 3 As shown in Figures 3(a) and 3(b), the PR curves of the prediction graphs on the ORSSD dataset and the EORSSD dataset, respectively; and compared with the results of other technical solutions. It can be observed that the present invention has achieved good performance on the ORSSD and EORSSD datasets, and has achieved significant improvements compared to other technical solutions. This also illustrates the advantageous results of the technical solution of the present invention.
Claims
1. A method for detecting salient objects in optical remote sensing images based on Transformer boundary perception, characterized in that: The following steps are involved: Step S1, using a feature encoder to encode and obtain multi-level features of a remote sensing image, and marking the remote sensing image features as f1-f4; Step S2: guide the shallow features f1 and f2 to interactively optimize through the spatial attention mechanism, and fuse and extract the common area feature information to obtain new shallow features F1 and F2; the specific method is as follows; First, calculate the spatial attention map A of the shallow features s : A s =SpatialAttn(f i ) SpatialAttn(f i )=σ(Conv(Cat(AvgPool(f i ),Maxpool(f i )))) Where i=1,2, σ(*) represents the sigmoid activation function, Conv represents the convolutional layer, Cat represents the concatenation of features by channel dimension, AvgPooling represents average pooling, and Maxpooling represents maximum pooling. The spatial attention map A s And the corresponding feature map f i Multiply, and get the new feature map f i ': f i '=A s ×f i Then, the two shallow features are multiplied and processed by 3*3 convolution to obtain the shallow fusion feature f fuse : f fuse =Conv(f1×f2) The shallow fusion feature f fuse The common region features are extracted by fusing with f1' and f2' through element-wise addition, and then new shallow features F1 and F2 are obtained; Among them, F1=f1'+f fuse , F2=f2'+f fuse ; Step S3, high-level features f3 and f4 are processed by the global information based on the Transformer global context module to obtain features F3 and F4 containing sufficient global context information respectively; Step S4: In the boundary-aware decoder, the boundary information in the features F1 to F3 is enhanced respectively, and the boundary features F1' to F3' are obtained through boundary truth supervision. The boundary features of each layer and the features of the emphasized salient areas are gradually fused through three boundary modules to obtain a complete salient map with clear boundaries. The specific process is as follows: S4.
1. The new shallow features F1, F2 and F3 are further mined through channel attention and residual connection to obtain boundary information, and F1', F2', F3' are obtained. The boundary features are obtained by boundary truth supervision: A c =ChannelAttn(F i ) ChannelAttn(F i )=σ(Maxpool(F i )) F i '=F i +A c ×F i ,i=1,2,3 S4.
2. In the boundary-aware decoder, the features F1 to F4 of each level and the boundary features F1' to F3' obtained by the encoder are input. The highest-level feature F4 is first upsampled to the same size as F3 and fused. Then, the boundary features F1' to F3' and F1 to F4 are interactively fused through three repeated boundary modules, and gradually refined to generate a more complete saliency map with clearer edges. Each boundary module in the decoder consists of two branches. One branch upsamples the boundary features twice to the same resolution as the next layer features. The two are multiplied pixel by pixel to obtain the edge area E1: E1=Up(F3')×F2 Subtract the boundary feature from 1 to get the region without the boundary The other branch upsamples feature F3 and The same size, then the border area will be removed Multiplying with F3 emphasizes the more distinct significant area E2: S4.3, fuse the salient region E2 and the complementary edge region E1 to obtain F2″, and fuse the new feature map F2″ with the next layer of features and repeatedly send it to the boundary module: F″2=E1+E2 F2=F″2+F2; Step S5: Combine the loss function L final Supervise the training network to obtain the final prediction graph; joint loss function L final for: L final =L s +L e Among them, L s is the significant loss function, L e is the boundary loss function; The significant loss function is specifically: L s =L bce +L dice +L smooth Among them, L bce is the binary cross entropy BCE loss function, L dice is the Dice loss function, L smooth is the smoothness loss function; Given a saliency map S = {S(j)|j = 1, ..., T} and a true value label Y = {Y(j)|j = 1, ..., T}, where j represents the jth pixel and T is the total number of pixels in the saliency map; The binary cross entropy BCE loss function is: Dice loss function L dice for: Smoothness loss function L s Specifically: Given a boundary map E = {E(k)|k = 1, ..., K} and a boundary truth label G = {G(k)|k = 1, ..., K}, where k represents the kth pixel and K is the total number of pixels in the boundary map; The feature map that focuses on the boundary Upsample to the resolution of the boundary truth map, and then supervised by the binary cross entropy BCE loss function, the boundary loss function L e Specifically: L e =α*L1+β*L2+γ*L3 Among them, α, β, and γ are weight parameters that control different losses, which are 0.5, 0.8, and 1 respectively.
2. The method for detecting salient objects in optical remote sensing images based on Transformer boundary perception according to claim 1, characterized in that: In the step S1, the feature encoder adopts a Swin-Transformer partial structure, including three stages; each stage includes multiple layers of repeated transformer modules, each transformer module includes window attention and shift window attention; and each stage can reduce the resolution of the input feature map.
3. The method for detecting salient objects in optical remote sensing images based on Transformer boundary perception according to claim 1, characterized in that: The specific method based on the Transformer global context module in step S3 is: Step S3.1: Downsample the high-level feature f3 to the same resolution as f4 and add them together to obtain the high-level fusion feature Step S3.2: High-level fusion features Acquire global context information through a dual-branch structure; Branch 1 obtains the global feature G1 with a global receptive field through the Transformer layer: Among them, T represents a group of Transformer layers, Conv1 represents a 1*1 convolution operation, BN represents a batch normalization layer, and ρ(*) represents a ReLU function; Branch 2 is obtained by average pooling The global feature G2 supplements branch 1 and enhances the target area: Step S3.3, fuse the two branches by element-wise addition, and then use a Sigmoid activation function to generate a weight φ to act on the high-level features f3 and f4 to obtain features F3 and F4 containing sufficient global context information: φ=Sigmoid(G1+G2) F3=f3×φ F4=f4×φ.
Citation Information
Patent Citations
Optical remote sensing image saliency target detection method
CN112347859A
Optical remote sensing image salient target detection method of double-flow decoding cross-task interaction network
CN113505634A