Lightweight significant target detection method for optical remote sensing image based on spatial semantic information interaction
By designing a lightweight detection method for spatial semantic information interaction in optical remote sensing image significance object detection, the problem of poor detection performance in complex scenarios in the prior art is solved, and high-precision significant object detection is achieved, which is better than the existing algorithm.
Patent Information
- Application Number
- CN202510237097.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-01
- Publication Date
- 2025-06-10
AI Technical Summary
The existing optical remote sensing image significance object detection model has poor detection performance in complex scenarios, and the existing natural image detection methods are not suitable for optical remote sensing tasks.
A lightweight significant object detection method for spatial semantic information interaction is designed, including a spatial information enhancement module, a semantic information enhancement module and a spatial semantic information interaction module. Through these modules, high-resolution spatial detail features and high-level semantic features are extracted and fused to enhance the spatial detail information and semantic information of each level of feature.
On the basis of low computing costs, the detection accuracy of significant object in optical remote sensing images is improved, which is better than existing advanced algorithms, especially in complex scenarios, and the detection performance is significantly improved.
Smart Images

Figure CN120126006A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and relates to an improvement of a saliency object detection model, specifically a lightweight saliency object detection method for optical remote sensing images with spatial semantic information interaction. Background Art
[0002] Saliency object detection aims to identify the most prominent or eye-catching objects or regions in an image or video. It has been applied in various computer vision tasks and achieved fruitful results, such as object segmentation, image quality assessment, image re-targeting, etc. In recent years, with the continuous development of deep learning, the research on saliency object detection in optical remote sensing images has received extensive attention. Optical remote sensing images refer to color images taken by satellite and aerial sensors in the range of 400 - 760nm, and there are only three optical bands, which is different from hyperspectral images that contain more spectral band information. The purpose of saliency object detection in optical remote sensing images is to highlight objects such as airplanes, islands, ships, buildings, and rivers that attract human attention at the pixel level of optical remote sensing images. Early methods aimed to utilize handcrafted low-level features such as texture, contrast, color, etc. These features have limitations in distinguishing foreground and background in complex scenes, and their performance also needs to be improved, far behind the current methods based on fully convolutional neural networks. Recently, a large number of natural image saliency object detection methods based on fully convolutional neural networks have also been proposed, and the detection accuracy has been improved, but there is less research on optical remote sensing images. Currently, existing research has proved that the methods of natural image saliency object detection are not suitable for optical remote sensing tasks. Therefore, the present invention designs three modules to improve the saliency object detection accuracy in optical remote sensing images. The specific method is as follows: In order to extract clear-edge spatial detail features, a spatial information enhancement module is designed; in order to extract richer high-level semantic features, a semantic information enhancement module is designed; in order to make full use of high-resolution spatial detail information and enhance the propagation of semantic information, a spatial semantic information interaction module is designed, enabling the network to fully explore the interaction information between saliency object spatial detail features and high-level semantic features. Summary of the Invention
[0003] (1) Technical Problems to be Solved
[0004] Aiming at the deficiencies of the prior art, the present invention provides a lightweight saliency object detection method for optical remote sensing images with spatial semantic information interaction. On the basis of low computational cost, it solves the problem that the current saliency object detection model for optical remote sensing images has poor detection performance in some complex scenes.
[0005] (2) Technical Solutions
[0006] To achieve the above objectives, the present invention proposes a lightweight saliency target detection method for optical remote sensing images with spatial semantic information interaction, namely SSINet. This method can fully explore the saliency target area by using spatial detail information and high-level semantic information, thereby generating saliency target features with high-quality edges and semantics. First, in order to extract spatial detail features with clear edges and rich high-level semantic features, the present invention proposes a Spatial-information Enhancement Module (SPEM) and a Semantic-information Enhancement Module (SEEM). The Spatial-information Enhancement Module fuses the shallowest layer features and the fourth layer features by means of channel connection; the Semantic-information Enhancement Module fuses the deepest layer features and the second layer features by means of channel connection. Second, in order to better utilize the extracted high-resolution spatial detail features and deep high-level semantic features, the present invention proposes a Spatial Semantic-information Interaction Module (SSIM). This module fuses the spatial detail features and high-level semantic features extracted by the above two modules with each layer of features through a correlation module, enhancing the spatial detail information and semantic information of each layer of features. Each Spatial Semantic-information Interaction Module obtains a saliency region feature map. In order to enhance the propagation of semantic information, the present invention proposes a simple Progressive decoder (PD). In this decoder, the output results of each layer of Spatial Semantic-information Interaction Module are fused with the semantic information of the output results of the Spatial Semantic-information Interaction Module of a deeper layer. Finally, a saliency region feature map is obtained, and this saliency region feature map is sent into a convolutional layer to obtain the final detection result of the saliency target.
[0007] A lightweight saliency target detection method for optical remote sensing images with spatial semantic information interaction according to the present invention. The overall architecture of the model of this method is an encoder-decoder network. This method includes the following steps:
[0008] S1. First, extract the spatial detail feature s from the shallowest layer features and the fourth layer features in the encoder network through the Spatial-information Enhancement Module si ; and extract the high-level semantic feature s from the deepest layer features and the second layer features in the encoder network through the Semantic-information Enhancement Module ci , and then use these two extracted features to interactively enhance the spatial details and semantic information of each layer of features.
[0009] S2. The spatial detail feature s siand high-level semantic features s ci are integrated into the spatial semantic information interaction module at each level. The spatial semantic information interaction module makes full use of the high-resolution spatial detail information and the deep high-level semantic information. Through the information interaction of this module, the output results of the spatial semantic information interaction module at each level are generated. The spatial semantic information interaction module makes full use of the high-resolution spatial detail information and the deep high-level semantic information.
[0010] S3. Under the action of the spatial semantic information interaction module, the spatial detail information and high-level semantic information at each level are continuously enhanced. Except for the deepest layer, the output results of the spatial semantic information interaction module at each level will pass through a progressive decoder. In this progressive decoder, the semantic information of the output results of the spatial semantic information interaction module at each level is fused with the semantic information of the output results of the spatial semantic information interaction module at a deeper level, enhancing the propagation of semantic information. Finally, a saliency region feature map is obtained, and this saliency region feature map is input into a convolutional layer to predict the saliency target.
[0011] Preferably, the spatial information enhancement module fuses the upsampled fourth-layer features and the shallowest-layer features through channel connection to extract the spatial detail feature s with clear edges si .
[0012] Preferably, the semantic information enhancement module fuses the downsampled second-layer features and the deepest-layer features through channel connection to extract the richer high-level semantic feature s ci .
[0013] Preferably, the specific implementation steps of S2 are as follows: First, the spatial semantic information interaction module first fuses the spatial detail information of the encoder network at a shallower level by element-wise multiplication, and then captures different context information of the saliency region features through 4 parallel convolutional branches. Then, the spatial attention features at each level are enhanced with the spatial detail information by improving the complementarity of these two features through the spatial correlation module, and at the same time, the channel attention features at each level are enhanced with the semantic information of the saliency target by improving the complementarity of these two features through the channel correlation module. Finally, the fusion results of the above two are fused by element-wise multiplication and residual connection to obtain the saliency target features with enhanced spatial details and semantics.
[0014] Preferably, in the progressive decoder, the semantic information of the output results of the spatial semantic information interaction module at each level is fused with the semantic information of the output results of the spatial semantic information interaction module at a deeper level by channel-wise multiplication to enhance the propagation of semantic information.
[0015] Preferably, the encoder network uses MobileNet V2 to extract the features of the salient region, and the decoder network is designed as a single-branch structure.
[0016] (III) Beneficial Effects
[0017] The present invention provides a lightweight salient object detection method for optical remote sensing images with spatial semantic information interaction, having the following beneficial effects:
[0018] The present invention extracts clear-edge spatial detail features and high-level semantic features by using a spatial information enhancement module and a semantic information enhancement module. The problem of insufficient utilization of high-resolution spatial detail information and deeper semantic information is solved through a spatial semantic information interaction module. The problem of dilution during the propagation of semantic information is alleviated through a progressive decoder.
[0019] The model proposed by the present invention has good performance. Experimental results on two optical remote sensing image datasets, EORSSD and ORSSD, show that the algorithm of the present invention is superior to existing advanced algorithms. Description of the Drawings
[0020] Figure 1 is the overall framework structure diagram of the present invention;
[0021] Figure 2 is the structure diagram of the spatial correlation module used in the present invention. Detailed Embodiments
[0022] The technical methods in the present invention will be clearly and completely described below in conjunction with the drawings. A lightweight salient object detection method for optical remote sensing images with spatial semantic information interaction, and its specific implementation steps are as follows:
[0023] (S1): Design an encoder-decoder network, a spatial information enhancement module, and a semantic information enhancement module.
[0024] The encoder network used in the present invention is the high-efficiency and high-performance MobileNet V2, and its output is {x i |i = 1, 2, 3, 4, 5}, and the decoder network is designed as a single-branch structure. First, as Figure 1 shown on the far left, the spatial information enhancement module (SPEM) proposed by the present invention fuses the shallowest features in the encoder network and the upsampled fourth-level features through channel concatenation, and then through two 3×3 convolutions and spatial attention, extracts clear-edge spatial detail features s si , as follows:
[0025]
[0026] where represents the channel splicing operation, conv represents the convolution operation, SA(·) represents the spatial attention operation, and ↑ represents the upsampling operation.
[0027] As Figure 1 shown on the far right, the semantic information enhancement module (SSEM) proposed by the present invention fuses the deepest features in the encoder network and the downsampled second-level features through channel splicing, and then passes through two 3×3 convolutions and channel attention to extract rich high-level semantic features s ci , as shown below:
[0028]
[0029] where CA(·) represents the channel attention operation and ↓ represents the downsampling operation.
[0030] (S2): Design the spatial semantic information interaction module.
[0031] As Figure 1 shown, the spatial semantic information interaction module (SSIM) proposed by the present invention fuses the spatial detail feature s si and the high-level semantic feature s ci . Substantially, the spatial semantic information interaction module performs three steps. First, it fuses the spatial detail information of a shallower level in the encoder network, then through a separate feature processing operation, and then a feature interaction operation.
[0032] First, the salient region feature fuses the downsampled spatial detail information of a shallower level in the encoder network by element-wise multiplication, and then uses a residual connection to retain the original information, that is, by element-wise addition, as shown below:
[0033]
[0034] where is element-wise multiplication and ⊕ is element-wise addition.
[0035] To capture multi-scale and multi-shape region features and obtain comprehensive information within a feature level, which is beneficial for capturing salient objects of various sizes and shapes in optical remote sensing images. As Figure 1 shown in the multi-scale feature extraction module, the features output above are first divided into 4 parts according to the number of channels, and four parallel branches are designed. Except for the first branch that only uses a 1×1 convolution to retain the original input feature information, each of the remaining branches uses a 1×1 convolution, a 1×(2k - 1) convolution, a (2k - 1)×1 convolution, and a dilated convolution with a dilation rate of 2k - 1, where k = 2, 3, 4. The output after the four branches is The four segmented features pass through these four different multi-scale feature extraction modules, and finally the output features of these different branches are integrated using the method of channel concatenation to generate the output features of the multi-scale feature extraction module as follows:
[0036]
[0037] Next, fuse the spatial detail feature s si extracted by the spatial information enhancement module ci and the high-level semantic feature s si extracted by the semantic information enhancement module. To better fuse the features, the present invention adopts a fusion method of a correlation module. Specifically, after performing the spatial attention mechanism at this level and s ci obtain more refined spatial detail features through the spatial correlation (SCorr) module, and at the same time, after performing the channel attention mechanism at this level and s
[0038]
[0039] such as Figure 2 shown, the spatial correlation module first uses element-wise multiplication to focus on and s si the common area of the spatial information, improve the complementarity of the two along the spatial characteristics, then use the softmax function in the spatial dimension to generate a weight of the spatially co-attended information, and then transfer this weight to and s si respectively, and introduce a short connection to s si in order to integrate the better original spatial features and achieve the purpose of enhancing the spatial information. Finally, the two outputs are concatenated to obtain the output features of the spatial correlation module Figure 2 also visualizes the input and output features of the spatial correlation module. It can be seen that after passing through the spatial correlation module, the spatial information of the features at this level is more refined. The channel correlation module is similar.
[0040] Finally, the output result of the spatial correlation module is interactively fused with the output result of the channel correlation module through element-wise multiplication to generate a stereo attention feature containing spatial detail information and high-level semantic information. After that, the stereo attention feature is fused into In it, a residual connection is used to retain the original information, that is, in the way of element addition, and finally the output features of SSIM-i (i = 1, 2, 3, 4, 5) are generated.
[0041]
[0042] where ⊙ is channel multiplication.
[0043] (S3): Design a progressive decoder.
[0044] As Figure 1 shown at the bottom, in order to enhance the propagation of semantic information, the progressive decoder (PD) module designed by the present invention fuses the output features of the spatial semantic information interaction module by multiplying channels to fuse the high-level semantic information of the deeper-level spatial semantic information interaction module, and then uses a residual connection to retain the original information, that is, in the way of element addition, as follows:
[0045]
[0046] Then, the output features of the deeper-level progressive decoder are fused with x i″ by using the method of channel splicing, and finally a 3×3 convolution is used to obtain the output features of the progressive decoder.
[0047] The effects of the present invention will be described in detail below in combination with experimental data and prediction diagrams.
[0048] Table 1 compares the computational efficiency and accuracy of the method proposed by the present invention with other methods on the EORSSD and ORSSD datasets, and the best scores are shown in bold. It can be seen from the experimental results in Table 1 that the method SSINet proposed by the present invention is superior to 18 other models, among which 4 indicators rank first and 1 indicator ranks second on all datasets. Specifically, on the ORSSD dataset, the model in this paper ranks first in both indicators. Compared with the second-best CorrNet in terms of comprehensive performance, the model in this paper is 0.02% and 0.12% higher than CorrNet in terms of and MAE scores respectively, and is only 0.04% lower than CorrNet in terms of S m , and the number of parameters and the amount of computation are 1.55M and 17.29G lower than CorrNet respectively. On the EORSSD dataset, except for S m which is slightly lower, the other scores are better than other models. Compared with lightweight methods (such as CSNet, SAMNet, MSCNet, FSMINet, CorrNet, and SeaNet), the model in this paper is comprehensively superior to these 6 methods in all indicators.
[0049] Table 1 Comparison of the present invention with advanced methods on the EORSSD and ORSSD datasets
[0050]
[0051] Table 2 Influence of the modules proposed by the present invention on the model performance
[0052]
[0053] Table 2 exemplifies the effectiveness of the modules proposed by the present invention. From the quantitative comparison shown in Table 2, as the proposed SPEM and SEEM increase, it can be seen that and S m scores increase and the MAE scores decrease, indicating the improvement of each proposed module on the overall model performance. In summary, the complete model of the present invention on the EORSSD dataset respectively improves "Baseline" by 0.92%, 1.19% and 0.19% in S m and MAE. On the ORSSD dataset, the performance is also significantly improved, that is, the complete model of the present invention on S m and MAE respectively improves "Baseline" by 0.89%, 1.07% and 0.13%.
[0054] The present invention proposes a lightweight saliency object detection method for optical remote sensing images with spatial semantic information interaction. It effectively extracts fine spatial detail features and rich high-level semantic features through a spatial information enhancement module and a semantic information enhancement module, and fuses them into each layer of the spatial semantic information interaction module designed by the present invention through a spatial correlation module and a channel correlation module respectively, making full use of the high-resolution spatial detail information, effectively alleviating the problem of dilution of deep semantic information during transmission, and further improving the detection accuracy. From a large number of experimental results, the method proposed by the present invention makes full use of and combines spatial detail information and high-level semantic information, and on the basis of low computational cost, makes the saliency object detection performance of optical remote sensing images superior to other advanced algorithms.
[0055] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made therein without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A lightweight salient object detection method for optical remote sensing images with spatial semantic information interaction, characterized by: The overall architecture of the model of this method is an encoder-decoder network; the method includes the following steps: S1. First, the shallowest level features and the fourth level features in the encoder network are extracted through the spatial information enhancement module to extract the spatial detail features s si The deepest and second-level features in the encoder network are extracted through the semantic information enhancement module to extract high-level semantic features s ci , and then use the two extracted features to interactively enhance the spatial details and semantic information of each layer of features; S2. The spatial detail feature s si and high-level semantic features ci It is integrated into the spatial semantic information interaction module at each level, and the output results of the spatial semantic information interaction module at each level are generated through the information interaction of this module; S3. Except for the deepest layer, the output results of each level of spatial semantic information interaction module are passed through a progressive decoder. In this progressive decoder, the output results of each level of spatial semantic information interaction module are fused with the semantic information of the output results of the deeper level of spatial semantic information interaction module to obtain a salient area feature map, and then this salient area feature map is input into a convolutional layer to predict salient targets.
2. The method for detecting lightweight salient objects in optical remote sensing images with spatial semantic information interaction according to claim 1, characterized in that: The spatial information enhancement module fuses the fourth-layer feature upsampling with the shallowest layer feature through channel connection to extract spatial detail features with clear edges. si .
3. The method for detecting lightweight salient objects in optical remote sensing images with spatial semantic information interaction according to claim 1, characterized in that: The semantic information enhancement module fuses the second-layer feature downsampling with the deepest-level features through channel connections to extract richer high-level semantic features. ci .
4. The method for detecting lightweight salient objects in optical remote sensing images with spatial semantic information interaction according to claim 1, characterized in that: The specific implementation steps of S2 are as follows: First, the spatial semantic information interaction module fuses the spatial detail information of the shallower encoder network by element-wise multiplication, and then captures different contextual information of the salient region features through four parallel convolution branches; Then, the spatial attention features of each level are combined with s si The spatial correlation module is used to fuse the channel attention features of each layer with s ci Fusion is performed through a channel correlation module; Finally, the above two fusion results are fused by element-wise multiplication and residual connection to obtain salient target features with spatial detail enhancement and semantic enhancement.
5. The method for detecting lightweight salient objects in optical remote sensing images with spatial semantic information interaction according to claim 1, characterized in that: In the progressive decoder, the output results of each level of spatial semantic information interaction module are fused with the semantic information of the output results of the spatial semantic information interaction module at a deeper level by channel multiplication to enhance the propagation of semantic information.
6. The method for detecting lightweight salient objects in optical remote sensing images with spatial semantic information interaction according to claim 1, characterized in that: The encoder network uses MobileNet V2 to extract salient area features, and the decoder network is designed as a single-branch structure.