An optical remote sensing image significant target detection method and related device

By using the combination method of feature encoding module, edge extraction module and label generation module in the significant object detection of optical remote sensing images, and using technical means such as attention mechanism and bottleneck structure, the existing methods have solved the defects of the details and edge blurring, complex background, and large change in target scales caused by long imaging distances, achieving a more efficient and significant object detection effect.

CN119919647BActive Publication Date: 2025-06-17GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510413227.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-17
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The existing optical remote sensing image significant object detection method has defects in dealing with the details and blurred edges caused by the long imaging distance, complex background, and large changes in the target scale, resulting in unsatisfactory detection results.

Method used

The feature encoding module is used to extract encoding features at different levels through five shared encoders, combine the edge extraction module and the label generation module to perform edge extraction and significant object detection, and use channel and space attention structure to extract content and position information, and reduce information redundancy and improve detection accuracy through bottleneck structure and pyramid attention structure.

Benefits of technology

It effectively improves the effect of significant object detection of optical remote sensing images, especially when dealing with complex backgrounds and multi-scale objects, and solves the shortcomings of existing methods in terms of blurred details and edges, complex backgrounds, and large changes in target scales.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919647B_ABST
    Figure CN119919647B_ABST
Patent Text Reader

Abstract

The present invention discloses an optical remote sensing image salient object detection method and related device. The method includes: inputting the optical remote sensing image to be detected into a feature encoding module to extract encoded features of different levels; inputting the encoded features of different levels into an edge extraction module to perform edge extraction operations to obtain a salient object edge map; and inputting the encoded features of different levels and the salient object edge map into a label generation module to perform salient object detection operations to obtain a salient object map. In the present invention, a more discriminative feature is learned through the feature encoding module to improve the salient detection performance, the edge information is effectively utilized through the edge extraction module to enhance the feature discrimination ability; and the edge and feature ratio relationship and information redundancy problems are effectively processed through the label generation module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision processing technology, and in particular to a method for detecting salient targets in optical remote sensing images and related devices. Background Art

[0002] Salient object detection is a key task in computer vision, which aims to detect the most visually distinctive object from a given image. In recent years, with the development of deep neural networks, salient object detection in natural scene images has become relatively mature, but salient object detection in optical remote sensing images remains to be explored. Salient object detection in optical remote sensing images has strong practical value. This task can be applied to various fields of computer vision to play a preprocessing role, such as object segmentation, visual tracking, image retrieval, cropping, image quality assessment, etc.

[0003] In recent years, many technologies have been introduced into optical remote sensing image salient object detection, such as scale fusion, edge guidance, complementary loss, etc. For example, LVNet uses convolutional networks, feature pyramids and specific modules to suppress background and highlight targets; DAFNet uses global context-aware attention (GCA) to capture semantic contextual relationships; ERPNet, composed of an encoder and two parallel decoders, uses edge information to calibrate the decoding process. However, existing optical remote sensing image salient object detection methods have defects in dealing with details and edge blur caused by long imaging distances, complex backgrounds, and large changes in target scales, resulting in unsatisfactory detection results. Summary of the invention

[0004] The present invention provides a method and related device for detecting salient targets in optical remote sensing images, which are used to solve the defects of existing methods for detecting salient targets in optical remote sensing images in processing the details and edges blurred, the background complex, the large change of the target scale caused by the long imaging distance, etc., improve the effect of detecting salient targets in remote sensing images, and make it more advantageous when processing complex backgrounds and multi-scale objects.

[0005] The present invention provides a method for detecting salient targets in optical remote sensing images, the method comprising:

[0006] Input the optical remote sensing image to be detected into the feature encoding module to extract the encoding features at different levels;

[0007] Inputting the encoding features of different levels into an edge extraction module to perform edge extraction operations to obtain a salient object edge map;

[0008] The encoding features at different levels and the salient object edge map are input into a label generation module to perform a salient object detection operation to obtain a salient object map.

[0009] Further, the feature encoding module is composed of five shared encoders; each of the shared encoders consists of a basic convolutional sub-module and a CASA sub-module;

[0010] Among them, the basic convolutional sub-module is constructed based on the ResNet-34 backbone network or the VGG-16 backbone network; the CASA sub-module includes a channel attention structure and a spatial attention structure, which are used to extract content information and position information respectively.

[0011] Further, the calculation and processing process of the CASA sub-module is as follows:

[0012]

[0013] In the formula: represents the encoded feature of the i-th layer, where ; represents the spatial feature obtained by performing spatial attention processing through the spatial attention structure of the i-th layer, represents the shallow feature output by the basic convolutional sub-module corresponding to the i-th layer; represents element-wise addition, represents element-wise multiplication; represents the channel feature obtained by performing channel attention processing through the channel attention structure of the i-th layer, represents the spatial weight of the spatial attention structure of the i-th layer; represents the channel weight of the channel attention structure of the i-th layer.

[0014] Further, the edge extraction module is composed of a significant convolution head and five cascaded edge decoders;

[0015] In the edge extraction module, edge extraction operations are performed on the encoded features of different layers through five cascaded edge decoders to generate the edge features of the first layer; the significant convolution head performs convolution operations on the edge features of the first layer to obtain the significant target edge map.

[0016] Further, the edge extraction operation process of the edge extraction module is as follows:

[0017]

[0018] In the formula: represents the feature map of the i-th layer, represents concatenation, represents the encoded feature of the fifth layer, represents the fifth-layer edge decoder performs edge feature extraction operations on the encoded feature of the fifth layer to obtain the fifth-layer edge feature;

[0019]

[0020] In the formula: represents the feature map of the i-th layer, represents the encoded feature of the i-th layer, represents the edge decoder of the i-th layer for the encoded feature of the i-th layer and the feature map of the (i + 1)-th layer The edge feature of the i-th layer obtained by performing an edge feature extraction operation;

[0021]

[0022] In the formula: represents the edge feature of the first layer obtained by the edge decoder of the first layer for the encoded feature of the first layer and the feature map of the second layer by performing edge feature extraction;

[0023]

[0024] In the formula: represents the significant object edge map, represents a convolution operation.

[0025] Furthermore, the label generation module consists of five decoders with bottleneck structures, five edge embedding structures, and a significant convolution head;

[0026] The significant object detection operation process of the label generation module is as follows:

[0027]

[0028] In the formula: represents the scale feature of the fifth layer, represents the edge capture operation of the edge embedding structure, represents the decoding operation of the decoder with a bottleneck structure of the fifth layer;

[0029]

[0030] In the formula: represents the scale feature of the i-th layer, represents the decoding operation of the decoder with a bottleneck structure of the i-th layer;

[0031]

[0032] In the formula: represents the significant object map; Represents the scale feature of the first level.

[0033] Furthermore, the edge embedding structure includes a residual connection structure and a pyramid attention structure;

[0034] In the residual connection structure, smooth the significant target edge map, downsample the result of the smoothing operation to obtain a downsampled feature; residually connect the downsampled feature and the decoded feature obtained by the decoding operation of the decoder with a bottleneck structure at different levels to obtain enhanced features at different levels; wherein, the pyramid attention structure is used to extract the scale information of the enhanced features at different levels;

[0035] The data processing process of the edge embedding structure is as follows:

[0036]

[0037] In the formula: Represents the enhanced feature of the i-th level, Represents the smoothing operation, Represents the downsampling operation, Represents the residual connection, where ;

[0038]

[0039] In the formula: Represents the scale feature of the i-th level, Represents the pyramid scale operation of the pyramid attention structure.

[0040] The present invention also provides an optical remote sensing image significant target detection system, and the system includes:

[0041] An encoding unit, configured to input an optical remote sensing image to be detected into a feature encoding module to extract encoded features at different levels;

[0042] An edge extraction unit, configured to input the encoded features at different levels into an edge extraction module to perform an edge extraction operation to obtain a significant target edge map;

[0043] A label generation unit, configured to input the encoded features at different levels and the significant target edge map into a label generation module to perform a significant target detection operation to obtain a significant target map.

[0044] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the optical remote sensing image significant target detection method as described above.

[0045] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the optical remote sensing image significant target detection method described above are implemented.

[0046] As can be seen from the above technical solutions, the present invention has the following advantages:

[0047] The present invention provides an optical remote sensing image significant target detection method and related device. The method includes: inputting an optical remote sensing image to be detected into a feature encoding module to extract encoded features of different levels; inputting the encoded features of different levels into an edge extraction module for edge extraction operations to obtain a significant target edge map; inputting the encoded features of different levels and the significant target edge map into a label generation module for significant target detection operations to obtain a significant target map.

[0048] In the present invention, the feature encoding module makes the learned features more discriminative by introducing channel information and spatial information; the edge extraction module effectively utilizes edge information to improve the significant target recognition ability; the label generation module reduces information redundancy, aligns edges and features, and improves the model performance and efficiency by introducing a bottleneck structure and an edge embedding structure. The present invention solves the defects of the existing optical remote sensing image significant target detection methods in dealing with details and edge blurring, complex background, large target scale change, etc. caused by a long imaging distance, improves the detection effect of significant targets in remote sensing images, and has more advantages in dealing with complex backgrounds and multi-scale objects. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0050] Figure 1 It is a flowchart of the steps of an optical remote sensing image significant target detection method provided by an embodiment of the present invention;

[0051] Figure 2 It is a framework diagram of the CSENet network structure provided by an embodiment of the present invention;

[0052] Figure 3 It is a block diagram of the structure of a shared encoder provided by an embodiment of the present invention;

[0053] Figure 4 It is a schematic diagram of the bottleneck structure provided by an embodiment of the present invention;

[0054] Figure 5Schematic diagram of the edge embedding structure provided by the embodiment of the present invention;

[0055] Figure 6 Schematic diagram of the pyramid attention structure provided by the embodiment of the present invention;

[0056] Figure 7 Block diagram of the structure of an optical remote sensing image salient object detection system provided by the embodiment of the present invention. Detailed implementation manners

[0057] The embodiment of the present invention provides an optical remote sensing image salient object detection method and related devices, which are used to solve the defects of the existing optical remote sensing image salient object detection methods in dealing with the details and edge blurring, complex background, large target scale variation, etc. caused by the far imaging distance, improve the detection effect of the salient object in the remote sensing image, and make it more advantageous in dealing with complex backgrounds and multi-scale objects.

[0058] In order to make the invention objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the embodiments described below are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0059] Please refer to Figure 1 and Figure 2 , the first aspect of the present invention provides an optical remote sensing image salient object detection method, and the method includes:

[0060] Step 101: Input the optical remote sensing image to be detected into the feature encoding module to extract encoded features of different levels.

[0061] It should be noted that, please refer to Figure 2 , the present invention is applied to the CSENet network structure; as Figure 3 shown, the feature encoding module is composed of five shared encoders; the optical remote sensing image to be detected can be a 256×256 RGB image, and then it is input to the five shared encoders in sequence according to different levels to generate encoded features of different levels .

[0062] Each shared encoder is composed of a basic convolution sub-module (Base-conv) and a CASA sub-module; the basic convolution sub-module can be constructed using a popular neural network as the backbone network, such as the ResNet-34 backbone network or the VGG-16 backbone network; the CASA sub-module includes a channel attention structure (CA) and a spatial attention structure (SA), which are used to extract content information and position information respectively.

[0063] 1) For the basic convolutional sub-module:

[0064] For the ResNet-34 backbone network, this implementation further modifies the structure of the ResNet-34 backbone network: Base-conv1 under the first layer is composed of a convolutional layer, a batch normalization layer, and a ReLU activation function. The convolutional layer has 3 input channels (I-C), 64 output channels (O-C), a 3×3 convolutional kernel size (K), and 1 padding (P); then the first layer (layer1) of ResNet34 is stacked, and the output size is 256×256. Under the second to fourth layers, namely Base-conv2, Base-conv3, and Base-conv4, the second, third, and fourth layers of the ResNet-34 model are directly used respectively, and the output sizes are 128×128, 64×64, and 32×32 respectively. The last Base-conv5 is composed of a max-pooling layer (kernel size = 2×2 and stride = 2) and three basic block modules; the max-pooling layer is used for downsampling to reduce the output size to 16×16; the basic block is composed of two 3×3 convolutional layers, with 512 input channels and 512 output channels.

[0065] For the VGG-16 backbone network, this implementation removes the fully connected layer and makes some other structural adjustments. Specifically: First, Base-conv1 consists of two convolutional operations. The first convolutional layer converts the input image with 3 channels into 64 channels (kernel size = 3×3, stride = 1, and padding = 1), followed by a batch normalization layer and a ReLU activation function; the input and output channels of the second convolutional layer are also 64, and the other settings are the same as those of the first convolutional layer; after that, batch normalization and ReLU activation are performed again. Then, Base-conv2 consists of a max-pooling layer (kernel size = 2×2 and stride = 2) and two convolutional operations. Under the third to fifth layers, namely Base-conv3, Base-conv4, and Base-conv5, they all consist of a max-pooling layer (kernel size = 2×2 and stride = 2) and three convolutional operations; these four convolutional modules have a total of 11 convolutional layers (each convolutional layer is followed by a batch normalization layer and a ReLU activation function); except for the number of input and output channels, the settings of the kernel size, stride, and padding are the same as those of the first convolutional layer in Base-conv1.

[0066] It can be understood that the backbone network of the shared encoder can be replaced by other network structures with similar feature extraction capabilities, such as ResNet-50, etc., instead of ResNet-34 and VGG-16, and appropriate adjustments are made to adapt to the module connection and function implementation of the present invention.

[0067] 2) For the CASA sub-module:

[0068] Please refer to Figure 3 , the CASA sub-module uses the shallow features output by the basic convolutional sub-module as the input, and the features sequentially enter the channel attention structure (CA) and the spatial attention structure (SA). Specifically: the shallow features are multiplied element-wise with the channel weights of the channel attention structure to obtain the channel features ; then, the spatial weights are multiplied element-wise with the channel features to obtain the spatial features ; finally, and the initial input are added element-wise to obtain the final output ; here is the encoded feature , and the above process is expressed by the formula as follows:

[0069]

[0070] In the formula: represents element-wise addition, represents element-wise multiplication.

[0071] It should be noted that the channel attention structure (CA) consists of an adaptive max pooling layer and two convolutional layers. After the first convolutional layer is the ReLU activation function, and after the second convolutional layer is the Sigmoid activation function; CA mainly extracts content information by determining the importance of different channels. Usually, an image will generate a series of feature maps through a convolutional network. Each channel of the feature map contains different levels of image features, such as edges, textures, or colors. The channel attention mechanism weights the channels, allowing the model to emphasize or suppress information from specific channels.

[0072] In the spatial attention structure (SA), first, the average value and the maximum value of the input feature map in the channel dimension are calculated, and then the two feature maps are concatenated to perform a convolutional operation using the Sigmoid activation function. SA mainly extracts the location information crucial for localization. It generates a spatial attention map by assigning weights to different spatial positions, allowing the model to focus on specific regions and ignore irrelevant background information.

[0073] Step 102, input the encoded features at different levels into the edge extraction module for edge extraction operations to obtain the significant object edge map.

[0074] It should be noted that the edge extraction module consists of a significant convolutional head and five cascaded edge decoders is composed of five cascaded edge decoders to perform edge extraction operations on the encoded features at different levels During the processing, the edge extraction module processes in the order from deep to shallow, that is, from the edge decoder at the deepest level to the edge decoder at the shallowest level to obtain the first-level edge features ; Finally, through the saliency convolution head perform convolution operations on the first-level edge features to obtain the saliency object edge map .

[0075] These five decoders are all composed of three convolutional layers, followed by a batch normalization layer and a ReLU activation function after each convolutional layer, and the parameter settings of its convolutional layers are shown in Table 1

[0076] Table 1 Parameter settings of convolutional layers in the edge decoder of the edge extraction module

[0077]

[0078] Among them using the encoded features of the fifth level as the input, output the fifth-level edge features ; Then connect and , and then through an upsampling operation to obtain a fifth-level feature map with the same size as the encoded features of the fourth level , and concatenate and as the input of , as follows :

[0079]

[0080] In the formula: represents concatenation

[0081] Except , the other four decoders use the encoded features of the same level and the edge features of deeper levels as the input, as follows

[0082]

[0083] In the formula: represents the feature map of the i-th level, represents the i-th level edge decoder to perform edge feature extraction operations. This fusion method is expected to improve accuracy and edge detection

[0084] Finally, through the saliency convolution head Obtain an edge map , as follows:

[0085]

[0086] Among them, is composed of four convolutional layers. All these convolutional layers do not require batch normalization, and the fourth convolutional layer has no activation function, while the other three convolutional layers are connected to the ReLU activation function; The specific parameter settings of are shown in Table 2, where x in Table 2 is the number of input channels.

[0087] Table 2 Parameters of the convolutional layer

[0088]

[0089] Step 103, input the encoded features and the significant object edge map at different levels into the label generation module for significant object detection operations to obtain a significant object map.

[0090] Among them, the label generation module mainly consists of five decoders with bottleneck structures , five edge embedding structures (EES) with pyramid attention structures (PA), and a significant convolution head. During the significant object detection operation, the label generation module is also processed in the order from deep to shallow, starting from the decoder and moving sequentially to the first-level decoder. After the decoded features at different levels generated by the decoder , the decoded feature is first connected through a residual connection to the output by the edge extraction module. Here, plays an auxiliary supervision role, guiding the edge embedding structure to better understand the shape and boundary of the object. Then, the connected features are passed through the pyramid attention module to further highlight the regions containing significant objects, obtaining scale features. Next, the scale features and the encoded features of the previous layer enter the previous layer decoder together to generate new scale features, and finally the first-level scale feature is obtained. The significant convolution head performs a convolution operation on the first-level scale feature to generate a significant object map .

[0091] Specifically, first, the decoder with a bottleneck structure at the fifth level performs a decoding operation on the encoded features at the fifth level and outputs the fifth-level decoded feature ; then, the decoded feature is fused through the edge embedding structure (EES) With the significant target edge map , the fifth-level scale features are generated , as shown in the following formula:

[0092]

[0093] Wherein, represents the edge capture operation of the edge embedding structure;

[0094] The other four decoders of the label generation module all use the encoded features of the same level and the scale features from deeper levels as inputs, that is:

[0095]

[0096] In the formula: represents the scale feature of the i-th level, represents the decoding operation of the decoder with a bottleneck structure at the i-th level.

[0097] Finally, through a convolution operation is performed to generate the significant target map , that is:

[0098]

[0099] It should be noted that the of the label generation module is similar to the in the edge extraction module, but the input channel number of the of the label generation module is 128.

[0100] In terms of structure, has a bottleneck structure with three convolutional layers, and the bottleneck structure is as Figure 4 shown. There is a batch normalization layer and a ReLU activation function behind each convolutional layer. The number of channels from its input to output is 512, 128, 128, 512 in sequence. This not only reduces the number of parameters and computational complexity, but also filters out redundant features. The remaining four decoders are all composed of upsampling, feature connection, channel dimension reduction, and bottleneck structure filtering.

[0101] Plays a crucial role in the data fusion of the significant target edge map and the decoded features . It helps the network capture better boundaries of the objects in the region of interest. As Figure 5 shown, is mainly composed of a residual connection structure (RC) and a pyramid attention structure (PA).

[0102] In the residual connection structure (RC), for the significant object edge map perform function smoothing operation, downsample ( ) the result of the function smoothing operation to the same size as the corresponding decoded feature to obtain the downsampled feature; using the downsampled feature as auxiliary information, perform residual connection between the downsampled feature and the decoded feature obtained through the decoding operations of the decoders with bottleneck structures at different levels to obtain enhanced features at different levels , as shown in the following formula:

[0103]

[0104] In the formula: .

[0105] To effectively process objects or regions with different scales, feed the enhanced feature into the PA structure to obtain the scale feature , that is:

[0106]

[0107] where, represents the pyramid scale operation of the pyramid attention structure.

[0108] Please refer to Figure 6 , the pyramid attention structure is set with a single-scale branch and a multi-scale branch, then the output of the pyramid attention structure can be single-scale information or multi-scale information , depending on the specific choice in the structure, as shown below:

[0109]

[0110] The single-scale branch introduces the spatial attention structure (SA) to obtain the spatial attention weight , and then obtains the single-scale output through element-wise multiplication and element-wise addition, that is:

[0111]

[0112] Similarly, the multi-scale branch also uses as the input, but uses the cascaded operation pyramid attention ( ) to obtain information of various scales, such as Figure 6As shown in the lower dotted box. This branch is divided into three streams. Except for the sizes of the input and output, the operation mode of each stream is similar to the previous single-scale operation. In addition, the output of the previous stream is also concatenated (C) with the subsequent stream after being upsampled (US).

[0113] Specifically, the first stream uses the features of four-fold downsampling (DS4) to generate new features to capture large objects or regions, that is:

[0114]

[0115] where .

[0116] Next, the second stream integrates the features of two-fold downsampling (DS2) with the output of the previous stream, that is:

[0117]

[0118] where,

[0119]

[0120] Finally, the last stream further integrates the features of the original size to adapt to objects or regions with different scales, that is:

[0121]

[0122] where,

[0123]

[0124] It can be understood that by changing the combination of the number of streams and the sampling multiple in the multi-scale branch, a better multi-scale feature extraction method can be explored while maintaining its function of integrating features of different scales. In addition, in the decoders of the edge extraction module and the label generation module, the parameter settings (such as the convolution kernel size, number, etc.) of the convolutional layer and the type of activation function can be adjusted. For example, different non-linear activation functions can be tried to replace ReLU to observe the impact on the model performance.

[0125] An optical remote sensing image salient object detection method provided by the present invention has the following advantages:

[0126] 1. Design a model that can effectively utilize edge information and enhance feature discrimination ability to solve the problems caused by imaging distance in the salient object detection of optical remote sensing images and accurately detect salient objects.

[0127] 2. Improve the cross-scene generalization and the ability to suppress irrelevant objects of the model through the attention mechanism and specific module design to cope with complex backgrounds; at the same time, use modules such as pyramid attention to effectively detect multi-scale objects and solve the problem of scale change.

[0128] 3. Improve the encoder and decoder structures to retain positional information, handle edge and feature scale relationships, and reduce information redundancy. Among them, in the encoder, CASA enables the encoder to learn more discriminative features, better retain positional information and extract effective features than existing encoders, so as to more accurately identify and locate targets in the detection of prominent targets in optical remote sensing images. In the decoder, the bottleneck structure and PA unit of the label generation module effectively handle edge and feature scale relationships and information redundancy problems, improve the model performance and efficiency, and have more advantages when dealing with complex backgrounds and multi-scale objects, showing obvious improvements compared with existing decoders.

[0129] Please refer to Figure 7 , the second aspect of the present invention provides an optical remote sensing image prominent target detection system, which includes:

[0130] An encoding unit 201, configured to input the optical remote sensing image to be detected into a feature encoding module to extract encoded features at different levels;

[0131] An edge extraction unit 202, configured to input the encoded features at different levels into an edge extraction module for edge extraction operations to obtain a prominent target edge map;

[0132] A label generation unit 203, configured to input the encoded features at different levels and the prominent target edge map into a label generation module for prominent target detection operations to obtain a prominent target map.

[0133] The third aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the optical remote sensing image prominent target detection method as described above.

[0134] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the optical remote sensing image prominent target detection method as described above.

[0135] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0136] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical, or other forms.

[0137] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0138] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0139] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0140] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A method for detecting salient objects in optical remote sensing images, characterized in that: The method comprises: Input the optical remote sensing image to be detected into the feature encoding module to extract the encoding features at different levels; Inputting the encoding features of different levels into an edge extraction module to perform edge extraction operations to obtain a salient object edge map; Inputting the encoding features of different levels and the salient object edge map into a label generation module to perform a salient object detection operation to obtain a salient object map; The feature encoding module is composed of five shared encoders; each of the shared encoders is composed of a basic convolution submodule and a CASA submodule; The basic convolution submodule is constructed based on the ResNet-34 backbone network or the VGG-16 backbone network; the CASA submodule includes a channel attention structure and a spatial attention structure for extracting content information and location information respectively; The edge extraction module consists of a significant convolution head and five edge decoders connected in series; In the edge extraction module, five edge decoders connected in series are used to perform edge extraction operations on the coding features of different levels to generate edge features of the first level; the salient convolution head is used to perform convolution operations on the edge features of the first level to obtain a salient target edge map; The label generation module consists of five decoders with bottleneck structures, five edge embedding structures and a significant convolution head; The salient object detection operation process of the label generation module is as follows: ; Where: Represents the scale characteristics of the fifth level, The edge capture operation representing the edge embedding structure, represents the decoding operation of the decoder with a bottleneck structure at the fifth level; ; Where: represents the scale feature of the i-th level, represents the decoding operation of the decoder with bottleneck structure at the i-th level; ; Where: represents a salient target map; Indicates the scale characteristics of the first level; The edge embedding structure includes a residual connection structure and a pyramid attention structure; In the residual connection structure, a smoothing operation is performed on the salient target edge map, and the smoothing operation result is downsampled to obtain downsampled features; the downsampled features are residually connected with decoded features obtained by decoding operations of decoders with bottleneck structures at different levels to obtain enhanced features at different levels; wherein the pyramid attention structure is used to extract scale information of the enhanced features at different levels; The data processing process of the edge embedding structure is as follows: ; Where: represents the enhanced features of the i-th level, represents smoothing operation, represents the downsampling operation, represents the residual connection, where ; ; Where: represents the scale feature of the i-th level, Representing the pyramid scale operation of the pyramid attention structure.

2. The optical remote sensing image salient object detection method according to claim 1, characterized in that: The calculation process of the CASA submodule is as follows: ; Where: represents the encoding feature of the i-th level, where ; represents the spatial features obtained by spatial attention processing through the i-th level spatial attention structure, Represents the shallow features output by the basic convolutional submodule corresponding to the i-th level; represents the addition of elements, Represents element-wise multiplication; represents the channel features obtained by channel attention processing through the i-th level channel attention structure, represents the spatial weight of the spatial attention structure at the i-th level; Represents the channel weight of the i-th level channel attention structure.

3. The optical remote sensing image salient object detection method according to claim 1, characterized in that: The edge extraction operation process of the edge extraction module is as follows: ; Where: represents the feature map of the i-th level, Indicates series connection, represents the encoding features of the fifth level, Represented by the fifth level edge decoder Encoding features for the fifth level The fifth level edge features are obtained by performing edge feature extraction operation; ; Where: represents the feature map of the i-th level, represents the encoding features of the i-th level, Represents the edge decoder at level i The encoding features of the i-th level and the feature map of level i+1 Perform edge feature extraction to obtain the edge features of the i-th level; ; Where: Represented by the first level edge decoder Encoding features of the first level and the feature map of the second level Perform edge feature extraction to obtain the first level of edge features; ; Where: represents the salient object edge map, Represents a convolution operation.

4. A salient target detection system for optical remote sensing images, characterized in that: The system comprises: The encoding unit is used to input the optical remote sensing image to be detected into the feature encoding module to extract the encoding features of different levels; An edge extraction unit, used for inputting the coding features of different levels into an edge extraction module to perform edge extraction operations to obtain a salient object edge map; A label generation unit, configured to input the encoding features of different levels and the salient object edge map into a label generation module to perform a salient object detection operation to obtain a salient object map; The edge extraction module consists of a significant convolution head and five edge decoders connected in series; In the edge extraction module, five edge decoders connected in series are used to perform edge extraction operations on the coding features of different levels to generate edge features of the first level; the salient convolution head is used to perform convolution operations on the edge features of the first level to obtain a salient target edge map; The label generation module consists of five decoders with bottleneck structures, five edge embedding structures and a significant convolution head; The salient object detection operation process of the label generation module is as follows: ; Where: Represents the scale characteristics of the fifth level, The edge capture operation representing the edge embedding structure, represents the decoding operation of the decoder with a bottleneck structure at the fifth level; ; Where: represents the scale feature of the i-th level, represents the decoding operation of the decoder with bottleneck structure at the i-th level; ; Where: represents a salient target map; Indicates the scale characteristics of the first level; The edge embedding structure includes a residual connection structure and a pyramid attention structure; In the residual connection structure, a smoothing operation is performed on the salient target edge map, and the smoothing operation result is downsampled to obtain downsampled features; the downsampled features are residually connected with decoded features obtained by decoding operations of decoders with bottleneck structures at different levels to obtain enhanced features at different levels; wherein the pyramid attention structure is used to extract scale information of the enhanced features at different levels; The data processing process of the edge embedding structure is as follows: ; Where: represents the enhanced features of the i-th level, represents smoothing operation, represents the downsampling operation, represents the residual connection, where ; ; Where: represents the scale feature of the i-th level, Representing the pyramid scale operation of the pyramid attention structure.

5. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the optical remote sensing image salient object detection method as described in any one of claims 1-3.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the optical remote sensing image salient object detection method as described in any one of claims 1 to 3 are implemented.

Citation Information

Patent Citations

  • Optical remote sensing image saliency target detection method

    CN119649233A

  • Generating synthesized digital images utilizing a multi-resolution generator neural network

    US20230053588A1