A Salient Object Detection Method for High-Resolution Remote Sensing Images Combining Frequency and Edge Learning

By combining edge learning and frequency channel attention, the problems of unclear edges and missing spectral features in high-resolution remote sensing images are solved, achieving higher detection accuracy and richer feature acquisition.

CN114529829BActive Publication Date: 2025-06-20ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210019714.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-10
Publication Date
2025-06-20
Estimated Expiration
2042-01-10

AI Technical Summary

Technical Problem

Due to the wide coverage area, complex targets and spectral complexity of high-resolution remote sensing images, ordinary deep learning networks find it difficult to obtain accurate edge features and rich remote sensing spectral information, resulting in insufficient detection accuracy.

Method used

Using a method that combines edge learning and frequency channel attention, we obtain precise edge features through edge learning, and use a frequency-based channel attention mechanism (FCA) to compensate for the loss of feature information caused by dimensionality reduction operations to obtain more frequency features.

Benefits of technology

The accuracy of the significance target detection of high-score remote sensing images is improved, the acquisition of image feature information is enhanced, and the problems of unclear edges and missing spectral features are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114529829B_ABST
    Figure CN114529829B_ABST
Patent Text Reader

Abstract

A method for detecting significant targets in high-resolution remote sensing images by combining frequency and edge learning, including: drawing samples according to the contours of target ground objects, making edge detection samples and semantic segmentation samples; building a neural network model that introduces edge learning and a spectrum-based channel attention mechanism, inputting high-resolution remote sensing images, edge detection samples, and semantic segmentation samples into training until they are in a fitting state to obtain a significant target detection model for remote sensing images; inputting high-resolution remote sensing images to obtain a predicted image. The present invention combines edge learning and a spectrum-based channel attention mechanism, solving the problem of unclear edges in significant detection in the traditional remote sensing field and the problem of inaccurate detection caused by the lack of spectral features when using deep learning methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of high-resolution remote sensing image processing and the technical field of saliency detection in the field of computer vision, and in particular to a method for realizing saliency target detection in high-resolution remote sensing images, which is suitable for saliency detection of ground objects in high-resolution remote sensing images. Background Art

[0002] The salient target detection technology can quickly and effectively extract the target area in the scene for further analysis based on the salient features (spatial domain, frequency domain, etc.). High-resolution remote sensing images have rich geographic information and spectral information. The computational complexity of the whole image information processing is huge. Combined with the salient target detection technology, only the target area of ​​most interest to people in the image can be extracted, thereby narrowing the scope of information processing and improving the efficiency of information processing. In addition, some deep learning networks can improve the target saliency by integrating local features into salient objects. However, high-resolution remote sensing images have the characteristics of wide coverage, complex targets, and spectral complexity. Ordinary deep learning networks cannot obtain accurate edge features of remote sensing images and extract rich remote sensing spectral information. At present, saliency detection in the remote sensing field mainly adopts the fusion of "bottom-up" and "top-down" models. Although this method takes into account both global information and local detailed information, and compensates for the incompleteness of target detection shapes and lack of accuracy to a certain extent, it is still insufficient for the detection ability of complex remote sensing images containing multiple types. In the future, it is still necessary to expand the types of remote sensing images and optimize the detection algorithm. The salient target detection method based on deep learning has the ability to learn multi-layer features, so it is easier to determine the salient area. Therefore, how to extract more remote sensing feature information and improve detection accuracy is the problem to be solved by the present invention. Summary of the invention

[0003] In order to solve the above problems, the present invention proposes a remote sensing salient target detection method that combines edge learning with frequency channel attention. In the model involved, the edge feature prediction results obtained by edge learning are fused with the prediction results of the semantic segmentation target to obtain a more accurate edge salient target; a frequency-based channel attention mechanism FCA (Frequency Channel Attention) is used to compensate for the loss of feature information caused by dimensionality reduction operations, and more frequency features are obtained to enhance the feature information of the image. Use data enhancement technology to expand the amount of data, and use transfer learning to enhance the training effect.

[0004] The method for detecting salient targets in high-resolution remote sensing images by combining frequency and edge learning of the present invention includes the processes of building, training, and predicting a salient detection network model, and includes the following steps:

[0005] Step 1: Make high-resolution remote sensing image samples, manually draw the precise boundaries of the targets and label the semantic types, perform rasterization processing to generate corresponding line label samples and surface label samples, and use data augmentation methods to expand the samples.

[0006] Step 2: Build a high-resolution remote sensing image saliency target detection network model;

[0007] Step 2.1: Select a suitable CNN network as the backbone. The data after convolution in the lower layers is used for the target edge feature extraction module, and the convolution in other layers is used for the saliency target feature extraction module;

[0008] Step 2.2: Connect the FCA attention mechanism after the convolution with dimensionality reduction. Its model is as Figure 3 shown, giving two-dimensional weights to the features of the image, which strengthens the channel information and the frequency information of the remote sensing image, where

[0009] att = sigmoid(fc(Freq)) (1)

[0010] Freq = [Freq 0 , Freq 1 , …, Freq n-1

[0011] Freq i = 2DDCT u,v (X i ), i ∈ (0, 1, …, n - 1)

[0012]

[0013] Step 2.3: Both the target edge feature extraction module and the saliency target feature extraction module are operated through the combination of three convolutions and activation functions. And starting from the lowest-layer saliency target feature extraction module, the data is passed to the upper layer to strengthen the feature information of the upper layer. And the feature information of the lowest layer is passed to the edge feature extraction module after one step of 2.2 to strengthen the feature information of this module;

[0014] Step 2.4: After each layer of the target edge feature extraction module and the saliency target feature extraction module goes through another combination operation of convolution and activation function to obtain the prediction image, use the line label samples to calculate the sum of the edge losses loss edge , where

[0015]

[0016] ​Step 2.5: After the target edge feature extraction module and the salient object feature extraction module repeat Step 2.2, they are connected to the edge fusion module for one-to-one fusion. After convolution, the predicted images of each layer are obtained, and the loss of semantic segmentation is calculated using the surface label. sal ;

[0017] Step 2.6: Each layer in the edge fusion module undergoes a combined operation of three convolutions plus an activation function to expand the field of view, and then undergoes another combined operation of convolution plus an activation function to obtain the predicted image. The features of each layer in Step 2.5 are fused, and the total loss is calculated.

[0018] loss = loss edge + loss sal (3)

[0019] Step 3: Use the samples in Step 1 to train the network model until it reaches the fitting state.

[0020] Step 4: Obtain the predicted image through the network model.

[0021] The present invention draws samples according to the contour of the target ground object, makes edge detection samples and semantic segmentation samples; builds a neural network model introducing edge learning and a spectrum-based channel attention mechanism, inputs high-resolution remote sensing images, edge detection samples, and semantic segmentation samples into training until the fitting state is reached, and obtains a salient object detection model for remote sensing images; inputs high-resolution remote sensing images to obtain the predicted image. The present invention combines edge learning and a spectrum-based channel attention mechanism, solves the problem of unclear edges in salient detection in the traditional remote sensing field, and the problem of inaccurate detection caused by the lack of spectral features when using deep learning methods.

[0022] The advantages of the present invention are:

[0023] (1) Edge learning is introduced into the network model, which strengthens the edge information of the target and fuses it with the target object to strengthen the target boundary.

[0024] (2) A spectrum-based channel attention mechanism is introduced into the network model to strengthen the channel information and the frequency information of the remote sensing image. Description of the Drawings

[0025] Figure 1 It is a schematic flowchart of the network model.

[0026] Figure 2 a~ Figure 2 c are experimental data graphs, Figure 2 a represents the original image sample graph of the high-resolution remote sensing image, Figure 2 b represents the surface label sample graph, Figure 2 c represents the line label sample graph.

[0027] Figure 3 It is a diagram of the FCA channel attention mechanism.

[0028] Figure 4 a~ Figure 4 b are experimental effect diagrams, Figure 4 a is a surface label sample diagram, Figure 4 b is a prediction result diagram. Specific implementation manners

[0029] Refer to Figure 1 As shown, it is a method for detecting salient objects in high-resolution remote sensing images by combining edge learning and frequency channel attention according to the present invention. The specific steps are as follows:

[0030] (1) Making high-resolution remote sensing image samples, including the following steps:

[0031] (1-1) Preparing high-resolution remote sensing images: Manually select salient regions, use ArcGIS software to manually annotate the edges of ground objects, strictly draw according to the real target boundaries, and assign semantic type attribute values to the ground objects. Among them, the salient regions are labeled as 1, and other regions are labeled as 0, obtaining the corresponding vector files and original images, and rasterizing the vector files into surface labels and line labels.

[0032] (1-2) Augmenting the data: Performing data augmentation operations such as mirroring, rotating, and cropping on the images obtained in step (1-1) to augment the surface label and line label samples. The original image is as shown in Figure 2 a, the surface label sample is as shown in Figure 2 b, and the line label sample is as shown in Figure 2 c;

[0033] (2) Building a high-resolution remote sensing image training model, including the following steps:

[0034] (2-1) Selecting ResNet50 as the backbone, where the first layer in ResNet50 is used for edge

[0035] feature extraction, and other layers are used for the salient object feature extraction module;

[0036] (2-2) Connecting the FCA attention mechanism after the convolution with dimensionality reduction. Its model is as shown in Figure 3 shown, assigning two-dimensional weights to the features of the image, which strengthens the channel information and the frequency information of the remote sensing image. Among them,

[0037]

[0038] (2-3) Start with the most bottom significant target feature extraction module. Each layer undergoes three operations of convolution plus activation function combination to expand the perception field. Among them, the number of input channels and output channels of the convolution is the same (n*n), and the relu function is used as the activation function. Then the data is passed to the upper layer to strengthen the feature information of the upper layer, and the feature information of the most bottom layer is passed to the edge feature extraction module after step 2.2 to strengthen the feature information of this module;

[0039] (2-4) After each layer of the target edge feature extraction module and the significant target feature extraction module undergoes another combination operation of convolution (n*1) plus activation function (relu), the predicted image is obtained. Calculate the sum of edge losses loss in each layer using the line label samples edge , where

[0040]

[0041] (2-5) After the target edge feature extraction module and the significant target feature extraction module repeat step (2-2), they are connected to the edge fusion module for one-to-one fusion, and the fusion method is linear addition. After convolution, the predicted images of each layer are obtained, and the loss of semantic segmentation is calculated using the surface label sal ;

[0042] (2-6) After each layer in the edge fusion module undergoes three operations of convolution plus activation function to expand the field of view, and then another combination operation of convolution (n*1) plus activation function (relu), the predicted image is obtained. The features of each layer in step 2.5 are fused, and the total loss is calculated

[0043] loss = loss edge + loss sal (3)

[0044] (3) Train this network model, including the following steps:

[0045] (3-1) Initialize the weights and set the hyperparameters of the model: The weights of the convolutional layers are all randomly initialized using truncated normal (σ = 0.01), and the biases are initialized to 0. The hyperparameters are set as follows: learning rate = 5e-5, weight decay = 0.0005, momentum = 0.9, the loss weight of each side output is equal to 1, and the Adam optimizer is used to update the network gradient;

[0046] (3-2) Set the data source to the prepared data source, train the network model to the fitting state, and obtain the final network model parameters;

[0047] (4) Set network parameters using the network model parameters obtained in step (3), and input test samples to obtain predicted images, as shown in Figure 4 shown in Figure 4 b, where

[0048] The content described in the embodiments of this specification is only an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms stated in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art according to the inventive concept of the present invention.

Claims

1. A method for detecting salient objects in high - resolution remote sensing images by combining frequency and edge learning, comprising the following steps: Step 1: Make high-resolution remote sensing image samples, manually draw the precise boundaries of the targets and label the semantic types, perform rasterization processing to generate corresponding line label samples and surface label samples, and use data augmentation methods to expand the samples; Step 2: Build a high-resolution remote sensing image saliency target detection network model; Step 2.1: Select a suitable CNN network as the backbone. The data after convolution in the lower layers is used for the target edge feature extraction module, and the convolution in other layers is used for the saliency target feature extraction module; Step 2.2: Connect the FCA attention mechanism after the convolution with dimensionality reduction to assign two-dimensional weights to the features of the image, which strengthens the channel information and the frequency information of the remote sensing image. Among them, Step 2.3: Both the target edge feature extraction module and the saliency target feature extraction module are operated through the combination of three convolutions and activation functions. Starting from the lowest-layer saliency target feature extraction module, the data is passed to the upper layer to strengthen the feature information of the upper layer. And the feature information of the lowest layer is passed to the edge feature extraction module after one operation of Step 2.2 to strengthen the feature information of this module; Step 2.4: After each layer of the target edge feature extraction module and the salient object feature extraction module undergoes a combined operation of convolution and activation function again, a predicted image is obtained, and the line label samples are used to calculate the sum of the edge losses loss in each layer edge , where where n is the number of CNN layers; Step 2.5: After the target edge feature extraction module and the salient object feature extraction module repeat Step 2.2, they are connected to the edge fusion module for one-to-one fusion. After convolution, the predicted images of each layer are obtained, and the loss of semantic segmentation is calculated using the surface labels. sal ; Step 2.6: Expand the field of view of each layer in the edge fusion module through the combination operation of three convolutions and activation functions, and then obtain the prediction image through another combination operation of convolution and activation function; fuse the features of each layer in Step 2.5 and calculate the total loss loss = loss edge + loss sal (3) Step 3: Use the samples in Step 1 to train the network model to the fitting state; Step 4: Obtain the prediction image through this network model.

2. The method for detecting salient objects in high - resolution remote sensing images by combining frequency and edge learning according to claim 1, characterized in that: Step 1 specifically includes: (1-1) Prepare high-resolution remote sensing images: Determine the salient regions, use ArcGIS software to manually annotate the edges of the ground objects, strictly follow the real target boundaries for drawing, and assign semantic type attribute values to the ground objects. Among them, the salient regions are labeled as 1, and other regions are labeled as 0, to obtain the corresponding vector file and the original image, and rasterize the vector file into surface labels and line labels; (1-2) Augment the data: Perform data augmentation operations of mirroring, rotating, and cropping on the image obtained in step (1-1) to expand the surface label and line label samples.

3. The method for detecting salient objects in high - resolution remote sensing images by combining frequency and edge learning according to claim 1, characterized in that: Step 3 specifically includes: (3-1) Initialize the weights and set the hyperparameters of the model: the weights of the convolutional layers are all randomly initialized using truncated normal with σ = 0.01, and the biases are initialized to 0; the hyperparameters are set as follows: learning rate = 5e -5 , weight decay = 0.0005, momentum = 0.9, the loss weight for each side output is equal to 1, and the Adam optimizer is used to update the network gradients; (3-2) Set the data source to the prepared data source, train the network model to the fitting state, and obtain the final network model parameters.

Citation Information

Patent Citations

  • A vehicle target identification method based on physical feature extraction and SVM

    CN109766899A

  • Semantic segmentation and edge fused high-resolution remote sensing target accurate detection method

    CN112084872A