Remote Sensing Image Shadow Detection Method Based on Fusion of Cross-Space and Channel Attention
By using a method of fusing cross space and channel attention in remote sensing image shadow detection, the problem of low accuracy of small-area shadow detection and false detection of non-hacked areas in the prior art is solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202211212081.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-09-30
AI Technical Summary
The existing remote sensing image shadow detection technology relies on a single segmentation threshold, is poorly robust, has low accuracy in shadow detection for small areas, and has the problem of false detection in non-shaded areas.
The remote sensing image shadow detection method that integrates cross space and channel attention is adopted. By constructing the cross space attention module and channel attention module, the horizontal and vertical direction information and channel characteristics of the remote sensing image are fused to improve the shadow feature extraction capability.
It improves the accuracy and robustness of shadow detection, improves the detection accuracy of shadows in small areas, and reduces false detection of non-shaded areas.
Smart Images

Figure CN115661638B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and further relates to a method for remote sensing image shadow detection based on the fusion of cross-space and channel attention in the field of image detection technology. The present invention can be used to detect the image parts containing shadows in remote sensing images. Background Art
[0002] High-resolution remote sensing images can capture the complex details and structures of land cover objects, such as building structures, road conditions, and vegetation textures. However, in reality, the occlusion of oblique light by tall objects causes the inevitable existence of shadows in remote sensing images. The existence of these shadows causes inconvenience to various computer vision tasks such as object detection, segmentation, and tracking of ground objects. The features of ground objects located in the shadows cannot be completely extracted, and local features are prone to damage, thus affecting the accuracy of object detection. Therefore, it is necessary to find a method to automatically detect and then remove shadows, which is beneficial to the detection of ground object information and is also very valuable for image editing, computational photography, and augmented reality.
[0003] China University of Geosciences proposed an unsupervised remote sensing image shadow detection method in its patent document "An Unsupervised Remote Sensing Image Shadow Detection Method" (Application No.: 202210638592.4, Application Publication No.: CN 115018791 A, Publication Date: September 6, 2022). The main steps implemented by this method are: (1) Convert the RGB color space of the remote sensing image to the HSI color space and perform Gaussian filtering for denoising; (2) Use the denoised image to form a sample set and obtain the optimal segmentation thresholds for each color channel of HSI; According to the optimal segmentation thresholds of each color channel of HSI, detect the image to be detected to obtain a preliminary shadow detection result; (3) Mark the preliminary shadow detection result to obtain the marked connected regions and calculate the areas of the connected regions; (4) Set a threshold for screening the areas of the connected regions and use morphological closing operations to fill the screened regions to obtain the final shadow region. The disadvantages of this method are: Since the detection accuracy of remote sensing images only depends on the selection of segmentation thresholds, and in engineering practice, the acquisition scenarios of remote sensing images are complex and the types of ground objects are diverse, a single segmentation threshold cannot be applied to multiple shadow situations. Therefore, the robustness of this method is poor, and it can only detect large areas of shadows in blocks, and the detection accuracy for small-area shadows is relatively low.
[0004] Han Hongyin et al. proposed a method for detecting shadows in remote sensing images based on the logarithmic shadow index in their published paper "Research on Shadow Detection and Compensation Technology for High-Resolution Optical Remote Sensing Images" (University of Chinese Academy of Sciences, doctoral thesis, 2021. DOI: 10.27522, publication date: March 15, 2021). The steps of this method are as follows: (1) Initially construct the initial shadow index; (2) Based on the initial shadow index, apply a natural logarithm operation to further enhance the distinguishability between shadows and non-shadows, and finally form the LSI; (3) Use the NVEM threshold method to binarize the LSI index image to finally obtain the shadow detection result image. The deficiencies of this method are as follows: When processing remote sensing images, this method needs to select appropriate thresholds for segmentation according to specific image scenes, and its generalization ability is poor. In actual remote sensing image processing applications, the extraction of shadow features is limited, and there is a problem of misdetection in non-shadow areas that are similar to shadows. Summary of the Invention
[0005] The purpose of the present invention is to address the above-mentioned deficiencies of the existing technologies, and provide a method for detecting shadows in remote sensing images based on the fusion of cross-space and channel attention, which is used to solve the problems of existing shadow detection technologies relying on a single segmentation threshold, poor robustness, low accuracy in detecting small-area shadows, limited extraction of shadow features, and misdetection in non-shadow areas.
[0006] The idea for achieving the purpose of the present invention is as follows: Aiming at the problems of difficult detection of small-area shadows and easy misjudgment of local dark areas as shadows in remote sensing images, the present invention constructs a cross-space attention module, which can fuse the horizontal and vertical information of each pixel point on the feature map of the remote sensing image into the current pixel point, so that each position point concentrates the features in the cross direction, enhancing the ability to extract the spatial information of the feature map, focusing on the context semantic information in the horizontal and vertical directions of a certain area, and making more accurate judgments for small areas and suspected shadow areas. Aiming at the problem of limited extraction of shadow features, considering the dark color feature of shadows and the existence of many multi-channel feature maps in the network, a channel attention module is designed specifically for extracting channel features. This module can assign weights to different channels, pay more attention to the channels that conform to the shadow color feature, thereby enhancing the ability to extract shadow features and reducing the missed detection and misdetection of shadow areas. After cascading these two attention modules, a sub-network that fuses cross-space and channel attention is formed, which is embedded in the decoding block part of the encoder-decoder network, and the ResNeXt50 is used as the backbone for the encoding part of the encoder-decoder network, thus completing the overall construction of the remote sensing image shadow detection network, enabling the network to have good feature extraction ability, being able to focus on the position and channel features of the shadow image simultaneously, thereby improving the accuracy and robustness of shadow detection, and improving the problems of difficult detection of small-area shadows and misdetection in non-shadow areas in the shadow detection results.
[0007] To achieve the above object, the technical solution of the present invention includes the following:
[0008] Step 1, generating a training set:
[0009] Step 1.1, selecting at least 500 remote sensing images, each of which contains at least 1 shadow area, and cropping each remote sensing image to a size of 512×512;
[0010] Step 1.2, annotating each shadow area in each remote sensing image, and generating a label image corresponding to each remote sensing image after annotation;
[0011] Step 1.3, forming all remote sensing images and their corresponding label images into a training set;
[0012] Step 2, constructing a remote sensing image shadow detection network:
[0013] Step 2.1, constructing a fusion cross-space and channel attention sub-network composed of a cascaded cross-space attention module and a channel attention module;
[0014] Step 2.1.1, constructing a cross-space attention module for extracting pixel features of the same row and the same column:
[0015] Build a cross-space attention module, where an output end of the input layer is sequentially connected to a feature extraction block, a multiplication layer, an output end of a normalization layer, a multiplication layer, a dot product layer, an addition layer, and an output layer; a transpose layer is connected between another output end of the normalization layer and the multiplication layer, and the multiplication layer is used to multiply the image after normalization processing by the image processed by the transpose layer; another output end of the input layer is connected to the dot product layer; the normalization layer is implemented by using the Softmax function;
[0016] The feature extraction block is composed of two branches in parallel with the same structure and different parameters. Each branch is composed of a cascaded convolutional layer and a feature extraction layer. The convolutional kernel size of the convolutional layer is set to 1×1. The feature extraction layers of the two branches use the pixel feature extraction formula to output the current position feature of the input picture pixel point and the pixel features of the same row and the same column in a set form;
[0017] The pixel feature extraction formula is as follows:
[0018]
[0019] Among them, Q i represents the feature of the i-th pixel extracted from the input image by the first branch, represents the one-dimensional feature of the i-th pixel extracted from the input image x by the first branch, Denote the feature of an image with the number of channels of the output of the first branch being C', the number of pixel rows being 1, and the number of pixel columns being 1, K i Denote the feature of the i-th pixel extracted by the second branch from the input image Denote the feature of the pixel in the same row and same column as the i-th pixel extracted from the input image x. H represents the total number of all pixel rows in the input image, and W represents the total number of all pixel columns in the input image Denote the feature of an image with the number of channels of the output of the second branch being C', the number of pixel rows being H + W - 1, and the number of pixel columns being 1
[0020] Step 2.1.2, construct a channel attention module for extracting channel features
[0021] Build a channel attention module. Among them, one output end of the input layer is successively connected to a dimension transformation block, a multiplication layer, a normalization layer, a first fully connected layer, a second fully connected layer, a dot product layer, and an output layer; the other output end of the input layer is connected to the dot product layer for weighting each channel of the input image; set the number of input nodes of the first fully connected layer to N, where the value of N is equal to the number of channels of the input image; set the number of output nodes of the first fully connected layer to N', and N' is equal to 1 / 16 of N; set the number of input nodes of the second fully connected layer to M, where the value of M is equal to N'; set the number of output nodes of the second fully connected layer to M', and the value of M' is equal to N. The normalization layer is implemented using the Softmax function
[0022] The dimension transformation block is composed of two branches in parallel with the same structure but different parameters. Each branch consists of a convolutional layer and a dimension transformation layer in cascade; set the convolutional kernel size of the convolutional layer to 1×1, the number of image channels after convolutional processing of the first branch to 1, and the number of image channels after convolutional processing of the second branch to C', where the value of C’ is equal to the number of channels of the input image; the dimension transformation layer is implemented using the reshape function
[0023] Step 2.1.3, cascade the cross-space attention module and the channel attention module, and connect the output layer of the channel attention module to the input layer of the cross-attention module to obtain a fusion cross-space and channel attention sub-network
[0024] Step 2.2, construct an encoder-decoder basic network
[0025] Construct an encoding and decoding basic network, the structure of which is in turn: input layer, the first scale reduction block, the second scale reduction block, the third scale reduction block, the fourth scale reduction block, the first scale restoration block, the second scale restoration block, the third scale restoration block, the fourth scale restoration block, output layer; among them, the structures of the first to fourth scale reduction blocks are the same, and its structure is composed of the first convolutional layer, the second convolutional layer and the downsampling layer connected in series. The structures of the first to fourth scale restoration blocks are the same, and its structure is composed of the first convolutional layer, the second convolutional layer and the upsampling layer connected in series. The output layer is composed of the first convolutional layer and the second convolutional layer in cascade;
[0026] Set the parameters of each layer of the encoding and decoding basic network as follows: Set the sizes of the convolutional kernels of the first to second convolutional layers in the first to fourth scale reduction blocks and the first to second convolutional layers of the convolutional layers in the first to fourth scale restoration blocks to 3×3. Set the downsampling kernel sizes of the downsampling layers in the first to fourth scale reduction blocks to 2×2. Set the upsampling kernel sizes of the upsampling layers in the first to fourth scale restoration blocks to 2×2;
[0027] Step 2.3, construct a remote sensing image shadow detection network:
[0028] Build a remote sensing image shadow detection network. Among them, connect 1 fusion cross-space and channel attention sub-network in parallel between the two convolutional layers of the fourth scale reduction block and the output of the first scale restoration block; connect 1 fusion cross-space and channel attention sub-network in parallel between the two convolutional layers of the third scale reduction block and the output of the second scale restoration block; connect 1 fusion cross-space and channel attention sub-network in parallel between the two convolutional layers of the second scale reduction block and the output of the third scale restoration block; connect 1 fusion cross-space and channel attention sub-network in parallel between the two convolutional layers of the first scale reduction block and the output of the fourth scale restoration block;
[0029] Step 3, train the remote sensing image shadow detection network:
[0030] Input the training set into the remote sensing image shadow detection network, and use the gradient descent method to iteratively update the parameters of the network until the loss function value of the network converges, and obtain a trained remote sensing image shadow detection network;
[0031] Step 4, detect the shadow in the remote sensing image:
[0032] After cropping the remote sensing image containing the shadow to be detected to a size of 512×512, input it into the trained remote sensing image shadow detection network, and output the shadow detection result.
[0033] Compared with the prior art, the present invention has the following advantages:
[0034] First, since the present invention constructs a cross - spatial attention module, which fuses the information in the cross direction of each pixel point on the remote sensing image feature map into the current pixel point, it can make more accurate judgments for small regions and suspected shadow regions, overcoming the problem of low detection accuracy of small - region shadows in the prior art, and thus improving the shadow detection accuracy of the present invention.
[0035] Second, since the present invention constructs a channel attention module, which pays more attention to the channels that conform to the shadow color characteristics, overcoming the problem of limited extraction of shadow features in the prior art, reducing the missed detection and false detection of shadow regions, and thus improving the shadow detection precision of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flowchart of the present invention;
[0037] Figure 2 is a schematic structural diagram of the cross - spatial attention module of the present invention;
[0038] Figure 3 is a schematic structural diagram of the channel attention module of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0039] The following describes the present invention in detail with reference to the accompanying drawings and specific embodiments.
[0040] Refer to Figure 1 and the embodiments to further describe the specific implementation steps of the present invention.
[0041] Step 1: Generate a training set and a test set.
[0042] Step 1.1: Select 500 remote sensing images from the high - resolution remote sensing image dataset open - sourced by the International Society for Photogrammetry and Remote Sensing. The image format is.tif, and the size of each remote sensing image is 512×512, and each remote sensing image contains at least 1 shadow region.
[0043] Step 1.2: Label each shadow region in each remote sensing image and generate a corresponding label image for each remote sensing image after labeling. The format of the label image is.tif.
[0044] Step 1.3: After pairing all the remote sensing images and their corresponding label images, randomly divide them. 80% is used as the training set and 20% is used as the test set.
[0045] Step 2: Construct a remote sensing image shadow detection network.
[0046] Step 2.1: Construct a fusion cross - spatial and channel attention sub - network composed of a cascaded cross - spatial attention module and a channel attention module.
[0047] Step 2.1.1: Construct a cross - spatial attention module for extracting pixel features in the same row and column.
[0048] Refer to Figure 2 for a further description of the cross - spatial attention module constructed in the present invention.
[0049] Build a cross - spatial attention module, where an output end of the input layer is successively connected to a feature extraction block, a multiplication layer, an output end of a normalization layer, a multiplication layer, a dot - product layer, an addition layer, and an output layer; a transpose layer is connected between the other output end of the normalization layer and the multiplication layer, and the multiplication layer is used to multiply the normalized image by the image processed by the transpose layer. The other output end of the input layer is connected to the dot - product layer. The normalization layer is implemented using the Softmax function.
[0050] The feature extraction block is composed of two branches in parallel with the same structure but different parameters. Each branch consists of a convolutional layer and a feature extraction layer. The convolutional kernel size of the convolutional layer is set to 1×1, and the number of channels of the image after convolutional processing is C'. The feature extraction layers of the two branches are respectively implemented by their respective pixel feature extraction formulas, which are used to extract the current position features, same - row and same - column pixel features of the pixel points of the input picture, and respectively output feature sets: Q = {Q1, Q2, …, Q i}, K = {K1, K2, …, K i}. Among them, Q represents the feature set obtained by the feature extraction layer of the first branch, K represents the feature set obtained by the feature extraction layer of the second branch, and i = 1, 2, …, H×W represents the pixel spatial position index.
[0051] The pixel feature extraction formulas are as follows:
[0052]
[0053] Among them, Q i represents the feature of the i - th pixel extracted by the first branch from the input image, represents the one - dimensional feature of the i - th pixel extracted by the first branch from the input image x, represents the feature of the image with the number of channels C', pixel rows 1, and pixel columns 1 output by the first branch, K i represents the feature of the i - th pixel extracted by the second branch from the input image, represents the feature of the pixels in the same row and column as the i - th pixel extracted from the input image x, H represents the total number of all pixel rows in the input image, W represents the total number of all pixel columns in the input image, represents the feature of the image with the number of channels C', pixel rows H + W - 1, and pixel columns 1 output by the second branch.
[0054] Step 2.1.2: Construct a channel attention module for extracting channel features.
[0055] Refer to Figure 3 for a further description of the channel attention module constructed in the present invention.
[0056] Build a channel attention module, where an output terminal of the input layer is successively connected to a dimension transformation block, a multiplication layer, a normalization layer, a first fully connected layer, a second fully connected layer, a dot product layer, and an output layer; another output terminal of the input layer is connected to the dot product layer for weighting each channel of the input image. Set the number of input nodes of the first fully connected layer to N, where the value of N is equal to the number of channels of the input image. Set the number of output nodes of the first fully connected layer to N', where N' is equal to 1 / 16 of N. Set the number of input nodes of the second fully connected layer to M, where the value of M is equal to N'. Set the number of output nodes of the second fully connected layer to M', where the value of M' is equal to N, and the normalization layer is implemented using the Softmax function.
[0057] The dimension transformation block is composed of two parallel branches with the same structure but different parameters. Each branch consists of a convolutional layer and a dimension transformation layer. The convolutional kernel size of the convolutional layer is set to 1×1. The number of image channels after convolution processing in the first branch is set to 1, and the number of image channels after convolution processing in the second branch is set to C', where the value of C' is equal to the number of channels of the input image. The dimension transformation layer is implemented using the reshape function, and the call format of this function is B = reshape(A, m, n, p), where A is the input matrix, B is the output matrix, m is the total number of rows of the output matrix, n is the total number of columns of the output matrix, and p is the dimension of the output matrix. The total number of rows of the image output by the dimension transformation layer of the first branch is H×W, the number of columns is 1, and the dimension is 1. The number of rows of the image output by the dimension transformation layer of the second branch is 1, the total number of columns is H×W, and the dimension is C'.
[0058] Step 2.1.3: Cascade the cross-space attention module and the channel attention module, and connect the output layer of the channel attention module to the input layer of the cross-attention module to obtain a fusion cross-space and channel attention sub-network.
[0059] Step 2.2: Construct an encoder-decoder basic network:
[0060] Construct an encoding and decoding basic network, whose structure is in sequence: input layer, the first scale reduction block, the second scale reduction block, the third scale reduction block, the fourth scale reduction block, the first scale restoration block, the second scale restoration block, the third scale restoration block, the fourth scale restoration block, output layer; wherein, the structures of the first to fourth scale reduction blocks are the same, and its structure is composed of a first convolutional layer, a second convolutional layer and a downsampling layer connected in series, the structures of the first to fourth scale restoration blocks are the same, and its structure is composed of a first convolutional layer, a second convolutional layer and an upsampling layer connected in series, and the output layer is composed of a first convolutional layer and a second convolutional layer.
[0061] Set the parameters of each layer of the encoding and decoding basic network as follows: set the sizes of the convolutional kernels of the first to second convolutional layers in the first to fourth scale reduction blocks and the first to second convolutional layers of the convolutional layers in the first to fourth scale restoration blocks to be 3×3, set the sizes of the downsampling kernels of the downsampling layers in the first to fourth scale reduction blocks to be 2×2, and set the sizes of the upsampling kernels of the upsampling layers in the first to fourth scale restoration blocks to be 2×2;
[0062] Step 2.3, construct a remote sensing image shadow detection network:
[0063] Embed 4 fusion cross-space and channel attention sub-networks with the same structure into the encoding and decoding basic network, and the module embedding connection method is as follows: connect 1 fusion cross-space and channel attention sub-network in parallel between the two convolutional layers of the fourth scale reduction block and the output of the first scale restoration block; connect 1 fusion cross-space and channel attention sub-network in parallel between the two convolutional layers of the third scale reduction block and the output of the second scale restoration block; connect 1 fusion cross-space and channel attention sub-network in parallel between the two convolutional layers of the second scale reduction block and the output of the third scale restoration block; connect 1 fusion cross-space and channel attention sub-network in parallel between the two convolutional layers of the first scale reduction block and the output of the fourth scale restoration block.
[0064] Step 3, train the remote sensing image shadow detection network.
[0065] Input the training set into the remote sensing image shadow detection network, and use the gradient descent method to iteratively update the parameters of the network until the value of the loss function of the network converges, and obtain the trained remote sensing image shadow detection network.
[0066] The loss function is as follows:
[0067]
[0068] Wherein, L represents the loss function, ∑ represents the summation operation, y i represents the true label corresponding to the i-th remote sensing image in the training set input to the remote sensing image shadow detection network, log(·) represents the logarithmic operation with base 2, It represents the predicted label of the i-th remote sensing image in the training set input to the remote sensing image shadow detection network.
[0069] Step 4: Detect remote sensing images with shadows.
[0070] Input the remote sensing image with shadows to be detected into the trained remote sensing image shadow detection network, and output the shadow detection result.
[0071] The following further illustrates the effect of the present invention in combination with simulation experiments:
[0072] 1. Simulation conditions:
[0073] The hardware test platform for the simulation experiment of the present invention is: the CPU is Intel(R) Xeon(R) Sliver 4110, the memory is 64GB, and the graphics card is NVIDIA RTX2080Ti.
[0074] The software platform for the simulation experiment of the present invention: a code running environment of the Pytorch 1.9.1 deep learning framework built in the Python 3.6.12 virtual environment of Anaconda.
[0075] The remote sensing image data used in the simulation experiment of the present invention is the AISD dataset released by the School of Resource and Environmental Sciences of Wuhan University in 2020. This dataset contains a total of 514 remote sensing images with shadows and their corresponding shadow labels, among which 463 are the sample set and 51 are the test set, and the image format is all.tif. The images and labels in the sample set are sequentially and simultaneously subjected to cropping, rotation, and mirror transformation methods to achieve data augmentation, resulting in remote sensing images with a size of 256×256×3 for each image, a total of 5019 remote sensing images. All the augmented images and their labels are combined to form a training set.
[0076] 2. Simulation content and result analysis:
[0077] The simulation experiment of the present invention is to train the training set on the remote sensing image shadow detection network constructed by the present invention and the networks constructed by two existing technologies (DSSDNet network, GSCA-UNet network) respectively, to obtain three trained networks, and then input the test set into each trained network for shadow detection, and output the results of the three networks for shadow detection of remote sensing images.
[0078] The prior art method DSSDNet network refers to the deeply supervised convolutional neural network for shadow detection proposed by Shuang Luo et al. in their published paper "Deeply supervised convolutional neural network for shadow detection based on a novel aerial shadow imagery dataset" (ISPRS Journal of Photogrammetry and Remote Sensing, 2020, 167: 443-457.), abbreviated as DSSDNet network.
[0079] The prior art method GSCA-UNet network refers to the automatic shadow detection network in urban aerial imagery based on the global-spatial-context attention module proposed by Yuwei Jin et al. in their published paper "GSCA-UNet: Towards Automatic Shadow Detection in Urban Aerial Imagery with Global-Spatial-Context Attention Module" (Remote Sens. 2020, 12, 2864), abbreviated as GSCA-UNet network.
[0080] To evaluate the shadow detection effects of the three networks, the F1-score of each method in the simulation results of the present invention is calculated through the following formula. The F1-score is the harmonic mean of the accuracy and recall rate of shadow detection. The calculation results of the three simulation methods are plotted in Table 1, where Ours in Table 1 represents the simulation experimental results using the method of the present invention.
[0081]
[0082]
[0083]
[0084] Table 1 Comparison of objective indicators between the present invention and other prior algorithms
[0085] Evaluation metrics DSSDNet GSCA-UNet Ours F1-score 91.79% 91.69% 92.60%
[0086] As can be seen from Table 1, the F1-score of the present invention is 92.60%, and the index is significantly higher than the two prior art methods, proving that the present invention can obtain higher accuracy in remote sensing image shadow detection.
Claims
1. A remote sensing image shadow detection method based on the fusion of cross-space and channel attention, characterized in that, Construct a fusion cross - space and channel attention sub - network composed of a cascaded cross - space attention module and a channel attention module; the steps of this detection method are as follows: Step 1, generate a training set: Step 1.1, select at least 500 remote sensing images, each of which contains at least 1 shadow area, and crop each remote sensing image to a size of 512×512; Step 1.2, annotate each shadow area in each remote sensing image, and generate a label image corresponding to each remote sensing image after annotation; Step 1.3, form a training set with all remote sensing images and their corresponding label images; Step 2, construct a remote sensing image shadow detection network: Step 2.1, construct a fusion cross - space and channel attention sub - network composed of a cascaded cross - space attention module and a channel attention module; Step 2.1.1, construct a cross - space attention module for extracting pixel features of the same row and the same column: Build a cross - space attention module, in which an output end of the input layer is sequentially connected to a feature extraction block, a multiplication layer, an output end of a normalization layer, a multiplication layer, a dot - product layer, an addition layer, and an output layer; a transpose layer is connected between another output end of the normalization layer and the multiplication layer, and the multiplication layer is used to multiply the normalized image by the image processed by the transpose layer; another output end of the input layer is connected to the dot - product layer; the normalization layer is implemented using the Softmax function; The feature extraction block is composed of two parallel branches with the same structure but different parameters. Each branch is composed of a cascaded convolutional layer and a feature extraction layer. The convolutional kernel size of the convolutional layer is set to 1×1. The feature extraction layers of the two branches use the pixel feature extraction formula to output the current position feature and the pixel features of the same row and the same column of the input picture pixels in a set form; The pixel feature extraction formula is as follows: Among them, Q i represents the feature of the i-th pixel extracted by the first branch from the input image, represents the one-dimensional feature of the i-th pixel extracted by the first branch from the input image x, represents the feature of an image with the number of channels C', the number of pixel rows 1, and the number of pixel columns 1 output by the first branch, K i represents the feature of the i-th pixel extracted by the second branch from the input image, represents the feature of the pixel in the same row and the same column as the i-th pixel extracted from the input image x. H represents the total number of all pixel rows in the input image, and W represents the total number of all pixel columns in the input image, represents the feature of an image with the number of channels C', the number of pixel rows H + W - 1, and the number of pixel columns 1 output by the second branch; Step 2.1.2, construct a channel attention module for extracting channel features: Build a channel attention module, in which an output end of the input layer is sequentially connected to a dimension transformation block, a multiplication layer, a normalization layer, a first fully - connected layer, a second fully - connected layer, a dot - product layer, and an output layer; another output end of the input layer is connected to the dot - product layer for weighting each channel of the input image; set the number of input nodes of the first fully - connected layer to N, where the value of N is equal to the number of channels of the input image; set the number of output nodes of the first fully - connected layer to N', and N' is equal to 1 / 16 of N; set the number of input nodes of the second fully - connected layer to M, where the value of M is equal to N'; set the number of output nodes of the second fully - connected layer to M', and M' is equal to N. The normalization layer is implemented using the Softmax function; The dimension transformation block is composed of two parallel branches with the same structure but different parameters. Each branch is composed of a cascaded convolutional layer and a dimension transformation layer; set the convolutional kernel size of the convolutional layer to 1×1, set the number of channels of the image processed by the convolution of the first branch to 1, and set the number of channels of the image processed by the convolution of the second branch to C', where the value of C' is equal to the number of channels of the input image; the dimension transformation layer is implemented using the reshape function; Step 2.1.3: Cascade the cross-space attention module and the channel attention module, connect the output layer of the channel attention module to the input layer of the cross-attention module, and obtain a sub-network that fuses cross-space and channel attention. Step 2.2: Construct an encoder-decoder basic network: Construct an encoder-decoder basic network with the following structure in sequence: input layer, the first scale reduction block, the second scale reduction block, the third scale reduction block, the fourth scale reduction block, the first scale restoration block, the second scale restoration block, the third scale restoration block, the fourth scale restoration block, output layer; among them, the structures of the first to fourth scale reduction blocks are the same, and their structures are composed of the first convolutional layer, the second convolutional layer, and the downsampling layer in series. The structures of the first to fourth scale restoration blocks are the same, and their structures are composed of the first convolutional layer, the second convolutional layer, and the upsampling layer in series. The output layer is composed of the first convolutional layer and the second convolutional layer in cascade. Set the parameters of each layer of the encoder-decoder basic network as follows: Set the kernel sizes of the first to second convolutional layers in the first to fourth scale reduction blocks and the first to second convolutional layers in the convolutional layers of the first to fourth scale restoration blocks to 3×3. Set the downsampling kernel sizes of the downsampling layers in the first to fourth scale reduction blocks to 2×2. Set the upsampling kernel sizes of the upsampling layers in the first to fourth scale restoration blocks to 2×2. Step 2.3: Construct a remote sensing image shadow detection network: Build a remote sensing image shadow detection network. Among them, connect a sub-network that fuses cross-space and channel attention in parallel between the two convolutional layers of the fourth scale reduction block and the output of the first scale restoration block; connect a sub-network that fuses cross-space and channel attention in parallel between the two convolutional layers of the third scale reduction block and the output of the second scale restoration block; connect a sub-network that fuses cross-space and channel attention in parallel between the two convolutional layers of the second scale reduction block and the output of the third scale restoration block; connect a sub-network that fuses cross-space and channel attention in parallel between the two convolutional layers of the first scale reduction block and the output of the fourth scale restoration block. Step 3: Train the remote sensing image shadow detection network: Input the training set into the remote sensing image shadow detection network, and use the gradient descent method to iteratively update the parameters of the network until the loss function value of the network converges, and obtain a trained remote sensing image shadow detection network. Step 4: Detect shadows in remote sensing images: After cropping the remote sensing image containing shadows to be detected to a size of 512×512, input it into the trained remote sensing image shadow detection network, and output the shadow detection result.
2. The remote sensing image shadow detection method based on fused cross-space and channel attention according to claim 1, wherein, The reshape(·) function described in Step 2.1.2 is B = reshape(A, m, n, p), where B is the output matrix, reshape(·) is a function of H×W, A is the input matrix, m is the total number of rows of the output matrix, n is the total number of columns of the output matrix, and p is the dimension of the output matrix; the total number of rows of the image output by the first branch dimension transformation layer is H×W, the number of columns is 1, and the dimension is 1; the number of rows of the image output by the second branch dimension transformation layer is 1, the total number of columns is H×W, and the dimension is C'.
3. The remote sensing image shadow detection method based on the fusion of cross-space and channel attention according to claim 1, wherein The loss function described in Step 3 is as follows: Among them, L represents the loss function, ∑ represents the summation operation, and y i represents the true label corresponding to the i-th remote sensing image in the training set input to the remote sensing image shadow detection network, log(·) represents the logarithmic operation with base 2, represents the predicted label of the i-th remote sensing image in the training set input to the remote sensing image shadow detection network.
Citation Information
Patent Citations
Unsupervised remote sensing image shadow detection method
CN115018791A