Aquatic vegetation classification method and device based on remote sensing image and storage medium
Through the deep convolutional neural network CAD-Unet model, combining channel and spatial attention mechanism and multi-scale hollow pooling, the problem of multi-scale information fusion in remote sensing classification of aquatic vegetation is solved, and classification accuracy and stability are improved, especially the recognition ability in complex water environments and sample imbalance.
Patent Information
- Application Number
- CN202510713138.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional methods are difficult to take into account multi-scale spatial details and large-scale context information in the remote sensing classification of aquatic vegetation. Especially when the boundaries between water and vegetation are blurred and sample categories are uneven, it is difficult to fully explore multi-scale context and focus on discriminant information of a small number of samples.
Deep convolutional neural network (CAD-Unet model) is adopted, through the combination of encoder module, bottleneck module and decoder module, it combines channel and spatial attention mechanism (CBAM) and multi-scale hollow space pyramid pooling (ASPP), and introduces Dropout layer to perform end-to-end training and dynamic learning rate scheduling to improve the model's generalization ability of complex backgrounds and few sample categories.
The accuracy and stability of aquatic vegetation classification is improved, especially under complex water environments and uneven sample sizes, and the ability to express small goals and complex backgrounds is enhanced, ensuring the fine recovery of boundary details and the continuity of classification results.
Smart Images

Figure CN120259889A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of remote sensing image classification, and particularly relates to a method for classifying aquatic vegetation based on remote sensing images. Background Art
[0002] The remote sensing classification of aquatic vegetation involves both complex spectral information and has to cope with the diversity of vegetation morphology and spatial distribution. Traditional methods usually rely on features such as texture, spectrum or vegetation index, and it is difficult to take into account multi-scale spatial details and large-scale context information. In recent years, the end-to-end learning framework based on deep convolutional neural network (CNN) has become the mainstream, which can automatically extract hierarchical features from large-scale data and reduces the dependence on expert experience for feature design. The U-Net structure is widely used due to its symmetric encoding-decoding path and skip connections, and this structure effectively integrates shallow high-resolution features and deep semantic features. However, when the standard U-Net model deals with the blurred boundary between water and vegetation and the imbalance of sample categories, it is often difficult to fully mine multi-scale context and focus on the discriminant information of a small number of samples. Summary of the Invention
[0003] To solve the above technical problems, the present invention provides a method, device and storage medium for classifying aquatic vegetation based on remote sensing images.
[0004] To achieve the above object, the present invention specifically adopts the following technical solutions: The present invention first provides a method for classifying aquatic vegetation based on remote sensing images, including the following steps: A method for classifying aquatic vegetation based on remote sensing images, characterized by including the following steps: Step S1, constructing a deep convolutional neural network, which is composed of an encoder module, a bottleneck module and a decoder module; The encoder module adopts multi-layer convolution and batch normalization operations, and embeds channel and spatial attention mechanisms in each part to realize the gradual extraction of image features and spatial downsampling; the encoder module has four parts, and the basic structure of each part is a convolutional layer, a secondary convolutional layer, a channel and spatial attention mechanism layer and a pooling layer in sequence. The fourth part adds a Dropout layer after the channel and spatial attention mechanism layer, and then completes the last-level spatial downsampling through the pooling layer; The bottleneck module uses the atrous spatial pyramid pooling structure to parallelly collect multi-scale context information; the structure of the bottleneck module is a convolutional layer, an atrous spatial pyramid pooling layer and a Dropout layer in sequence; The decoder module performs skip connections on the feature maps of the corresponding encoder layers through upsampling and feature fusion, gradually restoring spatial details, and finally generating a multi-class classification prediction map. The decoder module consists of four parts. The basic structure of each part is, in sequence, an upsampling layer, a skip connection splicing layer, a first convolutional layer, and a second convolutional layer. After the second convolution in the fourth part, a multi-class classification prediction map is generated through an output convolutional layer. The inputs of the skip connection splicing layers of each part are, in sequence, the output of the upsampling layer of the first part of the decoder module and the output of the fourth part of the encoder module; the output of the upsampling layer of the second part and the output of the third part of the encoder module; the output of the upsampling layer of the third part and the output of the second part of the encoder module; the output of the upsampling layer of the fourth part and the output of the first part of the encoder module. Step S2: Create an aquatic vegetation classification dataset, which is divided into a training set, a validation set, and a test set. Step S3: Use the dataset to train the constructed deep convolutional neural network. Step S4: Input the remote sensing image to be classified into the trained deep convolutional neural network, and output the classification result of the aquatic vegetation in the image.
[0005] In step S1, the structure of the deep convolutional neural network consists of an encoder module, a bottleneck module, and a decoder module. The specific implementation of each module is as follows: Encoder module: It consists of four parts. The basic structure of each part is, in sequence, a first convolutional layer → a second convolutional layer → a channel and spatial attention mechanism (CBAM) layer → a pooling layer. In the fourth part, a Dropout layer (dropout rate of 0.5) is added after the CBAM layer, and then the last-level spatial downsampling is completed through the pooling layer. The number of convolutional kernels in each convolutional layer is set to 64, 128, 256, and 512 respectively. The size of the convolutional kernels is set to 3×3, the stride is set to 1, and batch normalization and ReLU activation processing are performed after each convolution operation. All pooling layers use max pooling operations. The size of the pooling region kernel is set to 2×2, and the stride is set to 2.
[0006] Bottleneck module: Its structure is, in sequence, a convolutional layer → an atrous spatial pyramid pooling (ASPP) layer → a Dropout layer. The size of the convolutional kernel of the convolutional layer is set to 3×3, the stride is set to 1, the number of convolutional kernels is set to 1024, and batch normalization and ReLU activation processing are performed. The ASPP layer is set with five different dilation rates of (1, 6, 12, 18, 24), and the number of channels of the output feature map is set to 128. The dropout rate of the Dropout layer is set to 0.5.
[0007] Decoder module: It consists of four parts. The basic structure of each part is, in sequence, an upsampling layer → a skip connection splicing layer → a first convolutional layer → a second convolutional layer. After the second convolution in the fourth part, a multi-class classification prediction map is generated through an output convolutional layer. The upsampling layer uses 2x bilinear interpolation. After upsampling, a convolution operation is performed. The size of the convolution kernel is set to 2×2, and the stride is set to 1. The inputs to the channel splicing layers of each part are, in sequence, the output of the first upsampling layer of the decoder module and the output of the fourth part of the encoder module; the output of the second upsampling layer and the output of the third part of the encoder module; the output of the third upsampling layer and the output of the second part of the encoder module; the output of the fourth upsampling layer and the output of the first part of the encoder module. The number of convolution kernels in the first and second convolutional layers of the four parts of the decoder module is set to 512, 256, 128, and 64 in sequence. The size of the convolution kernels is all set to 3×3, and the stride is all set to 1. After the convolution operation, batch normalization processing and ReLU activation processing are performed. The size of the convolution kernel of the output convolutional layer is set to 1×1, the stride is set to 1, and Softmax activation is used.
[0008] In step S1, the CBAM layer structure includes a channel attention sub-module and a spatial attention sub-module. The channel attention sub-module performs global average pooling and global max pooling on the input feature map respectively to obtain vectors with a shape of C dimensions; the results of the two pooling operations are respectively input into a shared two-layer fully connected network. The number of nodes in the first layer is C / r, where r is set to 0.5, and ReLU activation processing is performed. The number of nodes in the second layer is C, and linear activation is performed; the outputs of the two fully connected layers are added together, and after Sigmoid activation, a channel attention coefficient with a shape of 1×1×C is obtained; the channel attention coefficient is multiplied by the input feature map channel by channel, and the output channel-refined feature map has a size of H×W×C. The calculation method of the channel attention sub-module is: In the formula, M C ( x ) is the channel attention weight; σ is Sigmoid function; ReLU ( ) is the ReLU activation function; W 1. W 2 is the weight of the fully connected layer; Avgpool is the average pooling operation; Avgpool is the average pooling operation; Maxpool is the max pooling operation.
[0009] The spatial attention sub-module performs average pooling and max pooling on the channel-refined feature map along the channel dimension respectively to obtain an average feature map and a max feature map, both with the size of H×W×1; after concatenating the above two feature maps along the channel dimension, a two-dimensional attention feature map with the size of H×W×2 is formed; the feature map is input into a convolution with a kernel size of 7×7 and a stride of 1, and after Sigmoid activation, a spatial attention coefficient with the shape of H×W×1 is generated; the spatial attention coefficient is multiplied element-wise with the channel-refined feature map, and the final fused feature map of CBAM is output, with the size of H×W×C. The calculation method is as follows: In the formula, M S ( x ) is the spatial attention weight, f 7×7 is the 7×7 convolution kernel; Concat ( ) is the concatenation operation.
[0010] In the encoder module, the CBAM layer is embedded in four parts respectively. Assuming the size of the original image is H×W×C o , the sizes of the corresponding feature maps of different parts are H×W, (H / 2)×(W / 2), (H / 4)×(W / 4) and (H / 8)×(W / 8) in sequence, and the corresponding number of channels C are 64, 128, 256 and 512 in sequence.
[0011] In step S1, the ASPP layer structure is five dilated convolution branches → global average pooling branch → feature concatenation → fused convolution in sequence. The number of channels of the feature map output by this module is set to 128. Five groups of different dilation rates (1, 6, 12, 18, 24) are set for the dilated convolution branches, and five dilated convolution operations are performed on the input feature map of the convolution branches in parallel. The size of the convolution kernel is set to 3×3 for all, and the stride is set to 1 for all. After convolution, batch normalization processing and ReLU activation processing are carried out immediately, and the output shape is (H / 16)×(W / 16)×128. The calculation formula is as follows: In the formula, H out and W out are the height and width of the output image; floo r( ) is the floor operation, which is used to ensure that the output size is an integer; H in and W in are the height and width of the input image; padding [0] and padding [1] are the number of pixels filled in the height and width directions; dilation [0] anddilation [1] is the dilation rate of the dilated convolution in the height and width directions; kernel _ size [0] and kernel _ size [1] are the sizes of the convolutional kernel in the height and width directions, both being 3 in a 3×3 convolutional kernel; stride [0] and stride [1] are the strides of the convolution operation in the height and width directions.
[0012] The global average pooling branch performs global pooling on the entire pooled branch input feature map of size (H / 16)×(W / 16)×1024 to obtain a 1×1×1024 vector; then through a convolution operation, the size of the convolutional kernel is set to 1×1, the stride is set to 1, and after convolution, batch normalization processing and ReLU activation are performed to extract global semantic features; finally, bilinear interpolation is used to upsample it to (H / 16)×(W / 16)×128. Feature concatenation concatenates the outputs of the above five dilated convolution branches and the global average pooling branch in the channel dimension to generate a multi-scale fusion feature map of size (H / 16)×(W / 16)×768. The fusion convolution performs a convolution operation on the concatenated feature map, the size of the convolutional kernel is set to 1×1, the stride is set to 1, and after convolution, batch normalization processing and ReLU activation are performed to achieve the final channel compression and multi-scale context information fusion, and the image size is (H / 16)×(W / 16)×128. The calculation formula of this module is: In the formula, Conv 1×1 ( ) is the 1×1 convolution operation; f i (x) is the pixel value of the output feature image of the th dilated convolution; k is the number of dilated convolutions; f gap (x) is the pixel value of the output feature image of the global average pooling branch; Conv 3×3,ri are different dilation rates r i of the 3×3 dilated convolution operation; r i is the dilation rate of the th dilated convolution.
[0013] In step S1, the calculation formula of the batch normalization processing of the CAD-Unet deep convolutional neural network is as follows: In the formula, is the mean within the batch; xi is the activation value of the th sample in the batch; m is the number of samples in the batch; is the variance within the batch; is the normalized activation value; ε is a very small positive number used to prevent the denominator from being zero.
[0014] In step S2, when making the aquatic vegetation classification dataset, the pixel bit depths of the remote sensing image and its corresponding classification label are both set to 8 bit. A sliding window of 256×256 is used to crop the image and label, and data augmentation operations such as horizontal flipping, vertical flipping, and diagonal mirroring are performed on them. Then, for the images and labels with fewer samples in the category, secondary augmentation operations such as random 90-degree rotation, random brightness, Gaussian noise, elastic transformation, translation scaling, and rotation are performed. The training set, validation set, and test set are divided according to the ratio of 8:1:1. The training set is normalized by the maximum value, and at the same time, a color dictionary is obtained to perform one-hot encoding on the label.
[0015] In step S3, the training set and validation set data are input into the CAD-Unet deep convolutional neural network. Through end-to-end backpropagation, the Adam optimization algorithm is used to perform gradient descent to update the network weights. The loss function uses the cross-entropy function, and a learning rate scheduling strategy is combined, that is, when the performance index of the validation set has not improved within the preset number of iteration rounds, the learning rate is reduced proportionally until the convergence condition is met, and finally, a trained CAD-Unet model is obtained.
[0016] The trained CAD-Unet model is evaluated for accuracy using the test set. The test set data is normalized and input into the model for forward inference to obtain the class probability distribution of each pixel; then, the argmax operation is applied to the model output to generate a discrete class index map, and the index is mapped to a visual label map using a predefined color mapping dictionary, and after restoring to the original resolution through nearest neighbor interpolation, it is saved as a predicted label image. Select accuracy verification metrics to systematically evaluate the classification performance of the model based on the ground truth labels of the test set.
[0017] In step S4, for large-scale prediction of remote sensing images, a sliding cropping strategy with a fixed-size sub-window is used to perform edge-ignoring block division and normalization operations on the remote sensing image to be classified. The size of the sub-window is set to 256×256, and each sub-window is input into the trained CAD-Unet model in parallel for inference, and the classification probability maps of each sub-window are fused and recombined based on an overlap compensation strategy to generate a complete multi-class aquatic vegetation classification result.
[0018] The present invention also provides an aquatic vegetation classification device based on remote sensing images, which includes a processor and a memory; programs or instructions are stored in the memory, and the programs or instructions are loaded and executed by the processor to implement the steps of the aquatic vegetation classification method based on remote sensing images.
[0019] The present invention also provides a computer-readable storage medium, on which programs or instructions are stored, and when the programs or instructions are executed by a processor, the steps of the aquatic vegetation classification method based on remote sensing images are implemented.
[0020] Advantages of the present invention: The aquatic vegetation classification method based on remote sensing images proposed by the present invention integrates the channel-spatial attention mechanism (CBAM) and multi-scale atrous spatial pyramid pooling (ASPP) into the model. While focusing on key information regions, it extracts context information of different receptive fields in parallel through multiple atrous convolutions, enhancing the model's expression ability for small targets and complex backgrounds; Dropout layers are introduced at the end of the encoder and in the bottleneck module to effectively suppress overfitting and improve the model's generalization ability for few-shot categories; The use of a symmetric encoder-decoder structure and skip connections effectively fuses shallow spatial details and deep semantic features, ensuring fine recovery of boundary details; During the training process, end-to-end backpropagation is combined with dynamic learning rate scheduling, which helps to alleviate the overfitting phenomenon of few-shot categories and improve the recognition accuracy of the model for different aquatic vegetation categories; In the prediction stage, an overlapping sliding window and fusion recombination strategy are adopted to eliminate edge artifacts in large-scale remote sensing images and ensure the continuity and consistency of classification results. The method of the present invention shows high classification accuracy and stability under complex water body environments and unbalanced sample sizes. Description of the Drawings
[0021] Figure 1 is a schematic flow chart of the aquatic vegetation classification method based on remote sensing images in the present invention; Figure 2 is a schematic structural diagram of the CAD-Unet model in the present invention; Figure 3 is a schematic structural diagram of the encoder module of the CAD-Unet model in the present invention; Figure 4 is a schematic structural diagram of the decoder module of the CAD-Unet model in the present invention; Figure 5 is a schematic structural diagram of the CBAM layer in the present invention; Figure 6 is a schematic structural diagram of the ASPP layer in the present invention; Figure 7These are the predicted images and label images of the CAD-Unet and U-Net models in the present invention. (a1-a3) are the results of the U-Net model, (b1-b3) are the label images, and (c1-c3) are the results of the CAD-Unet model. Detailed implementation manners
[0022] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0023] It should be understood that the specific implementation manners described herein are only used to explain the present invention and are not used to limit the present invention.
[0024] Embodiment 1 As Figure 1 shown, this embodiment provides a method for classifying aquatic vegetation based on remote sensing images, including the following steps: Step S1, construct a deep convolutional neural network (CAD-Unet model, where CAD respectively means: C-CBAM module, A-ASPP module, D-Dropout module); The deep convolutional neural network consists of an encoder module, a bottleneck module, and a decoder module. For details, see Figure 2 .
[0025] The encoder module adopts multi-layer convolution and batch normalization operations, and embeds channel and spatial attention mechanisms in each part to realize the hierarchical extraction of image features and spatial downsampling; The bottleneck module uses an atrous spatial pyramid pooling structure to collect multi-scale context information in parallel; Dropout layers are embedded at the end of the encoder and in the bottleneck layer to suppress overfitting; The decoder module restores the spatial details step by step by upsampling and feature fusion, performs skip connections on the feature maps of the corresponding encoder layers, and finally generates a multi-class classification prediction map.
[0026] The encoder module is constructed as follows: Encoder module: It has four parts, and each part is connected in sequence. For details, see Figure 3 . The basic structure of each part is a first convolutional layer → a second convolutional layer → a channel and spatial attention mechanism (CBAM) layer → a pooling layer. In the fourth part, a Dropout layer (dropout rate is 0.5) is added after the CBAM layer, and then the final stage of spatial downsampling is completed through the pooling layer. The number of convolution kernels in each part of the convolutional layer is set to 64, 128, 256, and 512 respectively. The size of the convolution kernels is set to 3×3, the stride is set to 1, and batch normalization processing and ReLU activation processing are performed after the convolution operation; the pooling layer all adopts max pooling operation, the size of the pooling area kernel is set to 2×2, and the stride is set to 2.
[0027] Among them, the CBAM layer structure consists of a channel attention sub-module and a spatial attention sub-module, as detailed in Figure 5 . The CBAM layer structure includes a channel attention sub-module and a spatial attention sub-module. The channel attention sub-module performs global average pooling and global max pooling on the input feature map respectively to obtain vectors with a shape of C dimensions; the two pooling results are respectively input into a shared two-layer fully connected network. The number of nodes in the first layer is C / r, where r is set to 0.5, and ReLU activation processing is performed. The number of nodes in the second layer is C, and linear activation is performed; the two fully connected outputs are added together, and after Sigmoid activation, a channel attention coefficient with a shape of 1×1×C is obtained; the channel attention coefficient is multiplied element-wise with the input feature map, and the output channel-refined feature map has a size of H×W×C. The calculation method of the channel attention sub-module is: In the formula, M C ( x ) is the channel attention weight; σ is Sigmoid function; ReLU ( ) is the ReLU activation function; W 1. W 2 is the weight of the fully connected layer; Avgpool is the average pooling operation; Maxpool is the max pooling operation.
[0028] The spatial attention sub-module performs average pooling and max pooling on the channel-refined feature map along the channel dimension respectively to obtain an average feature map and a max feature map, both with a size of H×W×1; after concatenating the above two feature maps along the channel dimension, a two-dimensional attention feature map with a size of H×W×2 is formed; this feature map is input into a convolution with a kernel size of 7×7 and a stride of 1, and after Sigmoid activation, a spatial attention coefficient with a shape of H×W×1 is generated; the spatial attention coefficient is multiplied element-wise with the channel-refined feature map, and the output CBAM final fusion feature map has a size of H×W×C. The calculation method is: In the formula, M S ( x ) is the spatial attention weight, f 7×7 is the 7×7 convolution kernel; Concat ( ) is the concatenation operation.
[0029] In the encoder module, the CBAM layer is embedded in four parts respectively. Assuming the original image size is H×W×C o, the sizes of the feature maps corresponding to different parts are successively H×W, (H / 2)×(W / 2), (H / 4)×(W / 4), and (H / 8)×(W / 8), and the corresponding number of channels C are successively 64, 128, 256, and 512.
[0030] The bottleneck module is composed as follows: Bottleneck module: Its structure is successively a convolutional layer → an atrous spatial pyramid pooling (ASPP) layer → a Dropout layer. The size of the convolutional kernel of the convolutional layer is set to 3×3, the stride is set to 1, the number of convolutional kernels is set to 1024, and batch normalization and ReLU activation processing are performed; the ASPP layer is set with five groups of different dilation rates (1, 6, 12, 18, 24), and the number of channels of the output feature map is set to 128; the dropout rate of the Dropout layer is set to 0.5.
[0031] Among them, the structure of the ASPP layer is successively five atrous convolution branches → a global average pooling branch → feature concatenation → a fusion convolution. The number of channels of the output feature map of this module is set to 128, and the ASPP layer is set with five groups of different dilation rates (1, 6, 12, 18, 24). Five atrous convolution operations are performed in parallel on the input feature map of the convolution branch. The size of the convolutional kernel is set to 3×3 for all, the stride is set to 1 for all, and after convolution, batch normalization processing and ReLU activation processing are immediately performed. The output shape is all (H / 16)×(W / 16)×128, and the calculation formula is as follows: In the formula, H out and W out are the height and width of the output image; floo r( ) is the floor operation, which is used to ensure that the output size is an integer; H in and W in are the height and width of the input image; padding [0] and padding [1] are the number of pixels filled in the height and width directions; dilation [0] and dilation [1] are the dilation rates of the atrous convolution in the height and width directions; kernel _ size [0] and kernel _ size [1] are the sizes of the convolutional kernel in the height and width directions, which are both 3 in a 3×3 convolutional kernel; stride [0] and stride [1] are the strides of the convolution operation in the height and width directions.
[0032] The global average pooling branch performs global pooling on the entire pooled branch input feature map of size (H / 16)×(W / 16)×1024 to obtain a 1×1×1024 vector; then through a convolution operation, the size of the convolution kernel is set to 1×1, the stride is set to 1, and after convolution, batch normalization processing and ReLU activation are performed immediately to extract global semantic features; finally, bilinear interpolation is used to upsample it to (H / 16)×(W / 16)×128. Feature concatenation concatenates the outputs of the above five dilated convolution branches and the global average pooling branch in the channel dimension to generate a multi-scale fusion feature map of size (H / 16)×(W / 16)×768. Fusion convolution performs a convolution operation on the concatenated feature map, the size of the convolution kernel is set to 1×1, the stride is set to 1, and after convolution, batch normalization processing and ReLU activation are performed immediately to achieve the final channel compression and multi-scale context information fusion, and the image size is (H / 16)×(W / 16)×128. The calculation formula of this module is: In the formula, Conv 1×1 ( ) is the 1×1 convolution operation; f i (x) is the pixel value of the output feature image of the th dilated convolution; k is the number of dilated convolutions; f gap (x) is the pixel value of the output feature image of the global average pooling branch; Conv 3×3,ri is the 3×3 dilated convolution operation with different dilation rates r i ; r i is the th dilation rate of the dilated convolution.
[0033] The decoder module is composed as follows: Decoder module: It has four parts, and each part is connected in sequence. See Figure 4。The basic structure of each part is: upsampling layer → skip connection splicing layer → first convolutional layer → second convolutional layer. After the second convolution in the fourth part, a multi-class classification prediction map is generated through an output convolutional layer. The upsampling layer uses 2x bilinear interpolation. After upsampling, a convolution operation is performed. The size of the convolution kernel is set to 2×2, and the stride is set to 1. The inputs of the channel splicing layers of each part are, in sequence, the output of the first upsampling layer of the decoder module and the output of the fourth part of the encoder module; the output of the second upsampling layer and the output of the third part of the encoder module; the output of the third upsampling layer and the output of the second part of the encoder module; the output of the fourth upsampling layer and the output of the first part of the encoder module. The number of convolution kernels in the first and second convolutional layers of the four parts of the decoder module is set to 512, 256, 128, and 64 in sequence. The size of the convolution kernels is all set to 3×3, and the stride is all set to 1. After the convolution operation, batch normalization processing and ReLU activation processing are both performed. The size of the convolution kernel of the output convolutional layer is set to 1×1, the stride is set to 1, and Softmax activation is used.
[0034] The formula for batch normalization processing is as follows: In the formula, is the mean within the batch; is the activation value of the th sample in the batch; m is the number of samples in the batch; is the variance within the batch; is the normalized activation value; ε is a very small positive number used to prevent the denominator from being zero.
[0035] Step S2, make an aquatic vegetation classification dataset; The dataset consists of multiple remote sensing images containing aquatic vegetation and their corresponding classification labels. Each image has undergone preprocessing, sliding window cropping, and data augmentation. The processed dataset is divided into a training set, a validation set, and a test set; Step S201, obtain and preprocess remote sensing images; The high-resolution remote sensing image data of the present invention comes from the surface reflectance (SR) Level 2 Tier 1 remote sensing images of Landsat-5 TM, Landsat-7 ETM+, Landsat-8 OLI, and Landsat-9 OLI2 in the Landsat Collection 2 dataset. On the Google Earth Engine (GEE) platform, according to the content in the code editor of the Earth Engine data catalog, the pixel values of different types of images are converted into surface reflectance. For some cloudy images, cloud masking functions are defined to remove clouds; for the strip missing problem of Landsat-7 images, Landsat-5 or Landsat-8 images with a relatively recent date are selected to fill the strips. The lake shape file is loaded from the GeoJSON file for lake area cropping. For Landsat-5 and Landsat-7 image data, bands B1, B2, B3, B4, B5, and B7 are selected to form a new image; for Landsat-8 and Landsat-9 image data, bands B2, B3, B4, B5, B6, and B7 are selected to form a new image. The bands of the new image are renamed as B1, B2, B3, B4, B5, and B6, and the range and function of each band are similar. After the above processing, the data is exported in the GEE platform. For subsequent processing convenience, the coordinate system is converted to EPSG:32650 in Python.
[0036] Step S202, making classification labels; According to the current situation and conditions of the lake, it is classified into five categories: water body, bare land, submerged vegetation, floating-leaved vegetation, and swamp / emergent vegetation. Based on field investigations and historical research data, combined with the characteristics of different aquatic vegetation, the Normalized Difference Water Index (NDWI), Red band (RED), the second component of the Tasseled Cap transformation (TC2), and the second component of the Principal Component Analysis (PC2) are selected as the classification basis. Batch calculations are performed on the original image in Python to obtain the corresponding NDWI, TC2, and PC2 images. A classification decision tree is constructed in ENVI 5.6, and the specific structure is as follows: if (NDWI > a) then (Class = water body) if not (Class = non-water body), if (Red > b) then (Class = bare land) if not (Class = wetland vegetation), if (TC2 > c) then (Class = swamp / emergent vegetation) if not (Class = other vegetation); if (PCA2 < d) then (Class = submerged vegetation) if not (Class = floating-leaved vegetation), where a, b, c, and d are segmentation thresholds and are adjusted according to the specific image characteristics. The classification map of typical lake features is obtained using the decision tree, and the swamp / emergent vegetation, floating-leaved vegetation, and submerged vegetation are selected and output as a single shape file. Then the file is transferred to Arcgis to add coordinate information and an id value field for different vegetation categories. The corresponding relationship is 0 - background, 85 - submerged vegetation, 170 - floating-leaved vegetation, 255 - swamp / emergent vegetation, and the image is converted into a raster file based on the id field.
[0037] Step S203, produce and divide the model training data set; Select the B2, B3, B4, NDWI, TC2, and PC2 band combinations as the training images, and change the pixel bit depth of the images to 8 bit in Arcgis. The produced images and label data correspond one by one. They are cropped into 256×256-sized training samples through a sliding window, and data augmentation operations such as horizontal flipping, vertical flipping, and diagonal mirroring are performed on them. Then, for the image data with a relatively high proportion of submerged, floating-leaved, and swamp / emergent vegetation, secondary augmentation operations such as random 90-degree rotation, random brightness, Gaussian noise, elastic transformation, translation scaling, and rotation are performed to obtain complete training samples. Then the sample data is divided into a training set, a validation set, and a test set according to the ratio of 8:1:1.
[0038] Step S3, train and validate the deep convolutional neural network model; Input the training set and the validation set into the constructed CAD-Unet model, update the network parameters through end-to-end backpropagation until the best classification accuracy is achieved, and select the accuracy evaluation index to verify the trained CAD-Unet model; Step S301, train the constructed deep CAD-Unet model; The training set and validation set data are input into the constructed CAD-Unet model. Through end-to-end backpropagation, the Adam optimizer is used for gradient descent to update the network weights. The cross-entropy function is used as the loss function, and the learning rate is dynamically adjusted in combination with the learning rate scheduling strategy. That is, when the validation set loss has not improved for 5 consecutive epochs, the learning rate is decayed by a factor of 0.7 but not lower than 1×10⁻ 6 ; the batch size is 2, and the initial learning rate is 1×10⁻ 4 , and it is trained for 100 epochs until the Accuracy of the validation set reaches the optimum, and the trained CAD-Unet model is obtained.
[0039] Step S302, evaluate and validate the deep convolutional neural network model; Use indicators such as Overall Accuracy (OA), Kappa coefficient, Producer's Accuracy (PA), and User's Accuracy (UA) for accuracy evaluation. The calculation formulas are as follows: In the formula, X KK is k the number of correctly classified N classes, and is k the total sum of reference data for the target object; is k the total sum of the evaluated data for
[0040] To verify the effectiveness of the present invention, the CAD-Unet model of the present invention is compared with the U-Net model, and the model accuracies are compared using the same data set and training strategy. The accuracy verification results of the U-Net and the CAD-Unet model of the present invention are shown in Table 1. It can be seen from the table that there are obvious differences between the U-Net and the CAD-Unet model of the present invention in vegetation classification. From the perspective of producer's accuracy, the CAD-Unet model of the present invention is superior to the U-Net in all three categories of background, floating-leaved vegetation, and marsh / emergent vegetation, and only the producer's accuracy of submerged vegetation slightly decreases. In terms of user's accuracy, the CAD-Unet model of the present invention is superior to the U-Net in all categories, and the improvement in submerged vegetation and marsh / emergent vegetation is particularly significant, indicating that it can effectively reduce the false detection rate and enhance the classification reliability. The overall accuracy and Kappa coefficient of the CAD-Unet model of the present invention are higher than those of the U-Net, further verifying its advantages in global classification consistency and class balance. Figure 7For the predicted images and labeled images of the U-Net and the CAD-Unet model of the present invention, as can be seen from the figure, the predicted images (c1, c2, c3) of the CAD-Unet model of the present invention are closer to the labeled images (b1, b2, b3), and there are obvious misclassification and missing classification phenomena in the predicted images (a1, a2, a3) of the U-Net model. The CAD-Unet model of the present invention is more accurate in classifying small and fragmented patches, with clearer contours, stronger generalization ability, and significantly fewer cases of incorrect recognition and missing classification. However, there are still a small number of missing classification phenomena for floating leaves and submerged vegetation. Generally speaking, the classification effect of the CAD-Unet model of the present invention is better, with less overall noise and more accurate recognition.
[0041] Step S4, perform prediction on the remotely sensed image to be classified; Perform large-image prediction on the remotely sensed image to be classified. Adopt a sliding cropping strategy with a fixed-size sub-window to perform edge-ignoring block division and normalization operations on the remotely sensed image to be classified. The size of the sub-window is set to 256×256. Input each sub-window into the trained CAD-Unet model of the present invention in parallel for inference, and fuse and reconstruct the classification probability maps of each sub-window based on an overlapping compensation strategy to generate a complete multi-class aquatic vegetation classification result.
[0042] Embodiment 2 This embodiment provides an aquatic vegetation classification device based on remotely sensed images, including a processor and a memory; programs or instructions are stored in the memory, and the programs or instructions are loaded and executed by the processor to implement the steps of the aquatic vegetation classification method based on remotely sensed images provided in Embodiment 1.
[0043] Embodiment 3 This embodiment provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is made to execute the aquatic vegetation classification method based on remotely sensed images provided in Embodiment 1.
[0044] The features and benefits of the present invention are illustrated by embodiments. The embodiments are only used to explain the technical solutions of the present invention and do not constitute a limitation thereto. The present invention is not limited to the specific feature combinations in the exemplified examples. The relevant features can exist alone or in other forms of combination. Those skilled in the art can make reasonable modifications, substitutions, or optimizations to the implementation schemes without departing from the basic concept of the present invention, and its essence still belongs to the technical scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. A method for classifying aquatic vegetation based on remote sensing images, characterized in that, It includes the following steps: Step S1, construct a deep convolutional neural network, which consists of an encoder module, a bottleneck module, and a decoder module; The encoder module adopts multi-layer convolution and batch normalization operations, and embeds channel and spatial attention mechanisms in each part to achieve gradual extraction of image features and spatial downsampling; The encoder module has four parts, and the basic structure of each part is a convolutional layer, a second convolutional layer, a channel and spatial attention mechanism layer, and a pooling layer in sequence. In the fourth part, a Dropout layer is added after the channel and spatial attention mechanism layer, and then the last-level spatial downsampling is completed through the pooling layer; The bottleneck module uses an atrous spatial pyramid pooling structure to parallelly collect multi-scale context information; The structure of the bottleneck module is a convolutional layer, an atrous spatial pyramid pooling layer, and a Dropout layer in sequence; The decoder module performs upsampling and feature fusion, makes skip connections for the feature maps of the corresponding encoder layers, gradually restores spatial details, and finally generates a multi-class classification prediction map; The decoder module has four parts, and the basic structure of each part is an upsampling layer, a skip connection splicing layer, a first convolutional layer, and a second convolutional layer in sequence. After the second convolution in the fourth part, a multi-class classification prediction map is generated through an output convolutional layer; The inputs of the skip connection splicing layers of each part are the output of the upsampling layer of the first part of the decoder module and the output of the fourth part of the encoder module; the output of the upsampling layer of the second part and the output of the third part of the encoder module; the output of the upsampling layer of the third part and the output of the second part of the encoder module; the output of the upsampling layer of the fourth part and the output of the first part of the encoder module; Step S2, make an aquatic vegetation classification dataset, which is divided into a training set, a validation set, and a test set; Step S3, use the dataset to train the constructed deep convolutional neural network; Step S4, input the remote sensing image to be classified into the trained deep convolutional neural network, and output the classification result of the aquatic vegetation in the image.
2. The aquatic vegetation classification method based on remote sensing images according to claim 1, wherein The channel and spatial attention mechanism layer includes a channel attention sub-module and a spatial attention sub-module; The channel attention sub-module performs global average pooling and global max pooling on the input feature map respectively to obtain vectors with a shape of C dimensions; Input the two-way pooling results into a shared two-layer fully connected network respectively. The number of nodes in the first fully connected network is C / r, where r is set to 0.5, and ReLU activation processing is performed. The number of nodes in the second fully connected network is C, and linear activation is performed; Add the two-way fully connected outputs, and after Sigmoid activation, obtain a channel attention coefficient with a shape of 1×1×C; Multiply the channel attention coefficient and the input feature map channel by channel, and output the channel-refined feature map; The calculation method of the channel attention sub-module is as follows: In the formula, M C ( x ) is the channel attention weight; σ is Sigmoid a function; ReLU ( ) is the ReLU activation function; W 1. W 2 is the weight of the fully connected layer; Avgpool is the average pooling operation; Maxpool is the max pooling operation; The spatial attention sub-module performs average pooling and max pooling on the channel-refined feature map along the channel dimension respectively to obtain the average feature map and the max feature map; the above average feature map and max feature map are concatenated in the channel dimension to form a two-dimensional attention feature map; the two-dimensional attention feature map is input into a convolution with a convolution kernel size of 7×7 and a stride of 1, and after being activated by Sigmoid, a spatial attention coefficient is generated; the spatial attention coefficient is multiplied element-wise with the channel-refined feature map, and the final fused feature map of the channel and spatial attention mechanism layer is output; The calculation method of the spatial attention sub-module is as follows: In the formula, M S ( x ) is the spatial attention weight, f 7×7 is a 7×7 convolution kernel; Concat ( ) is the concatenation operation.
3. The aquatic vegetation classification method based on remote sensing images according to claim 1, characterized in that The structure of the atrous spatial pyramid pooling layer is successively a five-way atrous convolution branch, a global average pooling branch, feature concatenation and fusion convolution; The number of channels of the feature map output by this module is set to 128, and five groups of different dilation rates are set for the dilated convolution branch. The five groups of different dilation rates are 1, 6, 12, 18, and 24 respectively. Five dilated convolution operations are performed on the input feature map of the convolution branch in parallel. The size of the convolution kernel is set to 3×3, and the stride is set to 1. After convolution, batch normalization processing and ReLU activation processing are performed immediately. The calculation formula is as follows: In the formula, H out and W out are the height and width of the output image; floo r( ) is the floor operation, which is used to ensure that the output size is an integer; H in and W in are the height and width of the input image; padding [0] and padding [1] are the number of pixels filled in the height and width directions; dilation [0] and dilation [1] are the dilation rates of the dilated convolution in the height and width directions; kernel _ size [0] and kernel _ size [1] are the sizes of the convolution kernel in the height and width directions, both of which are 3 in a 3×3 convolution kernel; stride [0] and stride [1] are the strides of the convolution operation in the height and width directions; The global average pooling branch performs global pooling on the input feature map of the pooling branch to obtain a vector of 1×1×1024; then through a convolution operation, the size of the convolution kernel is set to 1×1, the stride is set to 1, and after convolution, batch normalization processing and ReLU activation are performed immediately to extract global semantic features; finally, bilinear interpolation is used to upsample it; Feature concatenation concatenates the outputs of the above five-way atrous convolution branch and the global average pooling branch in the channel dimension to generate a multi-scale fused feature map; Fusion convolution performs a convolution operation on the concatenated multi-scale fused feature map, the size of the convolution kernel is set to 1×1, the stride is set to 1, and after convolution, batch normalization processing and ReLU activation are performed immediately to achieve the final channel compression and multi-scale context information fusion; The calculation formula of the hollow spatial pyramid pooling layer is as follows: In the formula, Conv 1×1 ( ) is a 1×1 convolution operation; f i (x) is the pixel value of the output feature map of the th dilated convolution; k is the number of dilated convolutions; f gap (x) is the pixel value of the output feature map of the global average pooling branch; Conv 3×3,ri is the 3×3 dilated convolution operation with different dilation rates r i ; r i is the th dilation rate of the dilated convolution.
4. The method for classifying aquatic vegetation based on remote sensing images according to claim 3, wherein The calculation formula for the batch normalization process is as follows: In the formula, is the mean within the batch; is the activation value of the th sample in the batch; m is the number of samples in the batch; is the variance within the batch; is the activation value after normalization; ε is a very small positive number used to prevent the denominator from being zero.
5. The aquatic vegetation classification method based on remote sensing images according to any one of claims 1-4, characterized in that, In step S1, the number of convolution kernels in each part of the convolutional layer of the encoder module is set to 64, 128, 256, 512 respectively, the size of the convolution kernels is all set to 3×3, and the stride is all set to 1; after the convolution operation, batch normalization processing and ReLU activation processing are performed; the pooling layer all adopts max pooling operation, the size of the pooling area kernel is all set to 2×2, and the stride is all set to 2; In step S1, the size of the convolution kernel of the convolutional layer of the bottleneck module is set to 3×3, the stride is set to 1, the number of convolution kernels is set to 1024, and batch normalization and ReLU activation processing are performed; the ASPP layer sets five groups of different dilation rates, and the dilation rates are 1, 6, 12, 18 and 24 respectively; the number of channels of the output feature map is set to 128; the dropout rate of the Dropout layer is set to 0.5; In step S1, the upsampling layer of the decoder module adopts 2-fold bilinear interpolation, and after upsampling, a convolution operation is performed, the size of the convolution kernel is set to 2×2, and the stride is set to 1; For the four parts of the decoder module, the number of convolutional kernels in the first and second convolutional layers are set to 512, 256, 128, and 64 in sequence. The size of the convolutional kernels is set to 3×3, the stride is set to 1, and batch normalization and ReLU activation are performed after the convolution operation. The size of the convolutional kernel of the output convolutional layer is set to 1×1, the stride is set to 1, and Softmax activation is used.
6. The method for classifying aquatic vegetation based on remote sensing images according to any one of claims 1-4, characterized in that, In step S2, when making the aquatic vegetation classification dataset, the pixel bit depths of the remote sensing images and their corresponding classification labels are both set to 8 bit. A sliding window of 256×256 is used for image and label cropping, and the data is subjected to a primary enhancement operation, which includes horizontal flipping, vertical flipping, and diagonal mirroring; then, for the images and labels with fewer samples in a category, a secondary enhancement operation is performed, which includes random 90-degree rotation, random brightness, Gaussian noise, elastic transformation, translation and scaling, and rotation; the training set, validation set, and test set are divided according to the ratio of 8:1:1; the training set and validation set are normalized by the maximum value, and at the same time, a color dictionary is obtained to perform one-hot encoding on the labels.
7. The aquatic vegetation classification method based on remote sensing images according to any one of claims 1-4, characterized in that, In step S3, the training set and validation set data are input into the deep convolutional neural network, and the network weights are updated by the gradient descent method using the Adam optimization algorithm through end-to-end backpropagation. The loss function uses the cross-entropy function, and a learning rate scheduling strategy is combined, that is, when the performance index of the validation set has not improved within the preset number of iterations, the learning rate is reduced proportionally until the convergence condition is met, and finally, a trained deep convolutional neural network is obtained.
8. The aquatic vegetation classification method based on remote sensing images according to claim 1, wherein In step S4, for large-scale prediction of remote sensing images, a sliding cropping strategy with a fixed-size sub-window is used to perform edge-ignoring block division and normalization operations on the remote sensing images to be classified. The size of the sub-window is set to 256×256, and each sub-window is input into the trained deep convolutional neural network in parallel for inference, and the classification probability maps of each sub-window are fused and recombined based on the overlapping compensation strategy to generate a complete multi-category aquatic vegetation classification result.
9. An aquatic vegetation classification device based on remote sensing images, characterized in that, It includes a processor and a memory; programs or instructions are stored in the memory, and the programs or instructions are loaded and executed by the processor to implement the steps of the aquatic vegetation classification method based on remote sensing images as described in any one of claims 1 to 8.
10. A computer-readable storage medium, on which programs or instructions are stored, and when the programs or instructions are executed by a processor, the steps of the aquatic vegetation classification method based on remote sensing images as described in any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Method and system for identifying forest stand tree species based on multi-source remote sensing data fusion
CN120544052A
Tobacco stem and leaf image segmentation method
CN120976549A
Remote sensing image water body and shadow classification method based on double-classification attention network
CN121033675A
Water body and shadow classification method for remote sensing image based on double classification attention network
CN121033675B