Coastal wetland classification method based on lightweight multidirectional compression attention network

By using a lightweight multi-directional compressed attention network to process remote sensing images, the problems of accuracy and lightweight in cross-climate belt tasks are solved, and efficient wetland classification and information mining are achieved.

CN119992214AActive Publication Date: 2025-05-13SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510158143.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-13
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

Traditional wetland classification models are difficult to accurately deal with wetland types at different latitudes and terrain in cross-climate zone tasks, and the model structure is complex and difficult to lighten.

Method used

The coastal wetland classification method based on a lightweight multi-directional compressed attention network is adopted to obtain and preprocess the remote sensing image, and wetland screening is carried out, and the wetland image is input into the lightweight multi-directional compressed attention network to obtain the final classification result.

Benefits of technology

The accuracy of wetland classification across climate zones and the lightweight model are achieved, the performance in complex tasks is improved, and the remote sensing data information mining capabilities are expanded.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992214A_ABST
    Figure CN119992214A_ABST
Patent Text Reader

Abstract

The invention discloses a coastal wetland classification method based on a lightweight multidirectional compression attention network. The method comprises the following steps. Acquiring a remote sensing image and preprocessing the remote sensing image to obtain a preprocessed remote sensing image; performing wetland screening on the preprocessed remote sensing image to obtain a wetland image; and inputting the wetland image into the lightweight multi-direction compression attention network to obtain a final classification result graph. According to the method, complexity and classification precision are balanced, and cross-climate zone coastal wetland classification and extraction are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning, and more specifically, to a coastal wetland classification method based on a lightweight multi-directional compressed attention network. Background Art

[0002] With the vigorous development of artificial intelligence, deep learning technology has been widely used to solve the problem of wetland classification.

[0003] Traditional wetland classification models are difficult to effectively handle the differences in latitudes and terrains of the same wetland type in cross-climate zone tasks, which leads to inaccurate wetland classification across climate zones. At the same time, general models rely on multi-head attention mechanisms to achieve axial attention, so the structure is complex and the model is difficult to lightweight.

[0004] The prior art discloses a coastal wetland remote sensing classification method based on a hierarchical strategy, including: step 1, preprocessing remote sensing data; step 2, sampling based on vegetation index obtained by remote sensing index, water index and brightness auxiliary data obtained based on tasseled cap transformation; step 3, coarse classification of multispectral images and auxiliary data; step 4, masking cultivated land, forest land and urban areas; step 5, using LBP local binary pattern operator to describe local texture features of the image; step 6, combining spectral information and spatial texture information to perform image segmentation on water bodies and wetland areas; step 7, obtaining coastal wetland vector data; step 8, making coastal wetland thematic maps. This method is based on traditional machine learning methods for wetland classification, and remote sensing data information mining is incomplete. Summary of the invention

[0005] The present invention aims to solve the defects of the prior art that the wetland classification across climate zones is inaccurate and the model is difficult to lightweight, and provides a coastal wetland classification method based on a lightweight multi-directional compressed attention network. The method has the characteristics of accurate classification of wetlands across climate zones and lightweight model.

[0006] The primary purpose of the present invention is to solve the above technical problems. The technical solutions of the present invention are as follows:

[0007] A coastal wetland classification method based on lightweight multi-directional compressed attention network, including:

[0008] S1: Acquire a remote sensing image and perform preprocessing to obtain a preprocessed remote sensing image;

[0009] S2: performing wetland screening on the pre-processed remote sensing image to obtain a wetland image;

[0010] S3: Input the wetland image into the lightweight multi-directional compressed attention network to obtain the final classification result map.

[0011] Furthermore, in step S1, preprocessing includes: cloud removal and median filtering.

[0012] Furthermore, the wetland screening includes:

[0013] S201: Calculating the normalized building index NDBI, the normalized vegetation index NDVI, and the improved normalized water index MNDWI for the preprocessed remote sensing image;

[0014] S202: according to the normalized building index NDBI, the normalized vegetation index NDVI, and the improved normalized water index MNDWI, using the Otsu algorithm to analyze and obtain a decision tree threshold;

[0015] S203: Using a decision tree threshold and a decision tree algorithm to determine whether each pixel in the preprocessed remote sensing image is a wetland, and summarizing the wetland determination results to obtain a wetland image.

[0016] Furthermore, the calculation formulas of the normalized building index NDBI, the normalized vegetation index NDVI, and the improved normalized water index MNDWI are as follows:

[0017]

[0018] NIR represents the near infrared band in the remote sensing image, RED represents the red light band summarized by the remote sensing image, MIR represents the mid-infrared band in the remote sensing image, and GREEN represents the green light band summarized by the remote sensing image.

[0019] Furthermore, the lightweight multi-directional compressed attention network includes: a lightweight attention enhancement convolution unit, a multi-directional axial compressed attention unit, a cross attention unit, and a classification result generation unit;

[0020] The wetland image is input to the input end of the lightweight attention enhanced convolution unit, the output end of the lightweight attention enhanced convolution unit is connected to the input end of the multi-directional axial compression attention unit, the output end of the multi-directional axial compression attention unit is connected to the input end of the cross attention unit, the output end of the cross attention unit is connected to the input end of the classification result generation unit, and the output end of the classification result generation unit outputs the final classification result map.

[0021] Furthermore, the lightweight attention-enhanced convolution unit includes: a first lightweight attention-enhanced convolution module, a second lightweight attention-enhanced convolution module, and a third lightweight attention-enhanced convolution module;

[0022] The first lightweight attention-enhanced convolutional module, the second lightweight attention-enhanced convolutional module, and the third lightweight attention-enhanced convolutional module have the same structure, and all include: a first convolutional layer, a second convolutional layer, a third convolutional layer, a first batch of normalization layers, a second batch of normalization layers, a third batch of normalization layers, a first multiplication point, a second multiplication point, a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a fourth batch of normalization layers, a fifth batch of normalization layers, a sixth batch of normalization layers, a first addition point, a second addition point, a seventh convolutional layer, a seventh batch of normalization layers, and a first activation layer;

[0023] The output end of the first convolutional layer is connected to the input end of the first batch of normalization layers, the output end of the second convolutional layer is connected to the input end of the second batch of normalization layers, the output end of the third convolutional layer is connected to the input end of the third batch of normalization layers, the output end of the first batch of normalization layers and the output end of the second batch of normalization layers are connected to the input end of the first multiplication point, the output end of the first multiplication point and the output end of the third batch of normalization layers are connected to the input end of the second multiplication point, the output end of the second multiplication point is connected to the input end of the fourth convolutional layer, the input end of the fifth convolutional layer, and the input end of the sixth convolutional layer, the output end of the fourth convolutional layer is connected to the fourth batch of normalization layers. The output end of the fifth convolutional layer is connected to the input end of the fifth batch normalization layer, the output end of the sixth convolutional layer is connected to the input end of the sixth batch normalization layer, the output end of the fourth batch normalization layer and the output end of the fifth batch normalization layer are connected to the input end of the first addition point, the output end of the first addition point and the output end of the sixth batch normalization layer are connected to the input end of the second addition point, the output end of the second addition point is connected to the input end of the seventh convolutional layer, the output end of the seventh convolutional layer is connected to the input end of the seventh batch normalization layer, and the output end of the seventh batch normalization layer is connected to the input end of the first activation layer.

[0024] Further, the multi-directional axial compression attention unit includes a first multi-directional axial compression attention module, a second multi-directional axial compression attention module, and a third multi-directional axial compression attention module;

[0025] The first multi-directional axial compression attention module, the second multi-directional axial compression attention module, and the third multi-directional axial compression attention module have the same structure, and all include: an average pooling layer, a skewness calculation layer, a kurtosis calculation layer, a variance calculation layer, a third addition point, and a fourth addition point;

[0026] The output end of the average pooling layer is connected to the input end of the skewness calculation layer, the input end of the kurtosis calculation layer, and the input end of the variance calculation layer. The output end of the skewness calculation layer, the output end of the kurtosis calculation layer, the output end of the variance calculation layer, and the input end of the average pooling layer are connected to the input end of the third addition point. The output end of the third addition point is connected to the input end of the fourth addition point.

[0027] Furthermore, the calculation formula of the skewness calculation layer is as follows:

[0028]

[0029] E represents expectation, i represents sequence number, x represents i represents the i-th value of the random variable X, μ represents the mean of the random variable values, and σ represents the standard deviation of the random variable values;

[0030] The calculation formula of the kurtosis calculation layer is as follows:

[0031]

[0032] X represents a random variable, n represents the total number of values ​​of the random variable, i represents the sequence number, and x i represents the i-th value of the random variable X, represents the mean of the random variable X;

[0033] The calculation formula of the variance calculation layer is as follows:

[0034]

[0035] p i Then the random variable X takes x i The probability of n is the total number of values ​​that the random variable can take, x i Represents the i-th value of the random variable X, where i represents the sequence number and X represents the random variable.

[0036] Further, the cross attention unit includes: a first cross attention module, a second cross attention module, and a third cross attention module;

[0037] The first cross attention module, the second cross attention module, and the third cross attention module have the same structure, and all include: an eighth batch normalization layer, a ninth batch normalization layer, a first linear layer, a second linear layer, a third linear layer, a fourth linear layer, an attention score layer, and a result summary layer;

[0038] The output end of the eighth batch normalization layer is connected to the input end of the first linear layer, the output end of the first linear layer is connected to the input end of the second linear layer, the output end of the ninth batch normalization layer is connected to the input end of the third linear layer, the output end of the third linear layer is connected to the input end of the fourth linear layer, the output end of the second linear layer and the output end of the fourth linear layer are connected to the input end of the attention score layer, and the output end of the attention score layer and the output end of the second linear layer are connected to the input end of the result summary layer.

[0039] A coastal wetland classification system based on lightweight multi-directional compressed attention network, including:

[0040] Preprocessing module: acquiring remote sensing images and preprocessing them to obtain preprocessed remote sensing images;

[0041] Wetland screening module: performing wetland screening on the pre-processed remote sensing image to obtain a wetland image;

[0042] Network reasoning module: The wetland image is input into the lightweight multi-directional compressed attention network to obtain the final classification result map.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] The present invention uses a lightweight multi-directional compressed attention network to reduce the complexity of the model structure, so that axial attention can play a greater role, effectively capture the detailed information in the image, directional dependencies, and the coordination between features, and has stronger adaptability and expressiveness, improving performance in complex tasks. At the same time, remote sensing images are acquired and preprocessed; the preprocessed remote sensing images are screened for wetlands to obtain wetland images; and the remote sensing data information mining capabilities are expanded. In summary, the present invention balances complexity and classification accuracy to achieve the classification and extraction of coastal wetlands across climate zones. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 A flowchart of a coastal wetland classification method based on a lightweight multi-directional compressed attention network provided in Example 1.

[0046] Figure 2 Flowchart for wetland screening provided for Example 1.

[0047] Figure 3 Structural diagram of the lightweight multi-directional compressed attention network provided in Example 1.

[0048] Figure 4 Schematic diagram of the principle of the lightweight multi-directional compressed attention network provided in Example 1.

[0049] Figure 5 Structural diagram of the lightweight attention-enhanced convolutional module provided in Example 1.

[0050] Figure 6 Schematic diagram of the principle of the lightweight attention-enhanced convolution module provided in Example 1.

[0051] Figure 7 Structural diagram of the multi-directional axial compression attention module provided in Example 1.

[0052] Figure 8 Schematic diagram of the principle of the multi-directional axial compression attention unit provided in Example 1.

[0053] Fig. 9 A structural diagram of the cross-attention module provided in Example 1.

[0054] Fig.10 Schematic diagram of the principle of the cross-attention module provided in Example 1.

[0055] Fig.11 This is a comparison chart of the classification results of coastal wetlands around the South China Sea provided in Example 1.

[0056] Fig.12 This is a comparison chart of the classification results of the first data set provided in Example 1.

[0057] Fig.13 This is a comparison chart of the classification results of the second data set provided in Example 1.

[0058] Fig.14 This is a comparison chart of the classification results of the third data set provided in Example 1. DETAILED DESCRIPTION

[0059] The drawings are for illustrative purposes only and should not be construed as limiting the present patent;

[0060] In order to better illustrate the present embodiment, some parts in the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product;

[0061] It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0062] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0063] Example 1

[0064] like Figure 1 As shown, a coastal wetland classification method based on a lightweight multi-directional compressed attention network includes:

[0065] S1: Acquire a remote sensing image and perform preprocessing to obtain a preprocessed remote sensing image;

[0066] S2: performing wetland screening on the pre-processed remote sensing image to obtain a wetland image;

[0067] S3: Input the wetland image into the lightweight multi-directional compressed attention network to obtain the final classification result map.

[0068] Furthermore, in step S1, the preprocessing includes: cloud removal and median filtering.

[0069] In a specific embodiment, the preprocessing method is: generate a shp file of the study area based on ArcGIS, supplement and update the geographic information such as projection, and upload the shp file to GEE Assets. Based on the GEE remote sensing database, decloud and median filter the Sentinel-2 image in the designated study area to obtain an image that can be processed by the network.

[0070] Furthermore, if Figure 2 As shown, the wetland screening includes:

[0071] S201: Calculating the normalized building index NDBI, the normalized vegetation index NDVI, and the improved normalized water index MNDWI for the preprocessed remote sensing image;

[0072] S202: according to the normalized building index NDBI, the normalized vegetation index NDVI, and the improved normalized water index MNDWI, using the Otsu algorithm to analyze and obtain a decision tree threshold;

[0073] S203: Using a decision tree threshold and a decision tree algorithm to determine whether each pixel in the preprocessed remote sensing image is a wetland, and summarizing the wetland determination results to obtain a wetland image.

[0074] It should be noted that the Otsu algorithm (OTSU) refers to an automatic threshold selection algorithm for image segmentation. Its core idea is to segment the image into foreground and background by traversing all possible thresholds so that the intra-class variance between the two parts is minimized and the inter-class variance is maximized. As shown in the following formula:

[0075]

[0076] in, is the between-class variance, defined as:

[0077]

[0078] Among them, w1(t) and w2(t) are the weights of the foreground and background segmented by threshold t. μ1(t) and μ2(t) are the means of the foreground and background. The NDBI value of the image NDBI(x,y) is [t L ,t U ] interval, and the NDVI value of the image NDVI(x,y) is in [t L ,t U ] interval, and the MNDWI value of the image MNDWI(x,y) is in [t L ,t U ] interval.

[0079] When the pixel value satisfies the condition that NDBI is less than the OTSU threshold and NDVI is greater than the OTSU threshold, the pixel is a vegetation or water area, otherwise the pixel is a building area or a bare land area without vegetation coverage. Considering that the tropical solar altitude angle leads to high reflectivity of water bodies, water bodies can be easily misjudged as buildings and eliminated. Therefore, MNDWI greater than the OTSU threshold is also classified as the maximum wetland range. This will effectively improve the classification efficiency.

[0080] Furthermore, the calculation formulas of the normalized building index NDBI, the normalized vegetation index NDVI, and the improved normalized water index MNDWI are as follows:

[0081]

[0082] NIR represents the near infrared band in the remote sensing image (wavelength is about 780-2526nm), RED represents the red light band summarized in the remote sensing image (wavelength is about 600-700nm), MIR represents the mid-infrared band in the remote sensing image (wavelength is about 3.0-20μm), and GREEN represents the green light band summarized in the remote sensing image (wavelength is about 492-577nm).

[0083] Furthermore, if Figure 3 As shown, the lightweight multi-directional compressed attention network includes: a lightweight attention enhancement convolution unit, a multi-directional axial compression attention unit, a cross attention unit, and a classification result generation unit;

[0084] The wetland image is input to the input end of the lightweight attention enhanced convolution unit, the output end of the lightweight attention enhanced convolution unit is connected to the input end of the multi-directional axial compression attention unit, the output end of the multi-directional axial compression attention unit is connected to the input end of the cross attention unit, the output end of the cross attention unit is connected to the input end of the classification result generation unit, and the output end of the classification result generation unit outputs the final classification result map.

[0085] It should be noted that if Figure 4 As shown in the figure, the input of the lightweight multi-directional compressed attention network is a data cube with a spatial range of 5x5 from the remote sensing image. First, the lightweight attention enhancement convolution unit (LAConv) is used for detail enhancement. Then, the data is enhanced in three axes by the multi-directional axial compressed attention unit (MASA) to generate three complementary cube features. Subsequently, these patch features are input into the cross attention unit (CA) for multi-feature collaborative fusion. Finally, the features are processed by the fully connected layer and the prediction results of the network are output. This architecture can effectively capture the detailed information, directional dependencies and mutual collaboration between features in the image by introducing a multi-level attention mechanism.

[0086] Furthermore, the lightweight attention-enhanced convolution unit includes: a first lightweight attention-enhanced convolution module, a second lightweight attention-enhanced convolution module, and a third lightweight attention-enhanced convolution module;

[0087] like Figure 5 As shown, the first lightweight attention enhanced convolution module, the second lightweight attention enhanced convolution module, and the third lightweight attention enhanced convolution module have the same structure, and all include: a first convolution layer, a second convolution layer, a third convolution layer, a first batch of normalization layers, a second batch of normalization layers, a third batch of normalization layers, a first multiplication point, a second multiplication point, a fourth convolution layer, a fifth convolution layer, a sixth convolution layer, a fourth batch of normalization layers, a fifth batch of normalization layers, a sixth batch of normalization layers, a first addition point, a second addition point, a seventh convolution layer, a seventh batch of normalization layers, and a first activation layer;

[0088] The output end of the first convolutional layer is connected to the input end of the first batch of normalization layers, the output end of the second convolutional layer is connected to the input end of the second batch of normalization layers, the output end of the third convolutional layer is connected to the input end of the third batch of normalization layers, the output end of the first batch of normalization layers and the output end of the second batch of normalization layers are connected to the input end of the first multiplication point, the output end of the first multiplication point and the output end of the third batch of normalization layers are connected to the input end of the second multiplication point, the output end of the second multiplication point is connected to the input end of the fourth convolutional layer, the input end of the fifth convolutional layer, and the input end of the sixth convolutional layer, the output end of the fourth convolutional layer is connected to the fourth batch of normalization layers. The output end of the fifth convolutional layer is connected to the input end of the fifth batch normalization layer, the output end of the sixth convolutional layer is connected to the input end of the sixth batch normalization layer, the output end of the fourth batch normalization layer and the output end of the fifth batch normalization layer are connected to the input end of the first addition point, the output end of the first addition point and the output end of the sixth batch normalization layer are connected to the input end of the second addition point, the output end of the second addition point is connected to the input end of the seventh convolutional layer, the output end of the seventh convolutional layer is connected to the input end of the seventh batch normalization layer, and the output end of the seventh batch normalization layer is connected to the input end of the first activation layer.

[0089] It should be noted that if Figure 6As shown in the figure, the lightweight attention enhancement convolution unit (LAConv) consists of three parallel branches and plays an important role, especially in detail enhancement, which has a significant impact on the improvement of subsequent features. Each branch contains a 2D convolution layer and a batch normalization layer, and generates three matrices respectively: the query matrix q, the key matrix k, and the value matrix v. Among them, the transpose of the key matrix k is multiplied by the query matrix q, and the resulting matrix is ​​multiplied by the value matrix v, and then the calculation of the attention mechanism is realized through the convolution operation. In order to further enhance the feature details, the model uses depthwise separable convolution for subsequent feature refinement. In the multi-directional compressed attention network, the first lightweight attention enhancement convolution module, the second lightweight attention enhancement convolution module, and the third lightweight attention enhancement convolution module are arranged in parallel and share weights. This design not only generates three detail-enhanced data cubes at the same time, providing a solid foundation for subsequent feature enhancement, but also effectively reduces the complexity of the model.

[0090] Further, the multi-directional axial compression attention unit includes a first multi-directional axial compression attention module, a second multi-directional axial compression attention module, and a third multi-directional axial compression attention module;

[0091] like Figure 7 As shown, the first multi-directional axial compression attention module, the second multi-directional axial compression attention module, and the third multi-directional axial compression attention module have the same structure, and all include: an average pooling layer, a skewness calculation layer, a kurtosis calculation layer, a variance calculation layer, a third addition point, and a fourth addition point;

[0092] The output end of the average pooling layer is connected to the input end of the skewness calculation layer, the input end of the kurtosis calculation layer, and the input end of the variance calculation layer. The output end of the skewness calculation layer, the output end of the kurtosis calculation layer, the output end of the variance calculation layer, and the input end of the average pooling layer are connected to the input end of the third addition point. The output end of the third addition point is connected to the input end of the fourth addition point.

[0093] It should be noted that if Figure 8As shown in the figure, the multi-directional axial compressed attention unit (MASA) optimizes the performance of the attention mechanism while maintaining effectiveness by enhancing the interpretability of the model and reducing the difficulty of training. The core task of the multi-directional axial compressed attention unit (MASA) is to process the three data cubes generated by the multi-directional axial compressed attention unit (LAConv) and convert them into three vectors in different directions through pooling operations. These vectors are enhanced along the three axes (i.e., x-axis, y-axis, and z-axis), respectively, where three statistics, skewness, kurtosis, and variance, are introduced. Specifically, variance is used to express the texture and edge information that may exist in the image, and in channel attention, it helps to identify channels that show significant changes in the image. Skewness helps to capture the possible asymmetry in the image, such as lighting shadows or reflections, thereby enhancing the robustness and generalization ability of the model in diverse environments. The application of skewness in channel attention further strengthens the nonlinear characteristics on the channel. Kurtosis is usually used for outlier detection, and its high peak value indicates that there may be significant features or noise in some positions of the image on a certain axis. These three statistics are weighted by three learnable parameters to ensure that the Multi-Directional Axial Compressed Attention Unit (MASA) has stronger adaptability and expressiveness while keeping the module lightweight.

[0094] In a specific embodiment, an activation layer may be added before the third addition point.

[0095] Furthermore, the calculation formula of the skewness calculation layer is as follows:

[0096]

[0097] E represents expectation, i represents sequence number, x represents i represents the i-th value of the random variable X, μ represents the mean of the random variable values, and σ represents the standard deviation of the random variable values;

[0098] The calculation formula of the kurtosis calculation layer is as follows:

[0099]

[0100] X represents a random variable, n represents the total number of values ​​of the random variable, i represents the sequence number, and x i represents the i-th value of the random variable X, represents the mean of the random variable X;

[0101] The calculation formula of the variance calculation layer is as follows:

[0102]

[0103] p i Then the random variable X takes x iThe probability of n is the total number of values ​​that the random variable can take, x i Represents the i-th value of the random variable X, where i represents the sequence number and X represents the random variable.

[0104] Further, the cross attention unit includes: a first cross attention module, a second cross attention module, and a third cross attention module;

[0105] like Fig. 9 As shown, the structures of the first cross attention module, the second cross attention module, and the third cross attention module are the same, and all include: an eighth batch normalization layer, a ninth batch normalization layer, a first linear layer, a second linear layer, a third linear layer, a fourth linear layer, an attention score layer, and a result summary layer;

[0106] The output end of the eighth batch normalization layer is connected to the input end of the first linear layer, the output end of the first linear layer is connected to the input end of the second linear layer, the output end of the ninth batch normalization layer is connected to the input end of the third linear layer, the output end of the third linear layer is connected to the input end of the fourth linear layer, the output end of the second linear layer and the output end of the fourth linear layer are connected to the input end of the attention score layer, and the output end of the attention score layer and the output end of the second linear layer are connected to the input end of the result summary layer.

[0107] It should be noted that if Fig.10 As shown in the figure, the cross attention unit (CA) is a technology for information interaction across modalities or different scales. In this model, a Transformer architecture based on cross attention, CrossTransformer, is adopted to achieve effective complementarity of the above three branch features. In summary, the three branches extract features in three directions respectively, and these features are patch features (Context Feature) and target features (Target Feature) for each other. By introducing bidirectional interactive attention, cross-channel information fusion can be achieved. This structural design enables cross attention to effectively capture the interdependence between features in different directions, thereby enhancing the feature expression ability of the model. Through this information interaction mechanism, the model can make more full use of the complementarity of features in different directions and improve its performance in complex tasks.

[0108] A coastal wetland classification system based on lightweight multi-directional compressed attention network, including:

[0109] Preprocessing module: acquiring remote sensing images and preprocessing them to obtain preprocessed remote sensing images;

[0110] Wetland screening module: performing wetland screening on the pre-processed remote sensing image to obtain a wetland image;

[0111] Network reasoning module: The wetland image is input into the lightweight multi-directional compressed attention network to obtain the final classification result map.

[0112] It should be noted that if Figure 11 to Figure 14 As shown in the figure, the lightweight multi-directional compressed attention network (LMDSAN-CCWC) is suitable for large-scale, cross-climate zones and can be deployed in the cloud. It has a clever engineering design that successfully minimizes the number of model parameters and maintains model accuracy. The model is based on a convolutional module and uses three visual Transformer structures with weight sharing to construct enhanced features. It calculates compressed attention in three axes based on the three features, and uses cross attention to coordinate the features of the three branches. The entire model combines mathematical models with networks, allowing this engineering design to perform better than other classic models in three coastal wetland features (i.e., mangroves, beach and swamp). In the classification task of coastal wetlands around the South China Sea, the overall accuracy OA of the present invention is 1.0099 times that of SeaFormer (LMDSAN-CCWC OA is 0.9500, SeaFormer is 0.9406), and the training time is 0.6176 times that of SeaFormer (LMDSAN-CCWC training time is 106.27s, SeaFormer is 172.06s), based on the parameter amount of only 0.009 times of the lightweight neural network SeaFormer (LMDSAN-CCWC parameter amount is 57KB, SeaFormer parameter amount is 6579KB), and the training time is 0.6176 times that of SeaFormer (LMDSAN-CCWC training time is 106.27s, SeaFormer is 172.06s). The method has been experimentally proven to be deployable in the cloud, and effectively solves the classification accuracy problem of large-scale research areas, constructs high-precision coastal wetland mapping, and provides technical support for the research on coastal wetland resource management and ecological protection.

[0113] The same or similar reference numerals correspond to the same or similar components;

[0114] The terms used in the drawings to describe positional relationships are only used for illustrative purposes and should not be construed as limiting this patent;

[0115] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.

Claims

1. A coastal wetland classification method based on lightweight multi-directional compressed attention network, characterized in that: include: S1: Acquire a remote sensing image and perform preprocessing to obtain a preprocessed remote sensing image; S2: performing wetland screening on the pre-processed remote sensing image to obtain a wetland image; S3: Input the wetland image into the lightweight multi-directional compressed attention network to obtain the final classification result map.

2. According to claim 1, a coastal wetland classification method based on lightweight multi-directional compressed attention network is characterized in that: In step S1, preprocessing includes: cloud removal and median filtering.

3. According to claim 1, a coastal wetland classification method based on lightweight multi-directional compressed attention network is characterized in that: The wetland screening includes: S201: Calculating the normalized building index NDBI, the normalized vegetation index NDVI, and the improved normalized water index MNDWI for the preprocessed remote sensing image; S202: according to the normalized building index NDBI, the normalized vegetation index NDVI, and the improved normalized water index MNDWI, using the Otsu algorithm to analyze and obtain a decision tree threshold; S203: Using a decision tree threshold and a decision tree algorithm to determine whether each pixel in the preprocessed remote sensing image is a wetland, and summarizing the wetland determination results to obtain a wetland image.

4. According to claim 1, a coastal wetland classification method based on a lightweight multi-directional compressed attention network is characterized in that: The calculation formulas of the normalized building index NDBI, the normalized vegetation index NDVI, and the improved normalized water index MNDWI are as follows: NIR represents the near infrared band in the remote sensing image, RED represents the red light band summarized by the remote sensing image, MIR represents the mid-infrared band in the remote sensing image, and GREEN represents the green light band summarized by the remote sensing image.

5. According to claim 1, a coastal wetland classification method based on lightweight multi-directional compressed attention network is characterized in that: The lightweight multi-directional compressed attention network includes: a lightweight attention enhancement convolution unit, a multi-directional axial compression attention unit, a cross attention unit, and a classification result generation unit; The wetland image is input to the input end of the lightweight attention enhanced convolution unit, the output end of the lightweight attention enhanced convolution unit is connected to the input end of the multi-directional axial compression attention unit, the output end of the multi-directional axial compression attention unit is connected to the input end of the cross attention unit, the output end of the cross attention unit is connected to the input end of the classification result generation unit, and the output end of the classification result generation unit outputs the final classification result map.

6. According to claim 5, a coastal wetland classification method based on lightweight multi-directional compressed attention network is characterized in that: The lightweight attention-enhanced convolution unit comprises: a first lightweight attention-enhanced convolution module, a second lightweight attention-enhanced convolution module, and a third lightweight attention-enhanced convolution module; The first lightweight attention-enhanced convolutional module, the second lightweight attention-enhanced convolutional module, and the third lightweight attention-enhanced convolutional module have the same structure, and all include: a first convolutional layer, a second convolutional layer, a third convolutional layer, a first batch of normalization layers, a second batch of normalization layers, a third batch of normalization layers, a first multiplication point, a second multiplication point, a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a fourth batch of normalization layers, a fifth batch of normalization layers, a sixth batch of normalization layers, a first addition point, a second addition point, a seventh convolutional layer, a seventh batch of normalization layers, and a first activation layer; The output end of the first convolutional layer is connected to the input end of the first batch of normalization layers, the output end of the second convolutional layer is connected to the input end of the second batch of normalization layers, the output end of the third convolutional layer is connected to the input end of the third batch of normalization layers, the output end of the first batch of normalization layers and the output end of the second batch of normalization layers are connected to the input end of the first multiplication point, the output end of the first multiplication point and the output end of the third batch of normalization layers are connected to the input end of the second multiplication point, the output end of the second multiplication point is connected to the input end of the fourth convolutional layer, the input end of the fifth convolutional layer, and the input end of the sixth convolutional layer, the output end of the fourth convolutional layer is connected to the fourth batch of normalization layers. The output end of the fifth convolutional layer is connected to the input end of the fifth batch normalization layer, the output end of the sixth convolutional layer is connected to the input end of the sixth batch normalization layer, the output end of the fourth batch normalization layer and the output end of the fifth batch normalization layer are connected to the input end of the first addition point, the output end of the first addition point and the output end of the sixth batch normalization layer are connected to the input end of the second addition point, the output end of the second addition point is connected to the input end of the seventh convolutional layer, the output end of the seventh convolutional layer is connected to the input end of the seventh batch normalization layer, and the output end of the seventh batch normalization layer is connected to the input end of the first activation layer.

7. According to claim 5, a coastal wetland classification method based on lightweight multi-directional compressed attention network is characterized in that: A multi-directional axial compression attention unit, comprising a first multi-directional axial compression attention module, a second multi-directional axial compression attention module, and a third multi-directional axial compression attention module; The first multi-directional axial compression attention module, the second multi-directional axial compression attention module, and the third multi-directional axial compression attention module have the same structure, and all include: an average pooling layer, a skewness calculation layer, a kurtosis calculation layer, a variance calculation layer, a third addition point, and a fourth addition point; The output end of the average pooling layer is connected to the input end of the skewness calculation layer, the input end of the kurtosis calculation layer, and the input end of the variance calculation layer. The output end of the skewness calculation layer, the output end of the kurtosis calculation layer, the output end of the variance calculation layer, and the input end of the average pooling layer are connected to the input end of the third addition point. The output end of the third addition point is connected to the input end of the fourth addition point.

8. According to claim 7, a coastal wetland classification method based on lightweight multi-directional compressed attention network is characterized in that: The calculation formula of the skewness calculation layer is as follows: E represents expectation, i represents sequence number, x represents i represents the i-th value of the random variable X, μ represents the mean of the random variable values, and σ represents the standard deviation of the random variable values; The calculation formula of the kurtosis calculation layer is as follows: X represents a random variable, n represents the total number of values ​​of the random variable, i represents the sequence number, and x i represents the i-th value of the random variable X, represents the mean of the random variable X; The calculation formula of the variance calculation layer is as follows: p i Then the random variable X takes x i The probability of n is the total number of values ​​that the random variable can take, x i Represents the i-th value of the random variable X, where i represents the sequence number and X represents the random variable.

9. According to claim 5, a coastal wetland classification method based on lightweight multi-directional compressed attention network is characterized in that: The cross attention unit includes: a first cross attention module, a second cross attention module, and a third cross attention module; The first cross attention module, the second cross attention module, and the third cross attention module have the same structure, and all include: an eighth batch normalization layer, a ninth batch normalization layer, a first linear layer, a second linear layer, a third linear layer, a fourth linear layer, an attention score layer, and a result summary layer; The output end of the eighth batch normalization layer is connected to the input end of the first linear layer, the output end of the first linear layer is connected to the input end of the second linear layer, the output end of the ninth batch normalization layer is connected to the input end of the third linear layer, the output end of the third linear layer is connected to the input end of the fourth linear layer, the output end of the second linear layer and the output end of the fourth linear layer are connected to the input end of the attention score layer, and the output end of the attention score layer and the output end of the second linear layer are connected to the input end of the result summary layer.

10. A coastal wetland classification system based on a lightweight multi-directional compressed attention network, applied to the classification method according to any one of claims 1 to 9, characterized in that: include: Preprocessing module: acquiring remote sensing images and preprocessing them to obtain preprocessed remote sensing images; Wetland screening module: performing wetland screening on the pre-processed remote sensing image to obtain a wetland image; Network reasoning module: The wetland image is input into the lightweight multi-directional compressed attention network to obtain the final classification result map.

Citation Information

Patent Citations

  • Hybrid precision quantification method of deep neural network and related device

    CN119204154A

  • Infrared spectrogram correlation intelligent detection method and apparatus

    WO2016106956A1