Sea-land segmentation method based on double attention mechanism
By adopting a land and sea segmentation method based on a dual attention mechanism in coastline extraction and dynamic monitoring, the problems of inefficiency and accuracy error in the existing technology are solved, and a more accurate and efficient coastline extraction and segmentation effect is achieved.
Patent Information
- Application Number
- CN202510062720.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-01-15
AI Technical Summary
The prior art has problems of inefficiency, accuracy error and difficulty in handling complex sea and land scenarios in coastline extraction and dynamic monitoring, especially when faced with multi-scale features, fuzzy information and small-scale targets.
The sea and land segmentation method based on the dual attention mechanism is adopted, and the encoder and decoder of the dual attention mechanism are constructed, combined with the convolutional block attention module, efficient multi-head attention mechanism, pyramid pooling module and adaptive feature fusion module, to realize multi-scale feature extraction and semantic segmentation of coastline images.
It improves the accuracy and efficiency of coastline extraction, can handle complex sea and land boundaries and small-scale targets more accurately, enhances the ability to capture key areas and edge details, and provides more accurate coastline details and segmentation results.
Smart Images

Figure CN120107779A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a land and sea segmentation method based on a dual attention mechanism. Background Art
[0002] The coastline is defined as "the trace of the boundary between land and sea at the average high tide level of many years of spring tide". In the field of land and sea segmentation of remote sensing images, it is called the instantaneous coastline, that is, the boundary between the sea and the land under specific time and tidal conditions. It not only has rich natural resources, but is also one of the areas with the most frequent human activities. However, under the dual influence of natural and human factors such as sea level rise, land subsidence, river sediment transport, reclamation projects and port construction, the coastline has become one of the most dynamic natural boundaries on the earth's surface. In China, the total length of the coastline is about 32,000 kilometers (18,000 kilometers of mainland coastline and 14,000 kilometers of island coastline), and more than one-third of the coastline is eroded. The dynamic changes of the coastline not only threaten the economic and social development of coastal areas, but also pose a challenge to the scientific management of the ecological environment and coastal resources. Therefore, the rapid and accurate extraction of the coastline and monitoring of its dynamic changes are of great significance for maintaining the ecological balance of coastal areas, assessing potential risks in extreme sea level events, and disaster prevention and mitigation.
[0003] Traditional coastline measurement methods are mainly manual field measurements, but manual field measurements take a long time, are labor-intensive, and are easily affected by subjective factors of surveyors, resulting in low efficiency and precision errors, making it difficult to achieve dynamic monitoring of the coastline. In contrast, remote sensing technology obtains information on the earth's surface from a long distance through detection equipment. It has the advantages of wide coverage, multi-temporal monitoring and high resolution. It is widely used in land resource surveys, agricultural development, and marine monitoring, and has become the main technical means for monitoring dynamic changes in coastlines. At present, the main methods for automatic coastline extraction based on remote sensing images at home and abroad include threshold segmentation methods, edge detection operator methods, and object-oriented methods. However, the above methods have the following disadvantages: 1) Although the threshold segmentation method is simple, it requires manual setting of the threshold. If the threshold is not set properly, it is easy to affect the accuracy of coastline extraction; 2) The edge detection operator method is greatly affected by noise, and the detected sea and land edges are not continuous enough. Subsequent processing is usually required after the coastline is extracted; 3) The object-oriented method is difficult to process high-resolution remote sensing images containing a large amount of data, and cannot fully utilize the useful information in the image. Therefore, these methods have certain limitations when dealing with multi-scale features, complex sea and land scenes, and coastline morphology.
[0004] In recent years, deep learning technology has developed rapidly. Relying on its powerful feature extraction capabilities, it has achieved remarkable success in various computer vision downstream tasks such as image classification, target detection and semantic segmentation. Deep learning technology has also made significant progress in the field of land and sea segmentation of remote sensing images. However, it is worth noting that in the land and sea segmentation task of remote sensing images, when faced with a large number of aquaculture areas, biological communities and plankton in coastal areas, deep learning models are difficult to clearly and smoothly divide the land and sea boundaries if they only rely on local feature extraction. In addition, due to the existence of artificial structures such as ports and docks and small-scale information such as narrow waterways, the model often faces challenges in processing edge details, making the land and sea segmentation results of existing land and sea segmentation models inaccurate. Therefore, it is very important to propose a land and sea segmentation method that can be applied to complex land and sea boundaries and interference factors such as small-scale targets. Summary of the invention
[0005] In view of the above-mentioned prior art, the present invention provides a land-sea segmentation method based on a dual attention mechanism, which mainly solves the technical problems existing in the above-mentioned background technology.
[0006] To achieve the above object, the technical solution of the embodiment of the present invention is implemented as follows:
[0007] A land-sea segmentation method based on a dual attention mechanism comprises the following steps:
[0008] S1, acquiring a remote sensing image, and preprocessing the remote sensing image to obtain a coastline remote sensing image to be segmented;
[0009] S2. Constructing a land-sea segmentation model with dual attention mechanism:
[0010] S2.1, construct the encoder of dual attention mechanism;
[0011] S2.2, construct an adaptive feature fusion module;
[0012] S2.3, constructing a multi-scale information fusion decoder based on the adaptive feature fusion module;
[0013] S3, training the land-sea segmentation model to obtain a trained land-sea segmentation model; inputting the coastline remote sensing image to be segmented into the trained land-sea segmentation model to obtain a land-sea segmentation result map of the remote sensing image.
[0014] Optionally, the encoder includes four Transformer Blocks, namely Transformer Block1, Transformer Block2, Transformer Block3 and Transformer Block4;
[0015] The Transformer Block1, Transformer Block2, Transformer Block3 and Transformer Block4 are connected in sequence;
[0016] Each of the Transformer Blocks includes a patch embedding layer, a Transformer layer, a convolutional attention module, and an Overlap Patch Merging layer.
[0017] Optionally, the Transformer layer includes an efficient multi-head attention mechanism and a hybrid feedforward network; the hybrid feedforward network includes a 1×1 convolution layer, a 3×3 convolution layer, an activation function and a Dropout layer;
[0018] The number of Transformer layers in the Transformer Block1 is three, the number of Transformer layers in the Transformer Block2 is four, the number of Transformer layers in the Transformer Block3 is six, and the number of Transformer layers in the Transformer Block4 is three.
[0019] Optionally, the expression of the efficient multi-head attention mechanism is:
[0020] Q=XW Q ,K=XW K ,V=XW V
[0021]
[0022] MultiHead(Q,K,V)=Concat(Z 1 ,Z 2 ,...,Z h )W O
[0023] X kv =Conv2d(X,kernel_Size=sr_ratio,stride=sr_ratio)
[0024] OutPut=X+Dropout(ProjDrop(MultiHead(Q,K,V)))
[0025] Among them, X represents the input feature, Q, K, V represent the query, key and value matrix respectively, and W Q , W Kand W V is the corresponding weight matrix; d k is the dimension of the key vector, h is the number of heads, and W O is the output linear transformation matrix.
[0026] Optionally, the expression of the convolutional attention module is:
[0027] M C (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F)))×F
[0028] F'=M C (F)×F
[0029] M S (F)=σ(Conv2d(Concat[AvgPool(F),MaxPool(F)]))
[0030] F”=M S (F')×F'
[0031] Among them, F is the input feature, M C is the channel attention weight, F' is the output after the channel attention module, σ represents the Sigmoid activation function, M S is the spatial attention weight, and F' is the final output after the spatial attention module.
[0032] Optionally, the decoder includes a pyramid pooling module, two residual convolution blocks, an adaptive feature fusion module, a 1×1 convolution layer and an upsampling layer; the connection between the two residual convolution blocks is a residual connection;
[0033] The input of the pyramid pooling module is connected to the output of the Transformer Block4, and the pyramid pooling module is used to extract multi-scale context information from the feature map output by the Transformer Block4;
[0034] The two residual convolution blocks are used to extract local features from the feature map output by the pyramid pooling module;
[0035] The adaptive feature fusion module is used to upsample the feature maps output by the two residual convolution blocks and fuse the features of the encoder;
[0036] The 1×1 convolutional layer is used to map the feature map output by the adaptive feature fusion module into a binary classification sea-land segmentation result map;
[0037] The upsampling layer is used to align the land-sea segmentation result image with the original input image in terms of spatial resolution.
[0038] Optionally, the adaptive feature fusion module includes dynamic snake convolution, 1×1 convolution, a channel attention mechanism based on global average pooling, a channel attention mechanism based on global maximum pooling, and a reverse residual block;
[0039] The expression of the adaptive feature fusion module is:
[0040] S1=Conv2d(Concat(S_high,DSC(S_low)))
[0041]
[0042] Among them, DSC represents dynamic snake convolution, Concat represents channel-level concatenation operation, GAP and GMP represent global average pooling and global maximum pooling respectively, MLP is multi-layer perceptron, σ is Sigmoid activation function, It is the feature map after IRB processing.
[0043] Optionally, the number of the adaptive feature fusion modules is three, namely a first adaptive feature fusion module, a second adaptive feature fusion module and a third adaptive feature fusion module;
[0044] The first adaptive feature fusion module is jump-connected with the Transformer Block1;
[0045] The second adaptive feature fusion module is jump-connected with the Transformer Block2;
[0046] The third adaptive feature fusion module is jump-connected to the Transformer Block3, and the outputs of the two residual convolution blocks are connected to the third adaptive feature fusion module;
[0047] The third adaptive feature fusion module, the second adaptive feature fusion module, and the first adaptive feature fusion module are connected in sequence.
[0048] Optionally, the two residual convolution blocks both include Conv3×3, batch normalization and ReLU activation functions;
[0049] The expressions of the two residual convolution blocks are:
[0050] Y=F(X)+X
[0051] Where X is the input feature map, and F(X) is the output after several layers of convolution, activation, normalization, etc.
[0052] The beneficial effects of the present invention are as follows: a trained sea-land segmentation model is obtained by training a coastline image through constructing a sea-land segmentation model, and the coastline image to be segmented is input into the trained sea-land segmentation model to obtain a sea-land segmentation result; wherein the sea-land segmentation model, in the encoding stage, introduces a convolutional block attention module and an efficient multi-head attention mechanism to form a dual attention encoder, which is used to fully combine global context with local information to enhance the ability to capture key areas and edge details; in the decoding stage, a pyramid pooling module, a residual convolution block and an adaptive feature fusion module are introduced to form a decoder with multi-scale information fusion capability, which can effectively retain low-level edge information, while gradually integrating semantic features from a high level, enhancing the association between features of different resolutions, and finally depicting more accurate coastline details, and finally obtaining an accurate sea-land segmentation result. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 A schematic flow chart of a land-sea segmentation method based on a dual attention mechanism provided in an embodiment of the present invention;
[0054] Figure 2 An overall block diagram of a land-sea segmentation method based on a dual attention mechanism provided in an embodiment of the present invention;
[0055] Figure 3 A schematic diagram of a convolutional attention module provided in an embodiment of the present invention;
[0056] Figure 4 A schematic diagram of a Transformer Block provided in an embodiment of the present invention;
[0057] Figure 5 A schematic diagram of an adaptive feature fusion module provided in an embodiment of the present invention;
[0058] Figure 6 The segmentation result diagram is based on the test image in the GF-HNCD dataset;
[0059] Figure 7 This is the segmentation result diagram based on the test image in the BSD dataset;
[0060] Figure 8 This is the result diagram of the ablation experiment. DETAILED DESCRIPTION
[0061] The technical solution of the present invention is further elaborated in detail below in conjunction with the drawings and specific embodiments of the specification. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by technicians in the technical field of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. In the following description, the expression "some embodiments" is related to a subset of all possible embodiments, but it should be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0062] In the following description, a large number of specific details are provided to provide a more thorough understanding of the present invention. However, it is apparent to those skilled in the art that the present invention can be implemented without one or more of these details. In other examples, in order to avoid confusion with the present invention, some technical features well known in the art are not described.
[0063] It should be understood that the present invention can be implemented in different forms and should not be interpreted as being limited to the embodiments proposed herein. On the contrary, providing these embodiments will make the disclosure thorough and complete, and the scope of the present invention will be fully conveyed to those skilled in the art. And the purpose of the terms used herein is only to describe specific embodiments and is not intended to be a limitation of the present invention. When used herein, the singular forms of "one", "one" and "said / the" are also intended to include plural forms, unless the context clearly indicates another way. It should also be understood that the terms "compose" and / or "include" when used in this specification determine the presence of the features, integers, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, parts and / or groups. When used herein, the term "and / or" includes any and all combinations of the relevant listed items.
[0064] It should also be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be a central element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be a central element at the same time. The terms "vertical", "horizontal", "inside", "outside", "left", "right" and similar expressions used herein are for illustrative purposes only and are not intended to be the only implementation method.
[0065] In order to fully understand the present invention, a detailed structure will be proposed in the following description to illustrate the technical solution proposed by the present invention. The optional embodiments of the present invention are described in detail as follows, but in addition to these detailed descriptions, the present invention may also have other implementations.
[0066] Example
[0067] Please refer to the attached Figure 1 and attached Figure 2 , the present application provides a land-sea segmentation method based on a dual attention mechanism, comprising the following steps:
[0068] S1, acquiring a remote sensing image, and preprocessing the remote sensing image to obtain a coastline remote sensing image to be segmented;
[0069] Specifically, images containing sea and land areas are obtained from satellites, drones or other remote sensing platforms, and pre-processing is performed on the acquired remote sensing images, such as denoising, correction and contrast enhancement, to improve image quality. Through pre-processing, a clearer coastline image is obtained;
[0070] S2. Constructing a land-sea segmentation model with dual attention mechanism:
[0071] S2.1, construct the encoder of dual attention mechanism;
[0072] S2.2, construct an adaptive feature fusion module;
[0073] S2.3, constructing a multi-scale information fusion decoder based on the adaptive feature fusion module;
[0074] S3, training the land-sea segmentation model to obtain a trained land-sea segmentation model; inputting the coastline remote sensing image to be segmented into the trained land-sea segmentation model to obtain a land-sea segmentation result map of the remote sensing image.
[0075] Specifically, a land-sea segmentation model (DA-MiTUNet) including an encoder and a decoder with a dual attention mechanism is constructed, wherein the encoder is used to extract features from the coastline image to obtain features with multiple resolutions; the decoder is used to fuse the multiple features with different resolutions; the remote sensing image is preprocessed to obtain the coastline image, which is then input into the land-sea segmentation model for training. After the trained land-sea segmentation model is obtained, the remote sensing image to be segmented is input into the trained land-sea segmentation model to obtain the segmentation result map of the remote sensing image.
[0076] As an optional implementation, the encoder includes four Transformer Blocks, namely Transformer Block1, Transformer Block2, Transformer Block3 and Transformer Block4;
[0077] The Transformer Block1, Transformer Block2, Transformer Block3 and Transformer Block4 are connected in sequence;
[0078] Each of the Transformer Blocks includes a patch embedding layer, a Transformer layer, a convolutional attention module, and an Overlap Patch Merging layer;
[0079] It should be noted that Vision Transformer (ViT) is different from the traditional image processing mode. It divides the input image into fixed-size image patches and converts them into sequence data for processing, relying on the powerful expression ability and self-attention mechanism of Transformer to learn and process image data. However, ViT uses fixed-size image patches and static position encoding, lacks a hierarchical structure, and fails to effectively utilize multi-scale features, resulting in limited feature expression capabilities when processing complex images and high computational complexity. Mix Transformer (MiT) has improved these problems by using a Transformer encoder with a hierarchical structure, which can generate multi-level and multi-scale features based on the input image, including high-resolution shallow features and low-resolution deep features, thereby improving the accuracy of semantic segmentation. Unlike ViT, MiT adopts a position-free encoding design and uses 3×3 convolutions to represent position information. This method better understands the spatial relationship between pixels and further improves efficiency, accuracy, and robustness. Therefore, the land and sea segmentation model in the present invention constructs an encoder based on MiT;
[0080] For details, please refer to the attached Figure 4, the preprocessed coastline image is input into the encoder. First, the preprocessed coastline image is downsampled in Transformer Block1, Transformer Block2, Transformer Block3 and Transformer Block4 in turn. In the Transformer layers of Transformer Block1, Transformer Block2, Transformer Block3 and Transformer Block4, different "patch_sizes" (7, 3, 3, 3) and "strides" (4, 2, 2, 2) are used respectively to obtain feature maps with resolutions of 1 / 4, 1 / 8, 1 / 16 and 1 / 32 respectively; in each Transformer Block, the image is first divided into small blocks and flattened into vectors through the patch embedding layer, and then processed through the Transformer layer. The feature map after processing by the Transformer layer will enter the Overlap Patch The merging layer merges overlapping small blocks. This layer introduces overlapping areas between small blocks in the feature map, performs downsampling and merging, so as to effectively retain the continuity of local features, generate richer and spatially consistent features, and provide better input for the next encoding stage. Based on this, the encoder can gradually extract and fuse multi-scale features of the image, providing rich feature information for downstream decoders and classification tasks.
[0081] The Transformer layer includes an efficient multi-head attention mechanism and a hybrid feedforward network; the hybrid feedforward network includes a 1×1 convolution layer, a 3×3 convolution layer, an activation function and a Dropout layer;
[0082] The number of Transformer layers in the Transformer Block1 is three, the number of Transformer layers in the Transformer Block2 is four, the number of Transformer layers in the Transformer Block3 is six, and the number of Transformer layers in the Transformer Block4 is three;
[0083] The expression of the efficient multi-head attention mechanism is:
[0084] Q=XW Q ,K=XW K ,V=XW V
[0085]
[0086] MultiHead(Q,K,V)=Concat(Z 1 ,Z 2 ,...,Z h )W O
[0087] X kv =Conv2d(X,kernel_Size=sr_ratio,stride=sr_ratio)
[0088] OutPut=X+Dropout(ProjDrop(MultiHead(Q,K,V)))
[0089] Where X represents the input feature, Q, K, and V represent the query, key, and value matrices respectively, and W Q , W K and W V is the corresponding weight matrix; d k is the dimension of the key vector, h is the number of heads, and W O is the output linear transformation matrix.
[0090] Specifically, each Transformer layer includes an efficient multi-head attention mechanism (Efficient MultiheadAttention) and a mixed feedforward network (MixFFN), which are used to capture long-distance dependencies between features and perform nonlinear transformations, respectively; among them, the efficient multi-head attention mechanism introduces the spatial downsampling ratio (sr_ratio), and reduces the dimension of the input by using a "sr_ratio×sr_ratio" convolution layer before processing the query and key-value pair, thereby reducing the computational complexity. When calculating attention, it can adaptively adjust the spatial resolution of the input features to ensure the extraction of effective long-distance dependencies at different scales, that is, to capture the relationship between features globally. MixFFN combines 1×1 convolution with 3×3 deep convolution to expand the feature dimension and enhance the ability to extract spatial features while maintaining computational efficiency, and then performs nonlinear transformation through activation function (GELU) and dropout.
[0091] As an optional implementation, please refer to the attached Figure 3 , the expression of the convolutional attention module is:
[0092] M C (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F)))×F
[0093] F'=M C (F)×F
[0094] M S (F)=σ(Conv2d(Concat[AvgPool(F),MaxPool(F)]))
[0095] F”=M S (F')×F'
[0096] Among them, F is the input feature, M C is the channel attention weight, F' is the output after the channel attention module, σ represents the Sigmoid activation function, M S is the spatial attention weight, and F” is the final output after the spatial attention module;
[0097] The convolutional block attention module combines channel attention and spatial attention in sequence to refine the input features in two stages. The convolutional block attention module performs weighted optimization of channel and spatial information in a local range, allowing the model to first focus on "which channels are important" and then focus on "which positions in space are important". By inferring attention from these two dimensions, the purpose of adaptively optimizing image features is achieved, thereby capturing the key information in the features more comprehensively.
[0098] Specifically, the convolutional block attention module is integrated into the end of the Transformer layer, enabling it to further improve its perception of details and boundaries while retaining global context information. The convolutional block attention module and the efficient multi-head attention mechanism form a dual attention encoder. The feature map processed by multiple Transformer encoder layers is adjusted in terms of channels and space, which enhances the model's ability to focus on important features and suppresses the interference of unimportant information to a certain extent, so that its output has more discriminative features.
[0099] As an optional implementation, the decoder includes a pyramid pooling module, two residual convolution blocks, an adaptive feature fusion module, a 1×1 convolution layer and an upsampling layer; the connection between the two residual convolution blocks is a residual connection;
[0100] The input of the pyramid pooling module is connected to the output of the Transformer Block4, and the pyramid pooling module is used to extract multi-scale context information from the feature map output by the Transformer Block4;
[0101] The two residual convolution blocks are used to extract local features from the feature map output by the pyramid pooling module;
[0102] The adaptive feature fusion module is used to upsample the feature maps output by the two residual convolution blocks and fuse the features of the encoder;
[0103] The 1×1 convolutional layer is used to map the feature map output by the adaptive feature fusion module into a binary classification sea-land segmentation result map;
[0104] The upsampling layer is used to align the land-sea segmentation result image with the original input image in terms of spatial resolution;
[0105] Specifically, after feature extraction by the four Transformer Blocks of the encoder, contextual information of different scales is obtained, namely, features of 1 / 4, 1 / 8, 1 / 16 and 1 / 32 resolutions, including high-resolution shallow features and low-resolution deep features. Shallow features usually have high spatial resolution, so they retain rich detail information, such as low-level visual features such as edges, textures and colors, which are crucial for the boundary positioning of segmented objects; low-resolution deep features come from the latter layers of the backbone network, have a larger receptive field, provide a wider range of global contextual information, and have a strong abstract ability for objects, which can help the sea and land segmentation model understand the complex remote sensing image content and make global semantic judgments. A hierarchical decoder is designed based on the multi-scale fusion strategy, including a pyramid pooling module, two consecutive residual convolution blocks, an adaptive feature fusion module, a 1×1 convolution layer and an upsampling layer, so that the model can retain fine boundary information and ensure the consistency of overall semantics.
[0106] As an optional implementation, please refer to the attached Figure 5 , the adaptive feature fusion module includes dynamic snake convolution, 1×1 convolution, channel attention mechanism based on global average pooling, channel attention mechanism based on global maximum pooling, and reverse residual block;
[0107] The expression of the adaptive feature fusion module is:
[0108] S1=Conv2d(Concat(S_high,DSC(S_low)))
[0109]
[0110] Among them, DSC represents dynamic snake convolution, Concat represents channel-level concatenation operation, GAP and GMP represent global average pooling and global maximum pooling respectively, MLP is multi-layer perceptron, σ is Sigmoid activation function, It is the feature map after IRB processing;
[0111] Specifically, the adaptive feature fusion module (AFFM) first introduces dynamic snake convolution to improve the spatial adaptability of shallow features (S_low), better capture irregular deformation areas, and make up for the limitations of fixed receptive fields. Subsequently, the upsampled deep features are concatenated with the shallow feature maps enhanced by dynamic snake convolution, and the redundancy of the channel dimension is reduced by 1×1 convolution to form a fused feature map S1. The fused feature map contains rich semantic information and fine spatial details, which makes up for the shortcomings of a single feature. Then, the adaptive feature fusion module (AFFM) introduces a channel attention mechanism based on global average pooling (GAP) and global maximum pooling (GMP), which dynamically adjusts the importance of different channels of features by generating attention weights, and promotes the flow of information between different feature layers. Finally, the inverse residual block (IRB) is applied to the feature map S1, and it is element-wise multiplied with the previous attention weights to obtain the final output to achieve effective fusion of multi-scale features.
[0112] The number of the adaptive feature fusion modules is three, namely a first adaptive feature fusion module, a second adaptive feature fusion module and a third adaptive feature fusion module;
[0113] The first adaptive feature fusion module is jump-connected with the Transformer Block1;
[0114] The second adaptive feature fusion module is jump-connected with the Transformer Block2;
[0115] The third adaptive feature fusion module is jump-connected to the Transformer Block3, and the outputs of the two residual convolution blocks are connected to the third adaptive feature fusion module;
[0116] The third adaptive feature fusion module, the second adaptive feature fusion module, and the first adaptive feature fusion module are connected in sequence;
[0117] Specifically, in the decoding process of the decoder, after the features output by the two residual convolution blocks are sent to the third adaptive feature fusion module, they are fused with the shallower high-resolution features (1 / 16 resolution features) obtained by inputting to Transformer Block3, and the fused features are sent to the second adaptive feature fusion module, where they are fused with the shallower high-resolution features (1 / 8 resolution features) obtained by inputting to Transformer Block2, and then the fused features are sent to the first adaptive feature fusion module, where they are fused with the shallower high-resolution features (1 / 4 resolution features) obtained by inputting to Transformer Block1. Through the above process, upsampling and feature fusion are performed layer by layer, and finally a 1x1 convolution layer is used to compress the number of channels to the number of categories, and a segmentation result is generated. At the same time, the result is upsampled through an upsampling layer (upSample) to ensure that the output has the same spatial size as the input (i.e., 512×512) to obtain the final result.
[0118] As an optional implementation, the two residual convolution blocks both include Conv3×3, batch normalization and ReLU activation functions;
[0119] The expressions of the two residual convolution blocks are:
[0120] Y=F(X)+X
[0121] Where X is the input feature map, and F(X) is the output after several layers of convolution, activation, normalization, etc.
[0122] Specifically, in the decoder, the pyramid pooling module first extracts multi-scale context information from the deepest 1 / 32 resolution features to further expand the model's receptive field and enhance the understanding of global information. The features processed by the pyramid pooling module are then sent to two consecutive residual convolution blocks for further processing. The residual convolution block contains Conv3×3, batch normalization, and ReLU activation function operations. The convolution block retains global feature information while extracting local features, thereby enhancing the expression of high-level semantic features. The design of residual connections also alleviates the gradient vanishing problem in network training to a certain extent and makes the transmission of information flow smoother. Subsequently, the decoder gradually upsamples the feature map through three adaptive feature fusion modules (AFFM) and fuses the shallow high-resolution features from the encoder to ensure that important detail information is not lost when reconstructing the spatial resolution of the features. Finally, a 1×1 convolution layer maps the feature map to a binary land-sea segmentation result map, and the final upsampling layer ensures that the segmentation result is fully aligned with the original input image in terms of spatial resolution.
[0123] In the present invention, four Transformer Blocks are used as encoders to enhance the model's ability to extract multi-scale features, and then a pyramid pooling module and dynamic snake convolution, two consecutive residual convolution blocks and three adaptive feature fusion modules (AFFM) form a decoder; in addition, the pyramid pooling module and dynamic snake convolution can further enhance the model's receptive field and the ability to capture spatial dependencies and refine feature representations. During the encoding process, by generating feature maps with 1 / 4, 1 / 8, 1 / 16 and 1 / 32 resolutions, multi-scale features are extracted layer by layer to ensure that rich contextual information is captured. At the same time, we integrate the convolutional block attention module into the end of the Transformer layer of each Transformer Block, so that it can further improve the perception of details and boundaries while retaining global context information; in the decoding process, the shallower high-resolution features (1 / 4, 1 / 8 and 1 / 16 resolution features) in the encoder are directly passed to the adaptive feature fusion module (AFFM) for feature fusion through skip connections, which helps to restore edge details, while the deepest 1 / 32 resolution features need to be aggregated by the pyramid pooling module (PPM) for multi-scale context to enhance the global receptive field of the features, and then sent to two convolutional blocks with residual connections (Conv Block1 and ConvBlock2) for further refinement, and the residual connection mechanism is used to alleviate the gradient vanishing problem. Subsequently, upsampling and feature fusion are performed layer by layer, and finally a 1x1 convolution layer is used to compress the number of channels to the number of categories and generate segmentation results. At the same time, the results are upsampled by the upsampling layer to ensure that the output has the same spatial size as the input (i.e., 512×512) to obtain the final segmentation result.
[0124] In order to verify the segmentation effect of a land-sea segmentation method based on a dual attention mechanism provided in the present invention, the hardware configuration includes an Intel Core i7-13700KF CPU, an NVDIA GeForce RTX 4080-16GB GPU, 32GB RAM and a Windows 10 operating system, and the software versions include PyTorch 1.10.0 and Python 3.8.18. During the training process, AdamW is used as the optimizer, the batch size is set to 6, the initial learning rate is 5e-4, and 200 epochs of training are performed, and the cross entropy loss function is also used;
[0125] Take the following dataset:
[0126] a) GF-HNCD
[0127] The GF-HNCD dataset is a dataset specifically used for Hainan land-sea segmentation and its subsequent tasks. The dataset was taken by the GF-1WFV satellite and covers the entire Hainan area. The imaging distance is 16 meters, and it contains 8 original remote sensing images (RSI), each with a resolution of 12000×13400. For ease of processing and analysis, the original images and their labels are cropped to images of size 512×512 pixels, totaling 5010 images. The pictures include three bands: red, green, and blue (corresponding to the 4-3-2 band combination), and the two categories of ocean and land are clearly marked.
[0128] b) Benchmark sea-land dataset
[0129] The Benchmark Sea-Land Dataset (BSD) is a high-quality Chinese coastal coastline dataset constructed based on Landsat-8 OLI images through preprocessing steps such as cropping, denoising and enhancement, as well as manual semantic annotation. The dataset includes 1950 training images of size 512×512 pixels and 1411 validation and test images of the same size. Rivers and lakes in the images are regarded as land, and ultimately contain two categories: land and sea. In the experiment, images with red-green-blue (4,3,2) bands were selected.
[0130] Since the land-sea segmentation model of the present invention is mainly used for land-sea segmentation of remote sensing images, it is a subtask of semantic segmentation. Therefore, we use common semantic segmentation evaluation indicators such as MIoU and mean F1 score (mF1-Score) as quantitative evaluation indicators to evaluate the segmentation performance of the model. At the same time, in order to objectively and comprehensively evaluate the effect of the model, the model is qualitatively evaluated in combination with the visual interpretation method to evaluate the model's ability to extract small-scale features of the land-sea boundary;
[0131] In order to better understand the model performance, we use the confusion matrix to calculate the MIoU and mean F1 score (mF1-Score) of the model. The evaluation index calculation formula is as follows:
[0132]
[0133] Among them, TP, FP, FN and TN are true positive, false positive, false negative and true negative respectively.
[0134] The specific experiments are as follows:
[0135] Based on the above two data sets a and b, experiments and comparisons were conducted using U-Net, PSPNet, UPerNet, SwinUnet, SegFormer, and TCUnet with the land-sea segmentation model provided by the present invention. Among them, U-Net and PSPNet are classic CNN-based segmentation models, UPerNet is a multi-task learning framework designed based on CNN, feature pyramid network (FPN) and pyramid pooling module (PPM), SegFormer and SwinUnet are models based on pure Transformer structure, and TCUnet is a hybrid model combining CNN and Transformer. U-Net uses the original fully convolutional network without a specific pre-trained backbone, while PSPNet and UPerNet use ResNet50 as their backbone network. SegFormer uses MiT-B0 as its backbone network, and the backbone network of SwinUnet is composed of multiple Swin Transformer Blocks. The CNN branch and Transformer branch of TCUnet use ResNet and PVT V2 as their backbone networks, respectively.
[0136] The experimental environment of the above models was kept consistent, and two different data sets were used to evaluate the results of each model to increase the diversity of the data sets. At the same time, in order to ensure the fairness of the experiment, these models were not pre-trained. Through experimental comparison, the performance of the land-sea segmentation model provided by the present invention in the land-sea segmentation task of remote sensing images can be more comprehensively evaluated, and the segmentation effect is shown in Table 1:
[0137] Table 1 Comparison of performance and segmentation effect of different models
[0138]
[0139] The data in Table 1 show the quantitative evaluation indicators of different models on the GF-HNCD and BSD datasets. The results show that the land-sea segmentation model provided by the present invention outperforms other comparative models in the two key indicators of mF1 and MIoU. In the above two datasets, mF1 reached 97.46% and 97.36%, respectively, while MIoU reached 96.96% and 96.77%, respectively. From the analysis of the data in the table, it can be seen that the hybrid model TCUnet that combines CNN with Transformer outperforms SegFormer and SwinUnet based on pure Transformer, and SegFormer shows higher classification accuracy than the CNN-based method. Although SwinUnet combines the excellent global modeling capabilities of the Swin Transformer block, the model ignores the fine-grained spatial position information during the stacking process, resulting in insufficient precise alignment of spatial information, resulting in loss of edge information. This makes SwinUnet's performance on the BSD dataset inferior to PSPNet and UPerNet, and only better than U-Net. On the GF-HNCD dataset, among the CNN-based methods, PSPNet showed the highest segmentation accuracy, which was better than U-Net and UPerNet, and only slightly lower than UPerNet on the BSD dataset. This is because it introduced the pyramid pooling module (PPM), which captures multi-scale contextual information through pooling operations of different scales, thereby enhancing the model's perception of objects of different scales. UPerNet forms a multi-task learning framework by combining the feature pyramid network (FPN) and the pyramid pooling module (PPM), but its performance is similar to that of PSPNet on a single semantic segmentation task.
[0140] At the same time, two indicators, parameters and floating-point operations (FLOPs), are introduced to measure the complexity and computational efficiency of each model. The parameter amount refers to the total number of all parameters that need to be trained in the model, usually expressed in units of "millions (M)" or "billions (B)". The floating-point operation number is an indicator to measure the computational complexity of the model. It indicates the number of floating-point operations required by the model in a forward propagation process, usually expressed as "G" (i.e., 1 billion floating-point operations). As can be seen from Table 1, the UPerNet model has the highest number of parameters and floating-point operations, reaching 64.12M and 238G. In contrast, the land-sea segmentation model provided by the present invention has achieved a balance in these two aspects. Although SegFormer based on MiT-B0 and TCUNet based on ResNet and PVT V2 both have fewer parameters and FLOPs, which makes them more advantageous in terms of portability. However, the experimental results show that the land-sea segmentation model provided by the present invention surpasses all comparative models in segmentation accuracy while maintaining a low computational overhead.
[0141] Please refer to the attached Figure 6 and Figure 7 , Figure 6 The segmentation results of all comparison methods based on the test images in the GF-HNCD dataset are shown. Figure 7 The segmentation results of all comparison methods in the test image based on the BSD dataset are shown; among them, (a) Unet, (b) PSPNet, (c) UPerNet, (d) SwinUnet, (e) SegFormer, (f) TCUNet and (g) DA-MiTUNet; it can be seen that the land-sea segmentation model provided by the present invention has better segmentation results than the other six comparison models, especially in the area marked by the orange rectangular box. Based on the differences in land-sea characteristics and segmentation results of remote sensing images, we further focus on the impact of the following five typical situations on extreme sea level event prediction and coastal risk assessment:
[0142] (1) Aquaculture areas, biomes and plankton
[0143] like Figure 7 As shown, Figure 7 The first and second images show that there is a lot of fuzzy information (biological communities or suspended solids) in the coastal area. In remote sensing images, since the color and texture of biological communities or suspended solids are similar to those of adjacent land, if the overall texture information is ignored and only the pixel features are used for classification, it is easy for the model to fail to effectively distinguish whether the area is water or land. In the semantic segmentation process, the model may over-rely on local features and lack the full integration of global context information, making it difficult to identify the environmental relevance and spatial distribution characteristics of the above-mentioned areas, resulting in fuzzy sea-land boundary segmentation. Through visual observation, it can be seen that our model can effectively avoid misclassifying them as land when dealing with these fuzzy information, and has better boundary detail preservation and segmentation integrity than other models. From the comparison results, it can be seen that only model f and model g did not mistakenly classify biological communities or plankton as land, and the other comparison models all had varying degrees of misclassification. Although model f did not have the above errors, it had a problem of missed detection when dealing with islands in the ocean, resulting in some islands not being correctly identified as land, showing the limitation of complex boundary recognition. The land-sea segmentation model provided by the present invention effectively reduces these misjudgments and also shows higher recognition in the processing of fuzzy areas, thereby generating a clearer land-sea segmentation boundary.
[0144] like Figure 7 As shown, Figure 7The third image shows a typical aquaculture area, which usually has obvious artificially planned boundaries. From the comparison results in the figure, models a, c and e did not correctly identify the aquaculture area. Models d and f had misclassification or omission in the boundary part, resulting in a lack of clarity and damage to the outline of the aquaculture area bordering the land. Although model b can correctly segment the sea-land boundary, our model further identifies small pieces of land in the water body on this basis, showing good boundary distinction ability and recognition ability of small-scale targets, highlighting its superiority in complex background and detail recognition.
[0145] (2) Complex mixed coastlines
[0146] like Figure 6 As shown, Figure 6 The first image depicts a coastline type with aquaculture areas mixed with silt, which is classified as land in the label. From the segmentation results, it can be seen that all models except model C can segment relatively complete sea-land boundaries. However, our model further demonstrates strong adaptability and processing effects on complex geomorphic environments in the waterway indicated by the yellow box in this area. Figure 7 As shown, Figure 7 The fourth picture shows a mixed coastline composed of a variety of complex geographical features such as aquaculture areas, biological communities and bedrock coastlines. It can be seen that when the land-sea segmentation model provided by the present invention processes fine structures such as aquaculture areas, the details of the land-sea edge remain clearer and more consistent with the shape of the label image. In contrast, other models such as a, b, c and e divide the coastal aquaculture areas and the land as a whole, resulting in a significant extension of the land-sea boundary to the sea. The land-sea edge of d and f both show a boundary jagged effect, presenting an irregular and rough boundary. This jagged effect is usually due to the fact that the model does not refine the features finely enough during the decoding stage, resulting in a lack of accuracy in classifying edge areas. In general, the land-sea segmentation model provided by the present invention can better capture the contours of the terrain and retain the details of tiny protrusions and depressions, which is crucial for the accurate identification of mixed coastlines.
[0147] (3) Narrow waterways, artificial structures and small harbors
[0148] like Figure 6 As shown, Figure 6 The second and third pictures Figure 7The fifth and sixth figures show the advantages of our model in narrow waterway recognition. As small-scale features in remote sensing images, narrow waterways are usually difficult to extract accurately due to their limited spatial range and slender, curved shapes, which poses a challenge to fine coastline segmentation. However, from the comparison of various segmentation results, it can be seen that the land-sea segmentation model provided by the present invention can capture the complete shape and clear boundaries of these narrow waterways more finely, showing stronger recognition ability. In addition, Figure 6 The fourth to sixth images show artificial coastlines built with concrete, bricks and other building materials, including breakwaters, docks and ports. Comparing the segmentation results of various models, our model can still accurately extract the outlines and boundaries of these artificial structures in areas with complex backgrounds and unclear color transitions, maintain the integrity of the artificial coastline, and accurately distinguish different types of artificial buildings, thereby more realistically presenting the spatial layout and form of these engineering structures. These results further demonstrate the advantages of our model in dealing with small-scale features and complex terrain environments.
[0149] Figure 6 and Figure 7 The yellow framed area in the seventh picture shows a semi-enclosed water area with a tortuous and complex coastline, belonging to a small harbor water area. Figure 6 In the data, models a and b showed certain capabilities in processing small harbor areas, but failed to fully identify the waters within the harbor, with some misclassification. Model c performed the worst. Models d to g were able to basically identify the morphological characteristics of the harbor, but our model performed the best. Figure 7 In the results, all models can correctly identify the harbor waters, but the completeness varies. Models a, b, c, and e all misclassify the green vegetation near the harbor as the ocean to varying degrees, and are insufficient in eliminating internal errors.
[0150] From the segmentation tasks of narrow waterways, small harbors and artificial coastlines mentioned above, it can be seen that our model shows obvious advantages in small-scale feature extraction. It can accurately identify and retain the complete form of the target in complex backgrounds and subtle boundaries, effectively avoid misclassification, and show strong adaptability and detail capture capabilities to a variety of complex landforms.
[0151] (4) Misjudgment of ships and green vegetation (internal error)
[0152] Figure 7 The eighth to tenth figures show a case where green vegetation and ships were misclassified. The three remote sensing images in this case all showed high clarity, but in the coastline extraction task, the green vegetation within the land area was misclassified as the ocean, and the cargo ship sailing in the sea was classified as land. Figure 7In the eighth and ninth images, the vegetation areas were misidentified as oceans by other models (such as a, b, c, and e) to varying degrees. Our model successfully avoided such misjudgments through more sophisticated feature extraction and effective use of contextual information, and accurately restored the green vegetation-covered areas to land. Figure 7 The yellow box in the tenth figure shows a cargo ship sailing in the ocean. Only models b and g do not classify it as land. This type of misjudgment may be caused by the lack of context information and confusion caused by the small size of the ship. From the above comparison results, it can be seen that the land-sea segmentation model provided by the present invention is more advantageous in eliminating internal errors.
[0153] (5) Impact of remote sensing image resolution
[0154] Figure 6 The eighth and ninth images show a case affected by insufficient remote sensing image resolution. In this case, we can see that the two remote sensing images have low resolution, blurred details, uneven lighting, obvious noise interference, and insufficient contrast, which limits the ability to identify ground features and makes it difficult to meet the needs of high-precision remote sensing analysis. Figure 6 In the eighth picture, the yellow frame of model a shows obvious grid-like and locally discontinuous features. This phenomenon is usually caused by the model's inability to effectively capture subtle boundaries and continuity features of objects under low-resolution and noisy inputs, resulting in its description of edges being too rough and lacking in details, resulting in obvious unevenness. Models b and f are both affected by insufficient image resolution, misclassifying high-reflection areas with uneven illumination and inconsistent reflection features in the image as land, while model c completely fails to effectively identify land and ocean. Figure 6 In the ninth picture, due to the interference of many factors such as clouds and atmospheric scattering, all models have different degrees of misclassification. Only models e and g can restore the sea and land areas more completely. From the comparison results, it can be seen that our model benefits from the ability to integrate multi-scale features and express deep features. Its feature extraction module can still effectively enhance feature contrast when dealing with noise and low resolution, effectively reducing the adverse effects of these low-quality images. Therefore, the sea and land segmentation model provided by the present invention shows higher noise resistance, effectively avoiding serious misclassification caused by insufficient resolution, and making it more reliable in actual remote sensing applications.
[0155] In the context of extreme sea level events, accurate land-sea segmentation is critical for monitoring and analyzing coastal zone changes. First, in areas such as aquaculture areas, biomes, and plankton, segmentation errors are prone to occur due to their similarity to the surrounding land features, which poses challenges to the accurate monitoring and assessment of phenomena such as seawater erosion and intrusion. Our model can accurately identify these complex areas through a sophisticated segmentation method, providing a reliable data basis for coastal environmental changes. In addition, artificial structures such as breakwaters and ports have an important impact on coastal flow patterns and coastal erosion during extreme sea level events such as heavy rains, tides, or storm surges. Accurately identifying the complete outlines of these artificial structures helps to evaluate their effectiveness in flood control and coastal erosion prevention. Our model helps to predict the protective role of these facilities under extreme climatic conditions by accurately extracting these artificial structures, thereby providing important support for coastal protection systems, such as evaluating their effectiveness in water flow control and coastal stability during heavy rains and storm surges. At the same time, narrow waterways and small harbors are often sensitive areas where water flows converge, which can easily lead to drastic fluctuations in water levels and expansion of flooding during extreme sea level events. Our model shows strong adaptability in identifying these small-scale features, and can make accurate assessments of possible flooding risks in extreme events, providing reliable reference information for flood prevention planning and response measures in coastal areas. Regarding the problem of misjudgment of green vegetation, traditional semantic segmentation models are often prone to misclassification, while the feature loss and noise interference of low-resolution images further affect the accuracy of classification. In the context of extreme sea level events, accurate identification of green vegetation is crucial for the assessment of ecosystem health and marine socioeconomic impacts. Our model significantly reduces misjudgment and overcomes the adverse effects of low resolution by making full use of contextual information and fusing multi-scale features. These advantages effectively support emergency response and assessment of the health status of coastal ecosystems.
[0156] In summary, through high-precision land-sea segmentation, our model can play an important role in disaster prediction, risk assessment and vulnerability analysis in coastal areas. It can not only provide strong data support for the effectiveness evaluation of coastal protection facilities, but also provide stable and reliable basic information for the long-term dynamic monitoring of the coastal environment, thereby providing a more solid guarantee for the risk response and prevention of extreme sea level events.
[0157] Ablation experiment:
[0158] Based on the BSD dataset, an ablation experiment is conducted to verify the effectiveness of the land-sea segmentation model and each module proposed in this invention, and the performance of each module is evaluated using the three indicators of MIoU, mF1 and Accuracy. We use UNet as the basic convolutional neural network, then replace the encoder of UNet with MiT, and then gradually add CBAM, (RCB+AFFM) and PPM modules. Among them, (RCB+AFFM) is a decoder composed of two consecutive residual convolution blocks (RCB) plus three adaptive feature fusion modules (AFFM). The details of the ablation experiment design are shown in Table 2:
[0159] Table 2 Overview of BSD-based ablation experiments
[0160]
[0161]
[0162] Please refer to the attached Figure 8 , Figure 8 Among them, (a) CNN; (b) MiT (Mix Transformer); (c) MiT + CBAM; (d) MiT + CBAM + (RCB + AFFM) and (e) MiT + CBAM + (RCB + AFFM) + PPM; Table 2 shows the evaluation indicators of the ablation experiment. It can be seen that the MIoU, mF1 and Accuracy of the basic convolutional neural network are 93.27, 95.07 and 95.14% respectively, and from Figure 8 It can be seen that UNet presents the worst classification results, reflecting its shortcomings in edge area processing and global accuracy. Then, after replacing the backbone network of UNet with MiT, the various indicators and segmentation accuracy of the model have been improved, indicating that the MiT encoder's efficient multi-head attention mechanism enhances feature extraction, enabling the model to better capture the complex information in remote sensing images in multi-scale scenarios. After further adding the CBAM module, the various evaluation indicators of the model increased by 1.25%, 0.66% and 0.54% respectively. By comparison Figure 8 (b) and (c), it can be seen that the misclassification in (c) is further reduced, and the model presents edge details more carefully. This is because the CBAM module enables the model to focus on important areas more effectively by modeling channel and spatial attention. Subsequently, we replaced the decoder with a new decoder consisting of a residual convolution block (RCB) and an adaptive feature fusion module (AFFM). Figure 8The segmentation results in part (d) show that the model exhibits higher smoothness and consistency in the edge transition area, and the segmentation accuracy of narrow waterways is further improved. Internal errors such as ship misclassification completely disappear, indicating that the decoder composed of RCB and AFFM can effectively fuse and enhance the details of features of different scales in the process of restoring resolution, making the segmentation results more accurate. Finally, the pyramid pooling module (PPM) is added to further utilize multi-scale context information to enhance the global perception ability of the model. In the end, the model is more accurate in the processing of edge details and has higher smoothness. Compared with the basic convolutional neural network, MIoU, mF1 and Accuracy are improved by 3.5%, 2.29% and 2.14% respectively, verifying the effectiveness of each module in the sea and land segmentation task of remote sensing images.
[0163] A new land-sea segmentation model (DA-MiTUNet) provided by the present invention improves the ability of existing models in processing complex scenes and capturing details by integrating multi-scale feature extraction and multiple attention mechanisms. Specifically, DA-MiTUNet uses Mix Transformer (MiT) as an encoder, and introduces a convolutional block attention module (CBAM) and an efficient multi-head attention mechanism (Efficient Multihead Attention) to form a dual attention encoder, which is used to fully combine global context and local information to enhance the ability to capture key areas and edge details. In the decoding stage, residual convolution blocks and adaptive feature fusion modules (AFFM) are introduced to form a decoder with multi-scale information fusion capabilities, which can effectively retain low-level edge information, while gradually integrating semantic features from high levels, enhancing the association between features of different resolutions, and ultimately depicting more accurate coastline details. At the same time, a pyramid pooling module (PPM) is integrated to expand the global receptive field of the model, so that the model can better capture multi-scale context information. Finally, a series of classic semantic segmentation models were compared and evaluated on the GF-HNCD and BSD datasets. The experimental results show that the land-sea segmentation model provided by the present invention has better segmentation results, highlighting its effectiveness. In summary, the DA-MiTUNet model can provide a more reliable scientific basis for the accurate extraction and dynamic monitoring of the coastline, assist in the ecological protection and resource management of coastal areas, and has higher application value in risk assessment, disaster prediction and environmental change monitoring related to extreme sea level events.
[0164] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. The protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A land-sea segmentation method based on a dual attention mechanism, characterized in that: The following steps are involved: S1, acquiring a remote sensing image, and preprocessing the remote sensing image to obtain a coastline remote sensing image to be segmented; S2. Constructing a land-sea segmentation model with dual attention mechanism: S2.1, construct the encoder of dual attention mechanism; S2.2, construct an adaptive feature fusion module; S2.3, constructing a multi-scale information fusion decoder based on the adaptive feature fusion module; S3, training the land-sea segmentation model to obtain a trained land-sea segmentation model; inputting the coastline remote sensing image to be segmented into the trained land-sea segmentation model to obtain a land-sea segmentation result map of the remote sensing image.
2. According to claim 1, a land-sea segmentation method based on a dual attention mechanism is characterized in that: The encoder includes four Transformer Blocks, namely Transformer Block1, Transformer Block2, Transformer Block3 and Transformer Block4; The Transformer Block1, Transformer Block2, Transformer Block3 and TransformerBlock4 are connected in sequence; Each of the Transformer Blocks includes a patch embedding layer, a Transformer layer, a convolutional attention module, and an Overlap Patch Merging layer.
3. The land-sea segmentation method based on the dual attention mechanism according to claim 2 is characterized in that: The Transformer layer includes an efficient multi-head attention mechanism and a hybrid feedforward network; the hybrid feedforward network includes a 1×1 convolution layer, a 3×3 convolution layer, an activation function and a Dropout layer; The number of Transformer layers in the Transformer Block1 is three, the number of Transformer layers in the Transformer Block2 is four, the number of Transformer layers in the Transformer Block3 is six, and the number of Transformer layers in the Transformer Block4 is three.
4. The land-sea segmentation method based on the dual attention mechanism according to claim 3 is characterized in that: The expression of the efficient multi-head attention mechanism is: Q=XW Q ,K=XW K ,V=XW V MultiHead(Q,K,V)=Concat(Z1,Z2,...,Z h )W O X kv =Conv2d(X,kernel_Size=sr_ratio,stride=sr_ratio) OutPut=X+Dropout(ProjDrop(MultiHead(Q,K,V))) Among them, X represents the input feature, Q, K, V represent the query, key and value matrix respectively, and W Q , W K and W V is the corresponding weight matrix; d k is the dimension of the key vector, h is the number of heads, and W O is the output linear transformation matrix.
5. The land-sea segmentation method based on the dual attention mechanism according to claim 2, characterized in that: The expression of the convolutional attention module is: M C (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F)))×F F'=M C (F)×F M S (F)=σ(Conv2d(Concat[AvgPool(F),MaxPool(F)])) F”=M S (F')×F' Among them, F is the input feature, M C is the channel attention weight, F' is the output after the channel attention module, σ represents the Sigmoid activation function, M S is the spatial attention weight, and F' is the final output after the spatial attention module.
6. The method for land-sea segmentation based on dual attention mechanism according to claim 2, characterized in that: The decoder includes a pyramid pooling module, two residual convolution blocks, an adaptive feature fusion module, a 1×1 convolution layer and an upsampling layer; the connection between the two residual convolution blocks is a residual connection; The input of the pyramid pooling module is connected to the output of the Transformer Block4, and the pyramid pooling module is used to extract multi-scale context information from the feature map output by the Transformer Block4; The two residual convolution blocks are used to extract local features from the feature map output by the pyramid pooling module; The adaptive feature fusion module is used to upsample the feature maps output by the two residual convolution blocks and fuse the features of the encoder; The 1×1 convolutional layer is used to map the feature map output by the adaptive feature fusion module into a binary classification sea-land segmentation result map; The upsampling layer is used to align the land-sea segmentation result image with the original input image in terms of spatial resolution.
7. The land-sea segmentation method based on the dual attention mechanism according to claim 6, characterized in that: The adaptive feature fusion module includes dynamic snake convolution, 1×1 convolution, channel attention mechanism based on global average pooling, channel attention mechanism based on global maximum pooling, and reverse residual block; The expression of the adaptive feature fusion module is: S1=Conv2d(Concat(S_high,DSC(S_low))) Among them, DSC represents dynamic snake convolution, Concat represents channel-level concatenation operation, GAP and GMP represent global average pooling and global maximum pooling respectively, MLP is multi-layer perceptron, σ is Sigmoid activation function, It is the feature map after IRB processing.
8. The land-sea segmentation method based on the dual attention mechanism according to claim 7, characterized in that: The number of the adaptive feature fusion modules is three, namely a first adaptive feature fusion module, a second adaptive feature fusion module and a third adaptive feature fusion module; The first adaptive feature fusion module is jump-connected with the Transformer Block1; The second adaptive feature fusion module is jump-connected with the Transformer Block2; The third adaptive feature fusion module is jump-connected to the Transformer Block3, and the outputs of the two residual convolution blocks are connected to the third adaptive feature fusion module; The third adaptive feature fusion module, the second adaptive feature fusion module, and the first adaptive feature fusion module are connected in sequence.
9. The method for land-sea segmentation based on dual attention mechanism according to claim 6, characterized in that: The two residual convolution blocks both include Conv3×3, batch normalization and ReLU activation functions; The expressions of the two residual convolution blocks are: Y=F(X)+X Where X is the input feature map, and F(X) is the output after several layers of convolution, activation, normalization, etc.
Citation Information
Patent Citations
Remote sensing image ocean and non-sea region segmentation method based on pyramid mechanism
CN113870281A
Remote sensing image semantic segmentation method based on double-branch feature fusion
CN115797931A
Adversarial hyperspectral and multispectral remote sensing fusion method
CN116468645A
Remote sensing image sea-land segmentation method and system
CN117726954A
Remote sensing image road extraction method based on multi-level deep and shallow feature matching
CN118230165A
Cited By
Coastline segmentation extraction method and system based on remote sensing data
CN120374987A
Remote sensing coastline automatic extraction method and system based on residual space pyramid segmentation and two-dimensional attention
CN120823520A
TITAN product thunderstorm forecast field correction method and system based on multi-scale integrated learning
CN121071356A
Space environment sensing method, device and equipment suitable for multi-agent system
CN121617096A