Yellow River delta tidal creek intelligent segmentation method based on deep learning

Through the deep learning-based trench segmentation method, combined with the coordinate attention mechanism and the U-Net architecture of multi-scale modules, the problem of time-consuming processing of complex data in traditional models is solved, and efficient and high-precision trench recognition is achieved, especially in complex environments, the segmentation effect of small targets is significant.

CN120298908APending Publication Date: 2025-07-11SHANDONG UNIV OF SCI & TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510261456.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing trench trench segmentation model requires manual selection and extraction of features, and it takes time and is difficult to process complex data. Traditional machine learning models have high requirements for feature engineering, making it difficult to achieve efficient and high-precision trench recognition.

Method used

A deep learning-based method is adopted to build a tidal groove segmentation extraction model, and a multi-scale module based on coordinate attention mechanism is introduced. The tidal groove features are sampled, feature fusion and average pooling are processed through the multi-scale module. Combined with the U-Net architecture optimization model, intelligent semantic segmentation of tidal grooves is automatically realized.

Benefits of technology

The efficiency and accuracy of trench trench recognition are improved, and the model's ability to capture spatial and directional features is enhanced. Especially in complex environments, the segmentation effect of small targets is significant, achieving efficient and high-precision trench recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298908A_ABST
    Figure CN120298908A_ABST
Patent Text Reader

Abstract

The invention discloses a Yellow River delta tidal creek intelligent segmentation method based on deep learning, and mainly relates to the technical field of ocean monitoring, and the method comprises the steps: obtaining a remote sensing image of a target region, carrying out the preprocessing of the remote sensing image, and constructing a tidal creek data set; constructing a tidal creek segmentation and extraction model, training the tidal creek segmentation and extraction model based on the tidal creek data set, and extracting tidal creek features through a model decoder; adding a multi-scale module based on a coordinate attention mechanism in the tidal creek segmentation extraction model, and performing sampling, feature fusion and average pooling processing on the extracted tidal creek features through the multi-scale module to obtain weighted tidal creek features; introducing a loss function to train and optimize the model, outputting a tidal creek semantic segmentation result through a model encoder, verifying the constructed tidal creek segmentation extraction model based on a tidal creek training set, and calculating the model accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of marine monitoring, mainly to the technical field of tidal channel segmentation, and specifically refers to an intelligent segmentation method for tidal channels in the Yellow River Delta based on deep learning. Background Art

[0002] Tidal channels are the most active micro-topographic units on tidal flats, generally showing a reticular distribution with uneven width and depth. Due to tidal action and the erosion and deposition processes of seawater, during the ebb and flow of tides, the flow of seawater will scour the surface and gradually form ditches. The Yellow River Delta region is the only piece of the warm temperate zone in China and even in the world that preserves the most complete, typical and youngest wetland ecosystem. Maintaining its health and stability mainly depends on the important role of tidal channels in supporting hydrological connectivity. In order to better analyze the impact of tidal channels on the coastal ecosystem, it is necessary to analyze the changes in tidal channels with high spatio-temporal resolution. Considering that tidal channels are greatly affected by tides and change rapidly in time and space, it is of great significance to efficiently and accurately extract the spatial distribution of tidal channels for the protection and rational development of resources in coastal intertidal areas.

[0003] In terms of geometric morphology, tidal channels show a meandering linear distribution. The widths of tidal channels are different, but they range from dozens of centimeters in the upper part of the intertidal zone to several kilometers in the lower part. The change in width will have a greater impact on the recognition accuracy. Moreover, tidal channels develop from the lower part of the intertidal zone and flow through different wetland types. The water content, background heterogeneity, sediment content and vegetation coverage of different wetland types are quite different. These characteristics pose great challenges to the accurate segmentation of tidal channels. Given the continuous change of the morphology of tidal channels and the important impact of the change situation on coastal engineering construction and the development of the tidal flat environment, it is urgent to find efficient and high-precision tidal channel recognition methods to realize long-term dynamic monitoring of the tidal flat situation and its evolution process. At present, traditional high-precision tidal channel recognition mainly relies on manual annotation, but this method is relatively dependent on experience and has poor recognition efficiency. With the increase in satellite remote sensing data, in view of the fact that the existing efficiency cannot meet the requirements of spatio-temporal monitoring of the tidal flat area, some traditional model methods and machine learning methods have emerged. However, traditional machine learning models have high requirements for feature engineering and need to manually select and extract features, which may become very time-consuming and difficult when dealing with complex data, and the accuracy is also affected by manual selection.

[0004] Compared with early machine learning models, deep learning can automatically learn high-level abstract features from raw data through multi-layer non-linear processing, reducing the workload of manual feature engineering and being able to process more complex patterns and data. The results of this study for the rapid and efficient extraction of tidal channels in the Yellow River Delta lay a foundation for in-depth understanding of the complex morphological structure and spatio-temporal analysis of tidal channels. Summary of the Invention

[0005] In view of the problem that the existing tidal channel segmentation model requires manual selection and extraction of features, which is time-consuming and difficult to process complex data, the present invention provides an intelligent segmentation method for tidal channels in the Yellow River Delta based on deep learning.

[0006] An intelligent segmentation method for tidal channels in the Yellow River Delta based on deep learning includes:

[0007] S1. Obtain the remote sensing image of the target area, perform image preprocessing on the remote sensing image, and construct a tidal channel data set;

[0008] S2. Construct a tidal channel segmentation and extraction model, train the tidal channel segmentation and extraction model based on the tidal channel data set, and extract tidal channel features through the model decoder;

[0009] S3. Add a multi-scale module based on the coordinate attention mechanism to the tidal channel segmentation and extraction model, perform sampling, feature fusion, and average pooling processing on the extracted tidal channel features through the multi-scale module to obtain weighted tidal channel features;

[0010] S4. The model encoder processes the output image based on the weighted tidal channel features and outputs the tidal channel semantic segmentation result.

[0011] S1 includes:

[0012] S1.1. Obtain the remote sensing image of the target area, perform image preprocessing such as atmospheric correction, radiometric calibration, orthorectification, and image fusion on the remote sensing image according to the characteristics of the target area to improve the quality of the remote sensing image;

[0013] S1.2. According to the characteristics of the tidal channels, manually mark the remote sensing image with a multi-spectral observation resolution of 4 meters of the PMS high-resolution camera to construct a tidal channel data set;

[0014] S1.3. Classify the tidal channel data set according to the type characteristics of the tidal channels in the target area, and divide the tidal channel data set into a salt marsh tidal channel data set and a mudflat tidal channel data set;

[0015] S1.4. Divide the salt marsh tidal channel data set and the mudflat tidal channel data set respectively. Divide the salt marsh tidal channel data set into a salt marsh tidal channel training set and a salt marsh tidal channel test set according to a ratio of 9:1, and divide the mudflat tidal channel data set into a mudflat tidal channel training set and a mudflat tidal channel test set according to a ratio of 9:1;

[0016] S1.5. The salt marsh tidal channel training set and the mudflat tidal channel training set constitute the tidal channel training set, and the salt marsh tidal channel test set and the mudflat tidal channel test set constitute the tidal channel test set.

[0017] S2 includes:

[0018] S2.1. Construct a tidal channel segmentation and extraction model, including a model encoder and a model decoder;

[0019] S2.2. Perform image data conversion processing on the remote sensing images in the tidal channel dataset, convert the remote sensing images from multi-spectral images to RGB images, and perform standardization and normalization processing on the RGB images;

[0020] S2.3. Use the processed RGB images as the input of the tidal channel segmentation and extraction model, train the tidal channel segmentation and extraction model based on the tidal channel training set, and extract tidal channel features through the model encoder and reduce the spatial resolution of the RGB images.

[0021] The model encoder performs four-dimensional multiplication operations on the input RGB images through the full-dimensional dynamic convolution block by dynamically adjusting the weights of the convolution kernels, including the multiplication operation of position information in the spatial dimension, the multiplication operation of channel information in the input channel dimension, the multiplication operation of filters in the output channel dimension, and the multiplication operation of kernels in the convolution kernel spatial dimension. The tidal channel features extracted by the model encoder are:

[0022] y = (α w1 ⊙α f1 ⊙α c1 ⊙α s1 ⊙W1 +... + α wn ⊙α fn ⊙α cn ⊙α sn ⊙W n )x;

[0023] Among them, y is the extracted tidal channel feature; α wi , i = 1, 2, 3…n, represents the position information in the spatial dimension of the i-th image in the tidal channel training set; α fi represents the input channel information of the i-th image in the tidal channel training set; α ci represents the filter of the output channel of the i-th image in the tidal channel training set; α s1 represents the kernel in the convolution kernel space of the i-th image in the tidal channel training set; W i represents the weight of the dynamic convolution kernel corresponding to the i-th image in the tidal channel training set; x represents the input feature map.

[0024] In S3, add a multi-scale module based on the coordinate attention mechanism between the model encoder and the model decoder, send the tidal channel features extracted by the model encoder into a multi-scale module based on the coordinate attention mechanism, and perform multi-scale processing on the extracted features through the multi-scale module.

[0025] Performing multi-scale processing on the extracted features through the multi-scale module includes:

[0026] S3.1, Input the tidal channel features extracted by the model encoder into the multi-scale module. By sampling the input tidal channel features at different sampling rates, extract input features of different scales for the images in the tidal channel training set that have passed through the model encoder, and then fuse the extracted features to obtain the feature extraction result;

[0027] S3.2, Based on the feature extraction result, obtain the width W and height H of the input feature map of the multi-scale module. Perform average pooling on the width W and height H respectively to obtain the average pooling output results of the tidal channel features in the vertical and horizontal directions;

[0028] S3.3, First increase the number of convolution channels in the spatial dimension of the output result, and then perform convolution to compress the channels. Aggregate the output result into complete tidal channel features through batch normalization BN and non-linear activation function Non-linear;

[0029] S3.4, Re-divide the feature vectors of the complete tidal channel features into two-direction feature vectors, re-adjust the number of channels of the two-direction feature vectors through convolution, and then process the two-direction feature vectors through the activation function Sigmoid;

[0030] S3.5, Based on the two-direction feature vectors, perform two-direction weighting on the tidal channel features of the original input multi-scale module of the multi-scale module, multiply the generated weight values with each feature of the original input features of the multi-scale module to obtain the weighted tidal channel features.

[0031] Based on the feature extraction result, the output results of the average pooling of the tidal channel features in the vertical and horizontal directions are:

[0032]

[0033] represents the average pooling result in the horizontal direction, i represents the i-th channel of the input data, that is, the channel number of the input data, h represents the horizontal position processed in the horizontal pooling, that is, the row number processed in the horizontal pooling, x i (h,m) represents the specific value of the input data, m represents the column number processed in the horizontal pooling, ∑ 0≤m<W x i (h,m) represents the sum of the specific values of the input data corresponding to m columns from 0 to W-1, represents the average pooling result in the vertical direction, w represents the column number processed in the vertical pooling, n represents the row number processed in the vertical pooling, and represent the average factor.

[0034] Aggregate the average pooling results obtained in the horizontal and vertical directions through batch normalization (BN) and a non-linear activation function (Non-linear) to obtain the complete tidal channel features as follows:

[0035] f = δ(F1([y h , y w ));

[0036] f represents the obtained complete feature, δ represents the non-linear activation function, and F1 represents the feature transformation and fusion function, which is used to combine and transform the average pooling output results in the horizontal and vertical directions. [y h , y w represents the concatenation operation on the average pooling output result y h in the horizontal direction and the average pooling output result y w in the vertical direction. [] represents the concatenation operation.

[0037] In S4, the model encoder processes the output image based on the obtained weighted tidal channel features. The feature map after being processed by the model encoder passes through a convolutional layer to output the tidal channel semantic segmentation result. Based on the weighted tidal channel features, the model encoder assigns a class label to each pixel of the output image, marks the area where the tidal channel is located as the target class, and marks other areas as the background class to obtain the binary image of the tidal channel segmentation.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] The present invention constructs a tidal channel segmentation and extraction model for the Yellow River Delta region, adds a multi-scale module based on the coordinate attention mechanism to the tidal channel segmentation and extraction model, trains and validates the tidal channel segmentation and extraction model based on the tidal channel dataset, uses the trained tidal channel segmentation and extraction model to perform intelligent semantic segmentation and extraction on salt marsh tidal channels and mud marsh tidal channels respectively, automatically realizes tidal channel recognition and segmentation, and the tidal channel segmentation and extraction model is a deep learning model based on the U-Net deep learning architecture, which is fast and accurate in tidal channel semantic segmentation;

[0040] The tidal channel segmentation and extraction model constructed by the present invention introduces directional depth convolution. On the basis of maintaining the advantages of the U-Net architecture, it further optimizes the performance of the model in terms of detail resolution, enhances the model's ability to capture spatial and directional features, especially enhances the model's ability to capture complex features, and has significant advantages in the extraction of directional features and detail recognition, especially in the segmentation effect of small targets in complex environments. Description of the Drawings

[0041] Figure 1 It is a schematic diagram of the overall network structure of the tidal channel segmentation and extraction model provided by the embodiment of the present invention;

[0042] Figure 2 Schematic diagram of the multi-scale module based on the coordinate attention mechanism provided by the embodiment of the present invention;

[0043] Figure 3 Schematic diagram of the processing flow of the coordinate attention mechanism provided by the embodiment of the present invention;

[0044] Figure 4 Original image of the mudflat tidal channel in the target area provided by the embodiment of the present invention;

[0045] Figure 5 For the model provided by the embodiment of the present invention Figure 4 Binary image of the mudflat tidal channel obtained after intelligent segmentation;

[0046] Figure 6 Original image of the salt marsh tidal channel in the target area provided by the embodiment of the present invention;

[0047] Figure 7 For the model provided by the embodiment of the present invention Figure 6 Binary image of the mudflat tidal channel obtained after intelligent segmentation. Detailed implementation manners

[0048] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0049] An intelligent segmentation method for tidal channels in the Yellow River Delta based on deep learning, comprising:

[0050] S1. Obtain the remote sensing image of the target area, perform image preprocessing on the remote sensing image, and construct a tidal channel dataset;

[0051] S2. Construct a tidal channel segmentation and extraction model, train the tidal channel segmentation and extraction model based on the tidal channel dataset, and extract tidal channel features through the model decoder;

[0052] S3. Add a multi-scale module based on the coordinate attention mechanism to the tidal channel segmentation and extraction model, perform sampling, feature fusion and average pooling processing on the extracted tidal channel features through the multi-scale module to obtain weighted tidal channel features;

[0053] S4. The model encoder processes the output image based on the weighted tidal channel features and outputs the tidal channel semantic segmentation result.

[0054] S1 includes:

[0055] S1.1. Obtain the remote sensing image of the target area, and perform image preprocessing on the remote sensing image, including atmospheric correction, radiometric calibration, orthorectification, and image fusion, according to the characteristics of the target area to improve the quality of the remote sensing image;

[0056] S1.2. According to the characteristics of tidal creeks, manually label the remote sensing image with a multispectral observation resolution of 4 meters from the PMS high-resolution camera to construct a tidal creek dataset;

[0057] S1.3. Classify the tidal creek dataset according to the type characteristics of tidal creeks in the target area, and divide the tidal creek dataset into a salt marsh tidal creek dataset and a mudflat tidal creek dataset;

[0058] S1.4. Divide the salt marsh tidal creek dataset and the mudflat tidal creek dataset respectively. Divide the salt marsh tidal creek dataset into a salt marsh tidal creek training set and a salt marsh tidal creek test set according to a ratio of 9:1, and divide the mudflat tidal creek dataset into a mudflat tidal creek training set and a mudflat tidal creek test set according to a ratio of 9:1;

[0059] S1.5. The salt marsh tidal creek training set and the mudflat tidal creek training set constitute the tidal creek training set, and the salt marsh tidal creek test set and the mudflat tidal creek test set constitute the tidal creek test set.

[0060] S2 includes:

[0061] S2.1. Construct a tidal creek segmentation and extraction model, including a model encoder and a model decoder;

[0062] S2.2. Perform image data conversion processing on the remote sensing images in the tidal creek dataset, convert the remote sensing images from multispectral images to RGB images, and perform standardization and normalization processing on the RGB images;

[0063] S2.3. Use the processed RGB images as the input of the tidal creek segmentation and extraction model, train the tidal creek segmentation and extraction model based on the tidal creek training set, and extract tidal creek features through the model encoder and reduce the spatial resolution of the RGB images.

[0064] Such as Figure 1As shown in the figure, the constructed tidal channel segmentation and extraction model is a U-Net model based on a convolutional neural network. The structure of the model is U-shaped, including a model encoder and a model decoder. The model encoder consists of multiple residual blocks and pooling layers. The dimensions of the feature maps output from the left to the right of the residual blocks are 512×512×32, 256×256×64, 128×128×128, 64×64×256, and 32×32×512 respectively. As the network depth increases, the width and height of the feature maps are halved, and the number of channels of the feature maps doubles. A pooling layer is connected after each residual block, using 2×2 max pooling with a stride of 2 to reduce the spatial dimension while retaining the feature information, that is, downsampling. Each residual block of the encoder consists of two convolutional blocks and a residual connection. Each convolutional block consists of a 3×3 full-dimensional dynamic convolutional block ODConv, batch normalization BatchNorm, and a ReLU activation function. The ReLU activation function is used to increase the non-linear expression ability of the network. The full-dimensional dynamic convolutional block ODConv performs dynamic attention allocation on the feature maps across different feature channels and spatial positions. The weights of the convolutional kernel are dynamically adjusted by the full-dimensional dynamic convolutional block ODConv, and four-dimensional multiplication operations are performed on the input RGB image, including the multiplication operation of position information in the spatial dimension, the multiplication operation of channel information in the input channel dimension, the multiplication operation of the filter in the output channel dimension, and the multiplication operation of the kernel in the convolutional kernel spatial dimension. The tidal channel features extracted by the model encoder are:

[0065] y = (α w1 ⊙α f1 ⊙α c1 ⊙α s1 ⊙W1 +... + α wn ⊙α fn ⊙α cn ⊙α sn ⊙W n )x;

[0066] Among them, y is the extracted tidal channel feature; α wi , i = 1, 2, 3…n, represents the position information in the spatial dimension of the i-th image in the tidal channel training set; α fi represents the input channel information of the i-th image in the tidal channel training set; α ci represents the filter of the output channel of the i-th image in the tidal channel training set; α s1 represents the kernel in the convolutional kernel space of the i-th image in the tidal channel training set; W i represents the weight of the dynamic convolutional kernel corresponding to the i-th image in the tidal channel training set; x represents the input feature map.

[0067] The model decoder consists of multiple 2x2 transposed convolution blocks, and each transposed convolution block contains a transposed convolution layer, batch normalization, and an activation function. A skip connection layer is provided after each transposed convolution block. Through the skip connection, the feature maps of each downsampling step in the encoder are concatenated with the feature maps of the corresponding upsampling step in the decoder to restore the detailed information of the feature maps and obtain a more accurate output.

[0068] In S3, a multi-scale module ASPP based on the coordinate attention mechanism CA is added between the model encoder and the model decoder. The tidal channel features extracted by the model encoder are used as the input Input and fed into a multi-scale module based on the coordinate attention mechanism. The tidal channel features are processed at multiple scales through the multi-scale module based on the coordinate attention mechanism.

[0069] Figure 2 It is a schematic diagram of the multi-scale module structure based on the coordinate attention mechanism provided by the embodiment of the present invention. As Figure 2 shown, the feature maps entering the multi-scale module respectively undergo multiple convolution operations and image pooling operations. The convolution operations include 1×1 convolution operations and 3×3 dilated convolution operations with dilation rates of 6, 12, and 18 respectively. Through the convolution operations, multi-scale tidal channel feature information is extracted to obtain a feature extraction result. The image pooling operation is to perform global average pooling on the feature maps to obtain global context information at the image level. Then the multi-scale module concatenates the feature maps obtained by the convolution operations and the feature maps obtained by the global pooling operation to form a comprehensive feature map containing different scale information, and introduces the coordinate attention mechanism during the concatenation process. A weight map with the same size as the input feature map is generated through the coordinate attention mechanism. By multiplying the input multi-scale feature maps and this weight map pixel by pixel, that is, performing bitwise multiplication, the features in the important regions are strengthened. Then a 1×1 convolution is used to reduce the dimension of the concatenated feature map, readjust the number of channels of the feature vectors, and generate the final output feature map.

[0070] The process of processing the extracted features at multiple scales through the multi-scale module is as Figure 3 shown, including:

[0071] S3.1, input the extracted tidal channel features into the multi-scale module. By sampling the input tidal channel features at different sampling rates, extract input features at different scales for the images in the tidal channel training set that have passed through the model encoder, and then fuse the extracted features to obtain a feature extraction result;

[0072] S3.2. Based on the feature extraction results, obtain the width W, height H, and number of channels C of the input feature map of the multi-scale module. Perform average pooling on the width W and height H respectively to obtain the average pooling output results in the vertical and horizontal directions; XavgPool performs average pooling on the width W, which is equivalent to performing average pooling in the horizontal direction, to obtain the average pooling output result in the horizontal direction, and the size of the output result is C×H×1. YavgPool performs average pooling on the height H, which is equivalent to performing average pooling in the vertical direction, to obtain the average pooling output result in the vertical direction, and the size of the output result is C×1×W;

[0073] S3.3. First, perform concatenation on the average pooling output results in the spatial dimension, which increases the number of convolutional channels. Then, perform convolution to compress the channels. Next, aggregate the output results into complete features through batch normalization BN and the non-linear activation function Non-linear function to obtain a complete feature vector;

[0074] S3.4. Split the complete feature vector, re-divide the complete feature vector into two-direction feature vectors, re-adjust the number of channels of the two-direction feature vectors through convolution, and then pass through the activation function Sigmoid function; the two-direction feature vectors respectively pass through 1×1 convolution and the activation function Sigmoid function;

[0075] S3.5. Based on the two-direction feature vectors, perform two-direction re-weighting on the tidal channel features of the original input multi-scale module of the multi-scale module, multiply the generated weight values with each input feature of the original input feature map of the multi-scale module to obtain the weighted input tidal channel features.

[0076] Based on the feature extraction results, the output results of the average pooling of the tidal channel features in the vertical and horizontal directions are:

[0077]

[0078] represents the average pooling result in the horizontal direction, i represents the i-th channel of the input data, that is, the channel number of the input data, h represents the horizontal position processed in the horizontal pooling, that is, the row number processed in the horizontal pooling, x i (h, m) represents the specific value of the input data, m represents the column number processed in the horizontal pooling, ∑ 0≤m<W x i (h, m) represents the sum of the specific values of the input data corresponding to m columns from 0 to W - 1, represents the average pooling result in the vertical direction, w represents the column number processed in the vertical pooling, n represents the row number processed in the vertical pooling, and represents the average factor.

[0079] Aggregate the average pooling results obtained in the horizontal and vertical directions through batch normalization (BN) and a non-linear activation function (Non-linear) to obtain the complete tidal channel features as follows:

[0080] f = δ(F1([y h , y w ));

[0081] f represents the obtained complete feature, δ represents the non-linear activation function, and F1 represents the feature transformation and fusion function, which is used to combine and transform the average pooling output results in the horizontal and vertical directions. [y h , y w represents the concatenation operation on the average pooling output result y h in the horizontal direction and the average pooling output result y w in the vertical direction. [] represents the concatenation operation.

[0082] In S4, the model encoder processes the output image based on the obtained weighted tidal channel features. The feature map after being processed by the model encoder then passes through a convolutional layer to output the tidal channel semantic segmentation result. Based on the weighted tidal channel features, the model encoder assigns a class label to each pixel of the output image, marks the area where the tidal channel is located as the target class, and marks other areas as the background class to obtain a binary image of the tidal channel segmentation. As Figure 1 described, after being processed by the model encoder, the Atrous Spatial Pyramid Pooling (ASPP) module, and the decoder, and then passing through a convolutional layer composed of a 1×1 convolution and a sigmoid function, it is equivalent to the output layer of the U-Net model. The convolutional layer maps the output of the decoder to a segmentation result with the same size as the input image to output the tidal channel semantic segmentation result. The sigmoid function maps the pixel value to a probability value between 0 and 1, marks the area where the tidal channel is located as the target class, and marks other areas as the background class to obtain a binary image of the tidal channel segmentation.

[0083] The tidal channel segmentation and extraction model ODResU-Net model constructed in the present invention adds a residual block (Resnet Block) and an Atrous Spatial Pyramid Pooling (ASPP) module based on the coordinate attention mechanism to the baseline U-Net model. Based on the salt marsh tidal channel dataset and the mudflat tidal channel dataset, ablation experiments are respectively carried out on the semantic segmentation of mudflat tidal channels and salt marsh tidal channels. Compare the semantic segmentation effects of the baseline U-Net model, the U-Net model with a residual block (Resnet Block) added, and the tidal channel segmentation and extraction model ODResU-Net model. The results of the performance evaluation indicators of each model obtained through the ablation experiment comparison are shown in Table 1:

[0084] Table 1 Comparison results of ablation experiments of different U-Net models

[0085]

[0086] Table 1 shows the performance evaluation metrics of the traditional U-Net model, the U-Net model with residual blocks Resnet Block added, and the tidal channel segmentation and extraction model ODResU-Net model in the mudflat tidal channel area and the salt marsh tidal channel area, including the F1 score, Kappa coefficient, recall rate, and accuracy of the model. As shown in Table 1, the traditional U-Net model performs well in simple scenarios but has limitations in multi-scale and small target recognition. After introducing the ResNet Block module, the model can better learn deep features, and the F1 score and Kappa coefficient have improved, but it is still slightly insufficient in dealing with directionality and spatial structure. The ODResU-Net model introduces directional depth convolution through the multi-scale module ASPP module based on the coordinate attention mechanism. On the basis of maintaining the advantages of the U-Net architecture, it further optimizes the performance of the model in terms of detail resolution, enhances the model's ability to capture spatial and directional features, especially enhances the model's ability to capture complex features, significantly improves the recall rate and F1 score of the model, and makes its segmentation effect on small targets better in complex environments. The ASPP module performs multi-scale processing on spatial information of different scales through dilated convolutions of different sizes, enhancing the network's perception ability for targets of different scales. This module can extract features at different scales of the tidal channel remote sensing image, effectively avoiding the problem of too small receptive field in traditional convolution operations, and improving the recall rate and accuracy of the model.

[0087] Figure 4 is the original input image of the mudflat tidal channel. Based on the tidal channel extraction and segmentation model, the tidal channel is intelligently segmented, and the binary image of the intelligent segmentation of the mudflat tidal channel is obtained as Figure 5 shown. Figure 6 is the original input image of the salt marsh tidal channel. Based on the tidal channel segmentation and extraction model, the tidal channel is intelligently segmented, and the binary image of the intelligent segmentation of the salt marsh tidal channel is obtained as Figure 7 shown.

[0088] The experimental results obtained by ablation experiments on the mudflat tidal channel and salt marsh tidal channel areas using different models are shown in Table 2:

[0089] Table 2 Ablation experiment results of different models in different regions

[0090]

[0091] As shown in Table 2, the PSPNet model mainly focuses on large-scale context information but performs poorly in terms of details and small object segmentation. The low recall rate means that it performs poorly in the detection of a small number of objects. The DeeplabV3+ model enhances multi-scale feature learning through dilated convolution and pyramid pooling, resulting in an improvement in the recall rate and F1 score, but there are still deficiencies in the small object segmentation task. The improved pooling strategy of the PoolFormer model makes it excellent in multi-scale feature extraction, with a significant improvement in the recall rate, but it is still not as strong as ODResU-Net in dealing with small objects. The SegNeXt model enhances the ability to recognize details and small objects through depth convolution and attention mechanism, with a relatively high recall rate and good performance, but there is still a gap in the comprehensive performance compared with ODResU-Net. After introducing directional depth convolution, the ODResU-Net model performs best in dealing with small objects and complex structures, and the F1 score and Kappa coefficient also reach the optimal. The model has significant advantages in the extraction of directional features and detail recognition, especially in complex scenarios.

[0092] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An intelligent segmentation method for tidal creeks in the Yellow River Delta based on deep learning, characterized in that, Including: S1. Obtain the remote sensing image of the target area, perform image preprocessing on the remote sensing image, and construct a tidal channel dataset; S2. Construct a tidal channel segmentation and extraction model, train the tidal channel segmentation and extraction model based on the tidal channel dataset, and extract tidal channel features through the model decoder; S3. Add a multi-scale module based on the coordinate attention mechanism to the tidal channel segmentation and extraction model, perform sampling, feature fusion, and average pooling processing on the extracted tidal channel features through the multi-scale module to obtain weighted tidal channel features; S4. The model encoder processes the output image based on the weighted tidal channel features and outputs the tidal channel semantic segmentation result.

2. The intelligent segmentation method of tidal creeks in the Yellow River Delta based on deep learning according to claim 1, characterized in that S1 includes: S1.

1. Obtain the remote sensing image of the target area, perform image preprocessing such as atmospheric correction, radiometric calibration, orthorectification, and image fusion on the remote sensing image according to the characteristics of the target area to improve the quality of the remote sensing image; S1.

2. According to the characteristics of the tidal channels, manually label the remote sensing image with a multi-spectral observation resolution of 4 meters from the PMS high-resolution camera to construct a tidal channel dataset; S1.

3. Classify the tidal channel dataset according to the type characteristics of the tidal channels in the target area, and divide the tidal channel dataset into a salt marsh tidal channel dataset and a mudflat tidal channel dataset; S1.

4. Divide the salt marsh tidal channel dataset and the mudflat tidal channel dataset respectively. Divide the salt marsh tidal channel dataset into a salt marsh tidal channel training set and a salt marsh tidal channel test set according to a ratio of 9:1, and divide the mudflat tidal channel dataset into a mudflat tidal channel training set and a mudflat tidal channel test set according to a ratio of 9:1; S1.

5. The salt marsh tidal channel training set and the mudflat tidal channel training set constitute the tidal channel training set, and the salt marsh tidal channel test set and the mudflat tidal channel test set constitute the tidal channel test set.

3. The intelligent segmentation method for tidal creeks in the Yellow River Delta based on deep learning according to claim 2, wherein S2 Including: S2.

1. Construct a tidal channel segmentation and extraction model, including a model encoder and a model decoder; S2.

2. Perform image data conversion processing on the remote sensing images in the tidal channel dataset, convert the remote sensing images from multi-spectral images to RGB images, and perform standardization and normalization processing on the RGB images; S2.

3. Use the processed RGB image as the input of the tidal channel segmentation and extraction model, train the tidal channel segmentation and extraction model based on the tidal channel training set, and extract tidal channel features through the model encoder and reduce the spatial resolution of the RGB image.

4. The intelligent segmentation method of tidal creeks in the Yellow River Delta based on deep learning according to claim 3, characterized in that, The model encoder performs four-dimensional multiplication operations on the input RGB image through a full-dimensional dynamic convolution block by dynamically adjusting the weights of the convolution kernels, including position information multiplication operations in the spatial dimension, channel information multiplication operations in the input channel dimension, filter multiplication operations in the output channel dimension, and kernel multiplication operations in the convolution kernel spatial dimension. The tidal channel features extracted by the model encoder are: y = (α w1 ⊙ α f1 ⊙ α c1 ⊙ α s1 ⊙ W1 +... + α wn ⊙ α fn ⊙ α cn ⊙ α sn ⊙ W n )x; Among them, y is the extracted tidal channel feature; α wi , i = 1, 2, 3…n, represents the position information of the spatial dimension of the i-th image in the tidal channel training set; α fi represents the input channel information of the i-th image in the tidal channel training set; α ci represents the filter of the output channel of the i-th image in the tidal channel training set; α s1 represents the kernel of the convolutional kernel space of the i-th image in the tidal channel training set; W i represents the weight of the dynamic convolutional kernel corresponding to the i-th image in the tidal channel training set; x represents the input feature map.

5. The intelligent segmentation method of tidal creeks in the Yellow River Delta based on deep learning according to claim 4, characterized in that, In S3, add a multi-scale module based on the coordinate attention mechanism between the model encoder and the model decoder, send the tidal channel features extracted by the model encoder into a multi-scale module based on the coordinate attention mechanism, and perform multi-scale processing on the extracted features through the multi-scale module.

6. The intelligent segmentation method of tidal creeks in the Yellow River Delta based on deep learning according to claim 5, characterized in that Performing multi-scale processing on the extracted features through the multi-scale module includes: S3.1, Input the tidal channel features extracted by the model encoder into the multi-scale module. By sampling the input tidal channel features at different sampling rates, extract input features of different scales for the images in the tidal channel training set that have passed through the model encoder, and then fuse the extracted features to obtain the feature extraction result; S3.2, Based on the feature extraction result, obtain the width W and height H of the input feature map of the multi-scale module, perform average pooling on the width W and height H respectively, and obtain the average pooling output results of the tidal channel features in the vertical and horizontal directions; S3.3, First increase the number of convolution channels in the spatial dimension of the output result, and then perform convolution to compress the channels. Aggregate the output result into a complete tidal channel feature through batch normalization BN and non-linear activation function Non-linear; S3.4, Re-divide the feature vector of the complete tidal channel feature into two-direction feature vectors, re-adjust the number of channels of the two-direction feature vectors through convolution, and then process the two-direction feature vectors through the activation function Sigmoid; S3.5, Based on the two-direction feature vectors, perform two-direction weighting on the tidal channel features of the original input multi-scale module of the multi-scale module, multiply the generated weight values with each feature of the original input features of the multi-scale module to obtain the weighted tidal channel feature.

7. A method for intelligent segmentation of tidal creeks in the Yellow River Delta based on deep learning according to claim 6, characterized in that, Based on the feature extraction result, the output results of the average pooling of the tidal channel features in the vertical and horizontal directions are: represents the average pooling result in the horizontal direction, i represents the i-th channel of the input data, that is, the channel number of the input data, h represents the horizontal position processed in the horizontal pooling, that is, the row number processed in the horizontal pooling, x i (h, m) represents the specific value of the input data, m represents the column number processed in the horizontal pooling, ∑ 0≤m<W x i (h, m) represents the summation of the specific values of the input data corresponding to m columns from 0 to W - 1, represents the average pooling result in the vertical direction, w represents the column number processed in the vertical pooling, n represents the row number processed in the vertical pooling, and represents the average factor.

8. The intelligent segmentation method of tidal creeks in the Yellow River Delta based on deep learning according to claim 7, characterized in that, Aggregate the average pooling results obtained in the horizontal and vertical directions through batch normalization BN and non-linear activation function Non-linear, and the complete tidal channel feature obtained is: f = δ(F1([y h , y w )); Let \(f\) denote the obtained complete tidal channel features, \(\delta\) denote the non-linear activation function, and \(F1\) denote the feature transformation and fusion function, which is used to combine and transform the output results of average pooling in the horizontal and vertical directions. \([y h ,y w represents the output result \(y h of average pooling in the horizontal direction and the output result \(y w of average pooling in the vertical direction for concatenation operation, and [] represents the concatenation operation.

9. The intelligent segmentation method of tidal creeks in the Yellow River Delta based on deep learning according to claim 8, characterized in that, In S4, the model encoder processes the output image based on the obtained weighted tidal channel features. The feature map after being processed by the model encoder passes through a convolution layer to output the tidal channel semantic segmentation result. Based on the weighted tidal channel features, the model encoder assigns a class label to each pixel of the output image, marks the area where the tidal channel is located as the target class, and marks other areas as the background class to obtain the binary image of the tidal channel segmentation.

Citation Information

Cited By

  • Unmanned aerial vehicle image collapse intelligent identification method based on coordinate attention mechanism

    CN120747802A

  • Unmanned aerial vehicle remote sensing extraction method for strip-shaped intercropping soybean early-stage seedling-lacking plant positions

    CN121482660A

  • A strip intercropping soybean early stage of missing plant position unmanned aerial vehicle remote sensing extraction method

    CN121482660B