Deep learning-based tidal creek segmentation and extraction method

By constructing a deep learning model based on MobileNetV2, combining spectral attention and lightweight convolution, the error and insufficient data in tidal groove extraction in tidal grooves are solved, and high-precision tidal groove segmentation and hydrological analysis are achieved.

CN120471934APending Publication Date: 2025-08-12DALIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510562711.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

When extracting tidal grooves on the tidal flats, the characteristics of the offshore water body are not obvious and similar to the road characteristics, and the characteristics of multi-scale are easily lost. Insufficient data sets lead to insufficient data volume during training.

Method used

A deep learning model based on MobileNetV2 was built, combining spectral attention module and lightweight separable cavity convolution, training the model through cross-entropy loss function, multi-scale feature extraction and segmentation, and vectorization was performed using ArcGIS.

Benefits of technology

It realizes high-precision automated segmentation of tidal grooves on tidal flats, reduces errors, and retains detailed information. It is suitable for tidal groove extraction and hydrological connectivity analysis in complex environments of estuary wetlands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005385061820000021
    Figure BDA0005385061820000021
  • Figure BDA0005385061820000031
    Figure BDA0005385061820000031
  • Figure BDA0005385061820000035
    Figure BDA0005385061820000035
Patent Text Reader

Abstract

The invention discloses a deep learning-based tidal creek segmentation and extraction method, which combines semantic segmentation and remote sensing images, constructs a tidal creek extraction model based on Sentinel-2 multispectral remote sensing and MobileNetV2, and comprises the steps of Sentinel-2 multispectral remote sensing image preprocessing, tidal creek data set making, model construction and testing, and tidal creek network topological graph acquisition. The method is good in performance in the aspects of automatic tidal creek extraction and hydrological quantitative analysis, a new quantitative tool is provided for dynamic monitoring, protection and recovery of estuary wetland hydrological connectivity, and effective management and sustainable utilization of a wetland ecosystem are supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing and target recognition, and specifically is a method for segmenting and extracting tidal flats and gullies based on deep learning. Background Art

[0002] Optical remote sensing imagery, with its advantages of fast update time, easy acquisition, high resolution, and wide image coverage, has been widely used in ecological monitoring and analysis. Automated information extraction from remote sensing images, combined with deep learning, has found widespread application across various industries, including the extraction of tidal flat and gully morphology using optical remote sensing imagery. Optical remote sensing imagery can effectively and quickly separate raster images and vector information from tidal flat and gully structures, providing a reliable tool for hydrological and other related analysis and processing.

[0003] Due to the unique geomorphological characteristics of estuarine wetlands, traditional LiDAR (LiDAR) DEM methods are inadequate for extracting topographic features of large-scale tidal gullies. With the advancement of remote sensing technology, the application of optical remote sensing imagery to extract tidal gully morphology has made significant progress. One approach utilizes image processing techniques to enhance the heterogeneity between tidal gullies and spatial background features, then extracts tidal gully morphology using threshold segmentation. This method offers good accuracy, but requires manual determination of thresholds for different study areas, resulting in spatial and temporal limitations. Another approach, developed in recent years, utilizes data-driven algorithms such as deep learning to achieve automated extraction. Deep learning uses supervised feature learning from manually annotated datasets generated by image annotation to better capture spatial and semantic information within images. Semantic segmentation models based on MobileNetV2 enable multi-scale feature extraction, capturing global context, recovering high-resolution features, and exhibiting strong robustness, making them excellent candidates for remote sensing data rich in detail and at multiple scales.

[0004] Currently, deep learning-based methods for segmenting and extracting mudflat creeks face several major challenges. Offshore creeks contain less water and less distinct water features. Their morphological characteristics are similar to those of roads, making extraction prone to errors. Tidal creeks contain multi-scale features, with smaller features occurring more frequently and easily lost during extraction. Furthermore, existing publicly available remote sensing image datasets for mudflat creeks are limited, resulting in insufficient data for training.

[0005] The Sentinel-2 optical remote sensing satellite image samples of the tidal flat and creek area in the present invention were obtained from the official website of the European Space Agency. Summary of the Invention

[0006] The existing technology has the problem that the offshore tidal gully has less water content and less obvious water body characteristics. The morphological characteristics of this part of the tidal gully are more similar to features such as roads, and extraction is prone to errors; the tidal gully of the mudflat contains multi-scale features, among which smaller features appear more frequently and are easily lost during extraction; among the existing public remote sensing image datasets, there are few datasets of tidal gullies of the mudflat, and the amount of data is insufficient for training.

[0007] In order to solve the above problems, the present invention provides a method for segmenting and extracting tidal flats and gullies based on deep learning, comprising:

[0008] S1: Construct a dataset of tidal flats and gullies;

[0009] Acquire Sentinel-2 optical remote sensing satellite image samples covering the tidal flats and gullies, and preprocess the optical remote sensing satellite image samples, including:

[0010] The geographic information system software ArcGIS was used to perform radiometric calibration, orthorectification, atmospheric correction, and image fusion. The images were then cropped to retain the visible, near-infrared, and far-infrared bands. The retained bands were automatically resampled using the nearest neighbor method to unify the resolution of each band and obtain the input image.

[0011] The input image is used to draw the tidal creek elements using the geographic information system software ArcGIS. By combining line and surface drawing, a vector file of the tidal creek is obtained and converted into a raster file with the same resolution as the optical remote sensing satellite image sample. The raster file with the information is assigned a value of 1, the rest is 0, and the other areas are Nodata, which is used as a model to extract the label image.

[0012] The pixel values of the input image and the label image are normalized by Python and mapped to the range of [0, 1]. The multi-band input image and the single-band label image are randomly cropped into images of a fixed size of 512*512.

[0013] S2: Construct a segmentation model for remote sensing images of tidal flats and gullies;

[0014] The lightweight convolutional neural network MobileNetV2 is used as the model's feature extractor to perform multi-layer feature extraction on the input image, including deep features from late layers, mid-layer features from middle layers, and shallow features from shallow layers.

[0015] A spectral attention module is added before inputting the input image. The channel attention mechanism is used to integrate the NDWI water index and increase the weight of channels related to water characteristics. Green is the green light band and NIR is the near infrared band. The formula is:

[0016]

[0017] A lightweight separable dilated convolution (DASPP) module is used to capture multi-scale contextual information through convolutions with different dilation rates to process deep features. A progressive feature fusion decoder is used to fuse mid-level and shallow-level features, introducing detail information layer by layer. Deconvolution is then used to restore the resolution and generate the final segmentation result. The different dilation rates of the convolutions are 1, 6, 12, and 18.

[0018] The cross entropy loss function is used for training feedback on the final segmentation result, where y and The range is 0-1, representing the true value and predicted value of the pixel respectively, N is the number of predictions, and the formula is:

[0019]

[0020] Among them, L represents the cross entropy, which measures the difference between the model's predicted distribution and the true distribution and guides the update of model parameters;

[0021] S3: Using the dataset consisting of the input image and the label image obtained in step S1, the tidal flat and tidal gully remote sensing image segmentation and extraction model constructed in step S2 is input for multiple rounds of training to obtain the optimal tidal flat and tidal gully segmentation model file. The specific steps include:

[0022] Y°=f decoder (f encoder (X)

[0023] Where X represents the input image, f encoder represents the encoder function, extracting features; f decoder Denotes the decoder function, which generates the segmentation result; Y° represents the output probability value, which assigns a value to each pixel according to the probability value. The final pixel value generated corresponds to the probability distribution result, and then the image segmentation result is obtained;

[0024] The tidal flat remote sensing image segmentation model is represented as an abstract function f, and the input image X is passed through the model f to obtain the prediction result Where θ represents the model parameter, the formula is:

[0025]

[0026] Calculate the predicted value using the cross entropy loss function L in step S2 The difference between the pixel's true value y and the gradient g of the loss function L to the model parameter θ is calculated as follows:

[0027]

[0028] Among them, g and t are the time step and Adam hyperparameters, which are input into the Adam optimizer to obtain the deviation correction value of the model, and are finally used to update the parameter θ;

[0029] After the training is completed, the obtained function f and parameter θ are saved as a model file;

[0030] S4: pre-process the new area image to be processed in step S1, load the model file trained in S3, extract the pre-processed new area image, and obtain a binary raster image of the mudflat tidal gully in the area;

[0031] S5: Use ArcGIS to reclassify the pixel values of the binary raster image of the mudflat tidal gully obtained in S4, perform pixel statistics, and divide the foreground and background. By identifying the strip area of the foreground and vectorizing the center line, a network topology map of the area is obtained. The length and position of the vector represent the data corresponding to this section of the tidal gully.

[0032] In the preferred mode, deep features have rich semantic information and are suitable for global feature analysis.

[0033] In an optimal manner, mid-level features balance semantic information and spatial resolution.

[0034] In the preferred mode, shallow features have a higher resolution and are suitable for capturing local details.

[0035] Beneficial Effects of the Invention: This invention discloses an automatic segmentation technique for mudflat and tidal gully images from Sentinel-2 optical remote sensing images. This technology addresses issues such as incomplete extraction and loss of detail in mudflat and tidal gully extraction within the complex environments of estuarine wetlands. By improving and applying deep learning models, a MobileNetV2-based deep learning model is constructed to extract multi-scale features from optical remote sensing images. This invention has practical application value, providing a standardized process and method for tidal gully extraction and hydrological connectivity analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a flow chart of the method of the present invention;

[0037] Figure 2 This is a model architecture diagram of the present invention;

[0038] Figure 3 This is a diagram of the prediction and extraction results of an example of the present invention. DETAILED DESCRIPTION

[0039] A deep learning-based tidal gully segmentation and extraction method, comprising:

[0040] S1: Construct a dataset of tidal flats and gullies;

[0041] Acquire Sentinel-2 optical remote sensing satellite image samples covering the tidal flats and gullies, and preprocess the optical remote sensing satellite image samples, including:

[0042] The geographic information system software ArcGIS was used to perform radiometric calibration, orthorectification, atmospheric correction, and image fusion. The images were then cropped to retain the visible, near-infrared, and far-infrared bands. The retained bands were automatically resampled using the nearest neighbor method to unify the resolution of each band and obtain the input image.

[0043] The input image is used to draw the tidal creek elements using the geographic information system software ArcGIS. By combining line and surface drawing, a vector file of the tidal creek is obtained and converted into a raster file with the same resolution as the optical remote sensing satellite image sample. The raster file with the information is assigned a value of 1, the rest is 0, and the other areas are Nodata, which is used as a model to extract the label image.

[0044] Since the pixel value range of remote sensing images is large, the computational complexity increases during the calculation process. The pixel values of the input image and the label image are normalized using Python and mapped to the range of [0, 1]. The multi-band input image and the single-band label image are randomly cropped into images of a fixed size. The cropped images are randomly divided into the training set and the test set in a ratio of 7:3. The images are serialized and written into the Tfrecords format using the relevant functions in TensorFlow for model training.

[0045] S2: Construct a segmentation model for remote sensing images of tidal flats and gullies;

[0046] The lightweight convolutional neural network MobileNetV2 is used as the model's feature extractor. The depthwise separable convolution in it can reduce the computational burden brought by high-resolution remote sensing images and perform multi-layer feature extraction on the input image. The deep features extracted from the late layers have rich semantic information and are suitable for global feature analysis; the middle-layer features extracted from the middle layers balance semantic information and spatial resolution; and the shallow features extracted from the shallow layers have higher resolution and are suitable for capturing local details.

[0047] A spectral attention module is added before inputting the input image, and a channel attention mechanism is used to enhance the feature extraction capability of key bands. The NDWI water index is integrated to increase the weight of channels related to water characteristics. Green is the green light band and NIR is the near infrared band. The formula is:

[0048]

[0049] A lightweight separable dilated convolution (DASPP) module is used to capture multi-scale contextual information through convolutions with different dilation rates to process deep features. A progressive feature fusion decoder is used to fuse mid-level and shallow-level features, introducing detail information layer by layer. Deconvolution is then used to achieve up-scaling and restore the resolution to generate the final segmentation result.

[0050] The lightweight separable atrous convolution DASPP module captures multi-scale contextual information through four-scale dilation rate convolution to process deep features and obtain multi-branch outputs. It also combines the global pooling branch to capture contextual information, concatenates the outputs of each branch, and performs convolution and normalization.

[0051] Through the progressive feature fusion decoder, shallow detail features and mid-level features are gradually fused, and high-resolution segmentation results are output after gradual upsampling and feature splicing, which restores high resolution while retaining details and semantic information.

[0052] The cross entropy loss function is used for training feedback on the final segmentation result, where y and The range is 0-1, representing the true value and predicted value of the pixel respectively, N is the number of predictions, and the formula is:

[0053]

[0054] Among them, L represents the cross entropy, which is used to measure the difference between the model prediction distribution and the true distribution and guide the update of model parameters;

[0055] The model uses the Adam optimizer to adaptively adjust the learning rate to promote rapid convergence and stability. To meet the needs of real-time applications, it is optimized through model pruning and quantization techniques, maintaining accuracy while reducing computational complexity.

[0056] S3: Using the dataset consisting of the input image and the label image obtained in step S1, the tidal flat and tidal gully remote sensing image segmentation and extraction model constructed in step S2 is input for multiple rounds of training to obtain the optimal tidal flat and tidal gully segmentation model file. The specific steps include:

[0057] Y°=f decoder (f encoder (X)

[0058] Where X represents the input image, f encoder represents the encoder function, extracting features; f decoder Denotes the decoder function, generating segmentation results; Y ° Represents the output probability value, assigns a value to each pixel according to the probability value, and the final generated pixel value corresponds to the probability distribution result, and then the image segmentation result is obtained;

[0059] The tidal flat remote sensing image segmentation model is represented as an abstract function f, and the input image X is passed through the model f to obtain the prediction result Where θ represents the model parameter, the formula is:

[0060]

[0061] Calculate the predicted value using the cross entropy loss function L in step S2 The difference between the pixel's true value y and the gradient g of the loss function L to the model parameter θ is calculated as follows:

[0062]

[0063] Among them, g and t are the time step and Adam hyperparameters, which are input into the Adam optimizer to obtain the deviation correction value of the model, and are finally used to update the parameter θ;

[0064] The model structure is the content in S2. The data passes through the model structure to extract features, calculate the difference between the predicted results and the true labels using a loss function, and use an optimizer to update the model parameters θ. After training is complete, the resulting function f and parameters θ are saved as a model file. This model file can be used to generate predictions for new, unseen images.

[0065] S4: Perform preprocessing such as cropping on the new area image to be processed, load the model file obtained by training in S3, extract the preprocessed new area image, and obtain a binary raster image of the mudflat tidal gully in the area.

[0066] S5: Use the road centerline method or RivGraph tool to extract the tidal gully network topology map. Through this map, you can get the curvature, nodes, center lines and other effective information that can be directly used for analysis.

[0067] Using ArcGIS, we reclassified the pixel values of the binary raster image of the mudflat tidal gully obtained from S4, calculated the pixel values, and divided the foreground and background into separate images. By identifying the foreground strips and vectorizing their centerlines, we generated a network topology map of the area. The length and position of the vectors represent the data corresponding to that section of the tidal gully. Another method involves processing the image using functions in the RivGraph library, which can also generate a tidal gully network topology map.

[0068] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A deep learning-based tidal gully segmentation and extraction method, characterized in that: include: S1: Construct a dataset of tidal flats and gullies; Acquire Sentinel-2 optical remote sensing satellite image samples covering the tidal flats and gullies, and preprocess the optical remote sensing satellite image samples, including: The geographic information system software ArcGIS was used to perform radiometric calibration, orthorectification, atmospheric correction, and image fusion. The images were then cropped to retain the visible, near-infrared, and far-infrared bands. The retained bands were automatically resampled using the nearest neighbor method to unify the resolution of each band and obtain the input image. The input image is used to draw the tidal creek elements using the geographic information system software ArcGIS. By combining line and surface drawing, a vector file of the tidal creek is obtained and converted into a raster file with the same resolution as the optical remote sensing satellite image sample. The raster file with the information is assigned a value of 1, the rest is 0, and the other areas are Nodata, which is used as a model to extract the label image. The pixel values of the input image and the label image are normalized by Python and mapped to the range of [0, 1]. The multi-band input image and the single-band label image are randomly cropped into images of a fixed size of 512*512. S2: Construct a segmentation model for remote sensing images of tidal flats and gullies; The lightweight convolutional neural network MobileNetV2 is used as the model's feature extractor to perform multi-layer feature extraction on the input image, including deep features from late layers, mid-layer features from middle layers, and shallow features from shallow layers. A spectral attention module is added before inputting the input image. The channel attention mechanism is used to integrate the NDWI water index and increase the weight of channels related to water characteristics. Green is the green light band and NIR is the near infrared band. The formula is: A lightweight separable dilated convolution (DASPP) module is used to capture multi-scale contextual information through convolutions with different dilation rates to process deep features. A progressive feature fusion decoder is used to fuse mid-level and shallow-level features, introducing detail information layer by layer. Upsampling is achieved through deconvolution, and the resolution is restored to generate the final segmentation result. The different dilation rates of convolution are 1, 6, 12, and 18 respectively. The cross entropy loss function is used for training feedback on the final segmentation result, where y and The range is 0-1, representing the true value and predicted value of the pixel respectively, N is the number of predictions, and the formula is: Among them, L represents the cross entropy, which measures the difference between the model's predicted distribution and the true distribution and guides the update of model parameters; S3: Using the dataset consisting of the input image and the label image obtained in step S1, the tidal flat and tidal gully remote sensing image segmentation and extraction model constructed in step S2 is input for multiple rounds of training to obtain the optimal tidal flat and tidal gully segmentation model file. The specific steps include: Y°=f decoder (f encoder (X)) Where X represents the input image, f encoder represents the encoder function, extracting features; f decoder Denotes the decoder function, which generates the segmentation result; Y° represents the output probability value, which assigns a value to each pixel according to the probability value. The final pixel value generated corresponds to the probability distribution result, and then the image segmentation result is obtained; The tidal flat remote sensing image segmentation model is represented as an abstract function f, and the input image X is passed through the model f to obtain the prediction result Where θ represents the model parameter, the formula is: Calculate the predicted value using the cross entropy loss function L in step S2 The difference between the pixel's true value y and the gradient g of the loss function L to the model parameter θ is calculated as follows: Among them, g and t are the time step and Adam hyperparameters, which are input into the Adam optimizer to obtain the deviation correction value of the model, and are finally used to update the parameter θ; After the training is completed, the obtained function f and parameter θ are saved as a model file; S4: pre-process the new area image to be processed in step S1, load the model file trained in S3, extract the pre-processed new area image, and obtain a binary raster image of the mudflat tidal gully in the area; S5: Use ArcGIS to reclassify the pixel values of the binary raster image of the mudflat tidal gully obtained in S4, perform pixel statistics, and divide the foreground and background. By identifying the strip area of the foreground and vectorizing the center line, a network topology map of the area is obtained. The length and position of the vector represent the data corresponding to this section of the tidal gully.

2. The method for segmenting and extracting tidal flats and tidal gullies based on deep learning according to claim 1, characterized in that: Deep features have rich semantic information and are suitable for global feature analysis.

3. The method for segmenting and extracting tidal flats and creeks based on deep learning according to claim 1, characterized in that: Mid-level features balance semantic information and spatial resolution.

4. The method for segmenting and extracting tidal flats and tidal gullies based on deep learning according to claim 1, characterized in that: Shallow features have higher resolution and are suitable for capturing local details.