Remote sensing water body extraction method and device based on U-Net and Transform fusion architecture

The remote sensing water extraction method based on the U-Net and Transformer fusion architecture solves the problems of insufficient multi-source data fusion and temporal modeling capabilities, and achieves high-precision and high-time-efficiency water extraction, generating a spatially detailed and temporally continuous water distribution dataset.

CN120913094AActive Publication Date: 2025-11-07HOHAI UNIV

Patent Information

Application Number
CN202511144553.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-11-07
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

Existing remote sensing water body extraction technologies suffer from problems such as difficulty in multi-source data fusion, insufficient temporal modeling capabilities, poor adaptability due to class imbalance, and limited spatial boundary identification, making it difficult to meet the comprehensive requirements of high timeliness, high accuracy, and strong adaptability.

Method used

A fusion architecture of U-Net and Transformer is adopted. Multispectral remote sensing images and knowledge product data are processed in blocks. Spatial features are extracted by the U-Net encoder, and the Transformer module captures the fusion features of temporal dependence and spatial interest. The binary water body mask map is generated by combining the Sigmoid activation function and the thresholding method, so as to achieve efficient generation of large-scale water body distribution datasets.

Benefits of technology

It has enabled the generation of high-resolution water body datasets with fine spatial detail and continuous temporal sequence, improving the boundary identification capability of small water bodies and complex backgrounds, and meeting the high precision and timeliness requirements of fields such as water resource monitoring and ecological assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913094A_ABST
    Figure CN120913094A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing water body extraction method and device based on a U-Net and Transform fusion architecture, and relates to the technical field of remote sensing image intelligent processing and artificial intelligence, and the method comprises the steps: firstly obtaining a multispectral remote sensing image and knowledge-based product data, and carrying out the blocking, so as to obtain a multispectral image block and a space attention layer block; a U-Net encoder extracts spatial features, a Transform module fuses the spatial and temporal features and focuses on a water body area, and a decoder outputs a probability graph and processes the probability graph to generate a binary mask graph. And finally, splicing the mask graphs to generate a month-by-month data set. According to the method, multi-source data fusion and space-time modeling can be realized, the small water body and boundary identification precision is improved, large-range efficient coverage is realized under high resolution, and the requirements of high precision and high timeliness are met.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent processing of remote sensing images and artificial intelligence, in particular to a remote sensing water body extraction method and device based on a U-Net and Transformer fusion architecture. BACKGROUND

[0002] As one of the important environmental elements on the ground, water body has important research and application value in remote sensing monitoring. Its spatio-temporal dynamic information is widely used in water resource management, flood disaster warning, water ecological environment assessment and other fields. In recent years, with the stable operation and data opening and sharing of medium and high resolution remote sensing satellites such as Landsat series and Sentinel series, it has become possible to obtain remote sensing image data covering a wide range and long time series, which has greatly promoted the development of automatic water body extraction technology.

[0003] Existing water body extraction methods mainly include spectral index methods based on threshold (such as NDWI, MNDWI, etc.), traditional machine learning methods (such as support vector machine, random forest, etc.) and deep learning methods widely studied in recent years. Among them, the convolutional neural network represented by U-Net is widely used in remote sensing image semantic segmentation tasks due to its strong spatial feature extraction capability, and is also used in related research on water body extraction. However, in practical applications, the existing model still generally lacks effective time series modeling capability and multi-source heterogeneous image adaptation capability, making it difficult to meet the current comprehensive needs of remote sensing water body extraction "high timeliness, high precision and strong adaptation". SUMMARY

[0004] The purpose of the present application is to provide a remote sensing water body extraction method and device based on a U-Net and Transformer fusion architecture, which can be applied to large-scale, high-temporal-resolution water body change monitoring tasks, can generate spatially fine, temporally continuous high-resolution water body data sets, and can realize high-precision, high-timeliness water body extraction.

[0005] To achieve the above purpose, the present application provides the following solutions:

[0006] In a first aspect, the present application provides a remote sensing water body extraction method based on a U-Net and Transformer fusion architecture, comprising:

[0007] obtaining multi-spectral remote sensing images of a target region to be extracted in a target month and corresponding knowledge product data, wherein the knowledge product data includes digital elevation model data and river network product vector data;

[0008] performing block processing on the multi-spectral remote sensing images and the knowledge product data to obtain a plurality of image blocks, wherein the image blocks include multi-spectral remote sensing image blocks and spatial attention map layer blocks;

[0009] The image block is taken as input, and a binary water mask graph is output by using a trained water extraction hybrid model, wherein the water extraction hybrid model is a deep learning model composed of a U-Net encoder module, a Transformer module containing a spatial attention mechanism and a U-Net decoder module, the U-Net encoder module is used to extract spatial features of the multispectral remote sensing image block, the Transformer module is used to output a fusion feature sequence that fuses time dependence and spatial attention according to the spatial features and the spatial attention layer block, and the U-Net decoder module is used to output a water distribution probability graph according to the fusion feature sequence; the water distribution probability graph is normalized by a Sigmoid activation function and processed by a threshold method to obtain the binary water mask graph.

[0010] Each of the binary water mask graphs is spliced according to spatial positions to obtain a complete water distribution graph layer covering the target region in a to-be-extracted month.

[0011] The complete water distribution graph layers of each to-be-extracted month are organized in time sequence to generate a monthly water distribution data set of the target region.

[0012] In a second aspect, the present application provides a computer device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the remote sensing water extraction method based on the U-Net and Transformer fusion architecture in the first aspect.

[0013] In a third aspect, the present application provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the remote sensing water extraction method based on the U-Net and Transformer fusion architecture in the first aspect.

[0014] In a fourth aspect, the present application provides a computer program product comprising a computer program, wherein the computer program is executed by a processor to implement the remote sensing water extraction method based on the U-Net and Transformer fusion architecture in the first aspect.

[0015] According to the embodiments provided in the present application, the present application has the following technical effects:

[0016] The application provides a remote sensing water body extraction method and device based on a U-Net and Transformer fusion architecture. First, multispectral remote sensing images and knowledge product data of a target region are acquired, and are subjected to block processing to obtain multispectral image blocks and spatial attention layer blocks, laying a foundation for subsequent feature fusion. Second, a trained hybrid deep learning model is used for inference: a U-Net encoder module extracts spatial detail features of the multispectral image blocks, solving the problem of insufficient spatial feature extraction; a Transformer module with a spatial attention mechanism combines the spatial attention layer blocks and spatial features, captures water body dependency in the time dimension through time series modeling, and guides the model to focus on significant water body areas, taking into account both time series dynamics and spatial attention; a U-Net decoder module outputs a water body distribution probability map based on the fused feature sequence, and generates a binary water body mask map through Sigmoid normalization and thresholding, thereby enhancing the recognition accuracy of small water bodies and boundaries. Finally, the binary water body mask maps output by the blocks are spliced according to spatial positions to form a complete water body distribution map layer, and a monthly data set is generated in chronological order. Through block processing and splicing strategies, high efficiency and coverage of large areas are achieved while ensuring high spatial resolution, balancing efficiency and accuracy.

[0017] In summary, the method realizes effective fusion of multi-source data and collaborative modeling of spatial and temporal features, can accurately capture spatial details and temporal dynamic changes of water bodies, and significantly improves the recognition ability of water body boundaries in small water bodies and complex backgrounds. At the same time, through block processing and splicing strategies, complete coverage of large areas is realized while ensuring high spatial resolution, and the final monthly water body distribution data set has the characteristics of spatial fineness and temporal continuity, meeting the demand for high-precision and high-timeliness water body data in water resource monitoring, ecological assessment and other fields. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0019] Figure 1 A flowchart of a remote sensing water body extraction method based on a U-Net and Transformer fusion architecture provided for Embodiment 1 of the present application;

[0020] Figure 2 A conceptual diagram of a remote sensing water body extraction method based on a U-Net and Transformer fusion architecture provided for Embodiment 1 of the present application;

[0021] Figure 3 A diagram illustrating training performance of a mixed deep learning model of Embodiment 1 of the present application;

[0022] Figure 4 A structural diagram of a computer device provided in Embodiment 2 of the present application. DETAILED DESCRIPTION

[0023] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0024] It is found through research that the existing remote sensing water body extraction technology has the following defects:

[0025] (1) Difficulty in multi-source data fusion: Multi-source remote sensing images have obvious differences in spatial resolution, time distribution, and spectral band structure, and direct fusion can easily cause information inconsistency, affecting the generalization ability of the model.

[0026] (2) Insufficient time series modeling capability: Most methods focus on static analysis of single-time-phase images, making it difficult to capture the dynamic evolution characteristics of water bodies in the time dimension.

[0027] (3) Poor adaptability to class imbalance: In remote sensing images, the water body area usually accounts for a very small proportion, causing a significant class imbalance problem. The existing methods have insufficient discrimination ability for small water bodies and boundary areas, and are prone to misidentification or missed detection.

[0028] (4) Limited spatial boundary recognition: The traditional U-Net structure has limited ability to express detailed edges when facing complex backgrounds (such as the interlaced distribution of water bodies, vegetation, bare land, and urban areas), and it is difficult to balance the extraction efficiency and spatial resolution capability under large-scale coverage in actual deployment.

[0029] Therefore, the present embodiment proposes a remote sensing water body extraction method that can fuse multi-time-phase and multi-source remote sensing data, has local spatial detail extraction and global time series feature modeling capabilities, and has class imbalance adaptability and model updating capability, to improve the stability, accuracy, and intelligent application level of the model in complex remote sensing scenarios.

[0030] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0031] Embodiment 1

[0032] See Figure 1 This embodiment provides a remote sensing water body extraction method based on a U-Net and Transformer fusion architecture, including:

[0033] S1, acquire multispectral remote sensing images of the target area for the month to be extracted and corresponding knowledge product data, wherein the knowledge product data includes digital elevation model data and river network product vector data;

[0034] S2, the multispectral remote sensing image and the knowledge product data are processed into blocks to obtain several image blocks, wherein the image blocks include multispectral remote sensing image blocks and spatial attention layer blocks;

[0035] S3, taking the image patch as input, the trained water body extraction hybrid model outputs a binary water body mask image. The water body extraction hybrid model is a deep learning model composed of a U-Net encoder module, a Transformer module with spatial attention mechanism, and a U-Net decoder module. The U-Net encoder module extracts the spatial features of the multispectral remote sensing image patch. The Transformer module outputs a fusion feature sequence that integrates time dependence and spatial attention based on the spatial features and the spatial attention layer patch. The U-Net decoder module outputs a water body distribution probability map based on the fusion feature sequence. The water body distribution probability map is then normalized using a Sigmoid activation function and processed using a thresholding method to obtain the binary water body mask image.

[0036] S4, each of the binary water body mask images is stitched together according to its spatial location to obtain a complete water body distribution layer covering the target area for the month to be extracted;

[0037] S5. Organize the complete water body distribution layers for each month to be extracted in chronological order to generate a monthly water body distribution dataset for the target area.

[0038] The following is combined Figure 2 The implementation process of the remote sensing water extraction method based on the U-Net and Transformer fusion architecture in this embodiment is described in detail.

[0039] like Figure 2 The remote sensing water extraction method based on the U-Net and Transformer fusion architecture shown includes the training process and application process of the water extraction hybrid model. Specifically, the method includes:

[0040] (1) Obtain 10m spatial resolution multispectral remote sensing images of each target month in the target area within the target year, published 30m spatial resolution JRCGSW water product images, and related knowledge product data.

[0041] The 10-meter spatial resolution multispectral remote sensing image is a monthly scale fused remote sensing image obtained by preprocessing, the preprocessing includes atmospheric correction, cloud detection and cloud removal processing, and median fusion operation of multiple scene images in the same month, and the multispectral includes but is not limited to the following spectral band information which has sensitive response to water body identification: blue light band (center wavelength about 0.45-0.52 μm), green light band (about 0.53-0.59 μm), red light band (about 0.64-0.67 μm), near-infrared band (about 0.85-0.88 μm), short-wave infrared 1 band (about 1.57-1.65 μm) and short-wave infrared 2 band (about 2.10-2.30 μm).

[0042] The 30-meter spatial resolution JRC global surface water (JRC GSW) water body product image is a water body identification time series product published by the Joint Research Centre of the European Union based on Landsat series remote sensing images, which has a monthly time resolution, and the pixels are divided into water body pixels, non-water body pixels and missing pixels according to the monthly water body identification results.

[0043] The knowledge product data is digital elevation model (DEM) data and river network product vector data.

[0044] (2) Preprocessing the multispectral remote sensing image, water body product image and knowledge product data obtained in step (1), and randomly selecting a sample set with complete spatial coverage and continuous time sequence, and dividing it into a training set and a test set according to a certain proportion.

[0045] The preprocessing in step (2) includes converting the river network product vector data into raster data, using bilinear interpolation method to resample the JRC GSW water body product image, DEM data and river network product data to the same 10m spatial resolution as the Sentinel-2 multispectral remote sensing image, eliminating the spatial scale difference, to ensure the spatial consistency of subsequent feature extraction.

[0046] The sample set is randomly selected from the monthly multispectral remote sensing image, which has complete spatial coverage and continuous time sequence. Each sample is composed of three parts: original image sequence, spatial attention map layer, and corresponding real label layer. The spatial size of the image sequence is 128x128 pixels, and the time length is 12 consecutive months; the corresponding tensor dimension is 12xCx128x128, where C represents the number of multispectral channels, the spatial attention map layer is DEM data and river network product data, and the corresponding tensor dimension is 2x128x128, and the real label layer is JRC GSW monthly water body product image, and the corresponding tensor dimension is 12x128x128.

[0047] The ratio of the training set and the test set is 70% training set and 30% test set.

[0048] (3) A hybrid deep learning model structure composed of a U-shaped network (U-Net) model structure and a transformer (Transformer) module containing a spatial attention mechanism is constructed, which receives data to be processed by the U-Net encoder module, is processed in turn by the Transformer module, and is finally output by the U-Net decoder module.

[0049] The hybrid deep learning model composed of the U-shaped network (U-Net) model structure and the transformer (Transformer) module containing a spatial attention mechanism includes:

[0050] 1) U-Net encoder module: spatial feature extraction is performed on each image sequence X t ∈R C×128×128 (t = 1, 2,..., 12), the encoder includes 3 layers of convolutional structure, each layer is composed of a single 3x3 convolution, batch normalization (Batch Normalization), ReLU activation function and 2x2 maximum pooling, the output feature expression of the lth layer is:

[0051]

[0052] wherein The final output expression of the 3rd layer is:

[0053]

[0054] 2) Transformer module with spatial attention mechanism: the spatial attention map information M ∈ R 2×128×128 is processed through 3 layers of 3x3 convolution and 2x2 maximum pooling downsampling operation to obtain spatial guidance features M' ∈ R 64×16×16 , and a spatial attention map is generated by 1x1 convolution and Sigmoid activation function, the expression is:

[0055] A guide = σ(Conv 1×1 (M′));

[0056] Where, σ(·) represents the Sigmoid activation function.

[0057] Each frame of encoded features F t is flattened into a vector z t ∈R D , D = 64x16x16, forming an input sequence matrix:

[0058] Z = [z1, z2, …, z 12 ] ∈ R 12×D ;

[0059] Input the standard Transformer encoder for sequence modeling. In order to guide the attention mechanism to focus on the spatial salient region, a spatial bias term B spatial ∈ R 12×12 is constructed by A guide , which is generated by average pooling, linear transformation, etc., to adjust the weight calculation process in the Transformer self-attention:

[0060]

[0061] where Q, K, V are the query (Query), key (Key), and value (Value) matrices, respectively, and d k is the key vector dimension.

[0062] Finally, the Transformer output is a feature sequence that integrates temporal dependence and spatial attention:

[0063] Z′ = [z′1, z′2, …, z′ 12 ] ∈ R 12×D .

[0064] 3) U-Net decoder module: The multi-frame fusion feature sequence Z′ output by the Transformer is reconstructed into a spatial feature map frame by frame. First, each frame of vector is reconstructed into a spatial feature:

[0065]

[0066] Then input the U-Net decoder, the decoder uses 3 layers of transpose convolution, batch normalization and ReLU activation function for upsampling, and outputs a probability map with a size of 1×128×128:

[0067]

[0068] After normalization by the Sigmoid activation function, a threshold method is used to generate a binary water mask map:

[0069]

[0070] The final output is a sequence of 12 binary water maps.

[0071] (4) using a weighted binary cross-entropy (WBCE) loss function combined with a Dice loss function, training the hybrid deep learning model using the training set to obtain a trained model, i.e., a water body extraction hybrid model.

[0072] The combination of the weighted binary cross-entropy (WBCE) and the Dice loss function as the expression of the total loss function is:

[0073] Loss total =λ1·WBCE+λ2·DicLoss;

[0074] Wherein, λ1 and λ2 are adjustable hyperparameters for controlling the weight of the two types of loss in training, satisfying λ1, λ2≥0; WBCE is the weighted binary cross-entropy loss value; DicLoss is the Dice loss value.

[0075] The weighted binary cross-entropy loss (WBCE) is used to alleviate the problem of extremely unbalanced positive and negative samples, and its expression is:

[0076]

[0077] Wherein, is the water body binary image sequence output by the hybrid deep learning model; M∈{0,1} is the corresponding real label layer in the sample set; β is a class weight factor for increasing the image of the minority class (water body), which is defined as:

[0078]

[0079] Wherein, N pos and N neg are the number of water body and non-water body pixels, respectively.

[0080] The Dice loss function is used to measure the area overlap between the predicted image and the real image, and is defined as:

[0081]

[0082] Wherein, ∈ is a constant smoothing term for avoiding zero denominator, which can effectively enhance the performance of the model in boundary discrimination and small target recognition.

[0083] (5) using the test set to evaluate the classification performance of the water body extraction hybrid model, using F1 score, precision (Precision) and recall (Recall) indicators to verify the accuracy and generalization ability of the model.

[0084] (6) The water body extraction mixed model is applied to the 10-meter spatial resolution multispectral remote sensing image in the target area, a sliding window mode is used for image blocking reasoning, image block splicing reconstruction is realized through overlapping area fusion and edge smoothing, water body distribution information of the target area is extracted, and a 10-meter spatial resolution monthly water body distribution data set is generated.

[0085] The trained model is applied to the 10-meter spatial resolution multi-source remote sensing image in the target area, a sliding window mode is used for image blocking processing, the size of the sliding window is 128x128 pixels, the time length is 12, and the window slides in space with a step of 10 pixels. Each image block includes corresponding multi-band remote sensing image information and corresponding knowledge product data, including DEM data and river network product data, which are input into the trained model as spatial guide layers. The water body distribution prediction results of each image block are spliced and reconstructed according to their spatial positions in the original image to form a water body layer that completely covers the target area; finally, the water body distribution layers corresponding to each time slice are organized in time sequence to generate a 10-meter spatial resolution monthly water body distribution data set.

[0086] The above steps (1)-(6) are realized by: multi-source data fusion: by uniformly resampling to 10m resolution, the spatial difference of multi-source data is solved, and the consistency of fusion is improved; spatio-temporal feature joint modeling: U-Net extracts spatial details, Transformer captures temporal dependence, and combines spatial attention mechanism to focus on key water body areas, improving dynamic monitoring capability; class imbalance optimization: WBCE and Dice loss joint optimization, alleviating the imbalance between positive and negative samples, and enhancing the identification accuracy of small water bodies and boundaries; efficient large-scale coverage: sliding window blocking reasoning realizes large-scale area coverage while maintaining 10m high resolution, balancing accuracy and efficiency.

[0087] The implementation process of the remote sensing water body extraction method based on the U-Net and Transformer fusion architecture in the embodiment is explained below with specific examples.

[0088] The existing 30m spatial resolution Sentinel-2 multispectral remote sensing image of the Yellow River Basin from 2018 to 2021, the 30m spatial resolution JRC GSW monthly water body product image of the Yellow River Basin from 2018 to 2021, the ASTER GDEM V3 digital elevation model data of the Yellow River Basin, and the GRWL river network product data are used to obtain the 10m spatial resolution water body distribution data set of the Yellow River Basin from 2018 to 2021:

[0089] SA, 10m spatial resolution Sentinel-2 multispectral remote sensing images of the Yellow River Basin from 2018 to 2021 are obtained through the Google Earth Engine (GEE) remote sensing cloud platform, including blue, green, red, near-infrared, short-wave infrared 1, and short-wave infrared 2 bands. The 30m spatial resolution JRC GSW monthly water body product images of the Yellow River Basin from 2018 to 2021, ASTER GDEM V3 digital elevation model data, and GRWL river network product data.

[0090] SB, the data obtained by SA is preprocessed, and the GRWL river network product vector data is converted into raster data. The JRC GSW water body product image, ASTER GDEM V3 digital elevation model data, and GRWL river network product data are resampled to the same 10m spatial resolution as the Sentinel-2 multispectral remote sensing image using the bilinear interpolation method. A sample set is constructed, each sample consisting of three parts: the original image sequence, the spatial attention layer, and the corresponding real label layer. The image sequence is randomly selected from the monthly multispectral remote sensing image with complete spatial coverage and continuous time phase. The spatial size of the image sequence is 128x128 pixels, and the time length is 12 consecutive months. The corresponding tensor dimension is 12x6x128x128. The spatial attention layer is the ASTER GDEM V3 digital elevation model data and the GRWL river network product data corresponding to the image sequence, and the corresponding tensor dimension is 2x128x128. The real label layer is the JRC GSW monthly water body product image corresponding to the image sequence, and the corresponding tensor dimension is 12x128x128.

[0091] SC, a hybrid deep learning model structure composed of a U-shaped network (U-Net) model structure and a transformer (Transformer) module containing a spatial attention mechanism is constructed, the model structure receives the data to be processed and is input by the U-Net encoder module, and is processed in turn by the Transformer module containing the spatial attention mechanism, and is finally output by the U-Net decoder module; the U-Net encoder module includes three layers of convolution structure, each layer is composed of a group of 3x3 convolution, batch normalization, ReLU activation function and 2x2 maximum pooling operation in turn, for extracting image features and downsampling layer by layer; the Transformer module containing the spatial attention mechanism includes, first extracting the spatial attention map layer information of the input encoded feature, downsampling through three layers of 3x3 convolution and 2x2 maximum pooling to generate spatial guide features, then generating a spatial attention map through 1x1 convolution combined with Sigmoid activation function; flatten each frame of encoded features into a vector to form an input sequence matrix, input into the standard Transformer encoder to realize sequence modeling, and at the same time, construct a spatial bias term through average pooling and linear mapping method, introduce the self-attention mechanism of Transformer to adjust its attention weight, so as to output a feature sequence that integrates time dependence and spatial attention; the U-Net decoder module includes, multi-frame reconstruction of the fused feature sequence, using three layers of upsampling structure, each layer is composed of transposed convolution, batch normalization and ReLU activation function, realizing spatial resolution recovery, finally outputting a probability map with a size of 1x128x128, after normalization by Sigmoid activation function, generating a binary water mask map by threshold method, and finally outputting a sequence of 12 binary water maps.

[0092] SD, using a weighted binary cross-entropy (WBCE) loss function and a Dice loss function to jointly optimize, training the hybrid deep learning model on the training set, as shown in Figure 3 The AdamW optimizer is used in the training process, the initial learning rate is set to 3e-4, and the cosine decay learning rate scheduling is combined until the error converges to the preset threshold, the training is completed, and the trained model, i.e. the water extraction hybrid model, is obtained.

[0093] SE, the classification performance of the water extraction hybrid model is evaluated using the test set, and the F1 score, precision (Precision) and recall (Recall) indicators are used to verify the model accuracy and generalization ability; the evaluation results show that the F1 score, precision and recall indicators of the filled data are within the expected range, verifying the accuracy and reliability of the filled data.

[0094] SF, the trained model is applied to the Sentinel-2 monthly multi-source remote sensing images of the Yellow River Basin with a spatial resolution of 10 meters. The entire image is processed in a sliding window manner, with the size of the sliding window being 128x128 pixels and the time length being 12. The window slides in space with a step size of 10 pixels. Each image block includes corresponding multi-band remote sensing image information and corresponding knowledge product data, including ASTER GDEM V3 digital elevation model data and GRWL river network product data, which are input into the trained model as spatial guide layers. The water body distribution prediction results of each image block are spliced and reconstructed according to their spatial positions in the original image to form a complete water body layer covering the target area. Finally, the water body distribution layers corresponding to each time slice are organized in chronological order to generate the 2018-2021 Yellow River Basin 10m spatial resolution water body distribution dataset.

[0095] The method provided by the embodiment combines the spatial feature extraction capability of U-Net and the time series modeling advantage of Transformer, realizes fine and dynamic extraction of water body information, enhances boundary perception capability by introducing a spatial attention mechanism, and effectively alleviates the class imbalance problem by using weighted cross-entropy and Dice loss function joint optimization, thereby improving the recognition accuracy of small-scale water bodies and complex background areas. At the same time, the embodiment has a high time update frequency and spatial resolution, can generate high-time-efficiency and monthly-scale water body distribution data products, has good generalization ability and actual application value. The method is suitable for large-scale and high-time-resolution water body change monitoring tasks, and can generate spatially fine and temporally continuous high-resolution water body dataset.

[0096] Embodiment 2

[0097] The computer device provided in the embodiment can be a server or a terminal, and its internal structure diagram can be as shown in Figure 4As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through the system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capability. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the data involved in the U-Net and Transformer fusion architecture based remote sensing water body extraction method. The input / output interface of the computer device is used to exchange information between the processor and the external device. The communication interface of the computer device is used to communicate with the terminal outside through the network connection. The computer program is executed by the processor to implement the U-Net and Transformer fusion architecture based remote sensing water body extraction method in embodiment 1.

[0098] Those skilled in the art can understand that, Figure 4 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement. In one exemplary embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps in each method embodiment described above.

[0099] Embodiment 3

[0100] The embodiment provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the U-Net and Transformer fusion architecture based remote sensing water body extraction method in embodiment 1.

[0101] Embodiment 4

[0102] The embodiment provides a computer program product, which includes a computer program, and the computer program is executed by a processor to implement the U-Net and Transformer fusion architecture based remote sensing water body extraction method in embodiment 1.

[0103] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.

[0104] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing related hardware through a computer program, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments of each method. Among them, any reference to memory, database or other medium used in each embodiment provided by the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. Volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc.

[0105] The database involved in each embodiment provided by the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in each embodiment provided by the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.

[0106] Each technical feature of the above embodiments can be combined arbitrarily. In order to make the description simple, not all possible combinations of each technical feature in the above embodiments are described, but as long as the combination of these technical features does not exist contradictory, it should be considered as the scope of the present application.

[0107] The principles and implementations of the present application are described in the specific examples herein, and the above examples are only used to help understand the method of the present application and its core idea; meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation and application range will be changed. In summary, the content of the specification should not be understood as a limitation of the present application.

Claims

1. A remote sensing water body extraction method based on a U-Net and Transformer fusion architecture, characterized in that, The remote sensing water body extraction method based on the U-Net and Transformer fusion architecture comprises: acquiring multi-spectral remote sensing images of a target region in a to-be-extracted month and corresponding knowledge product data, wherein the knowledge product data comprises digital elevation model data and river network product vector data; performing block processing on the multi-spectral remote sensing images and the knowledge product data to obtain a plurality of image blocks, wherein the image blocks comprise multi-spectral remote sensing image blocks and spatial attention map layer blocks; inputting the image blocks to output a binary water body mask image by using a trained water body extraction hybrid model, wherein the water body extraction hybrid model is a deep learning model composed of a U-Net encoder module, a Transformer module with a spatial attention mechanism and a U-Net decoder module, the U-Net encoder module is used to extract spatial features of the multi-spectral remote sensing image blocks, the Transformer module is used to output a fusion feature sequence that fuses time dependence and spatial attention according to the spatial features and the spatial attention map layer blocks, and the U-Net decoder module is used to output a water body distribution probability map according to the fusion feature sequence; the water body distribution probability map is normalized by a Sigmoid activation function and processed by a threshold method to obtain the binary water body mask image; splicing each binary water body mask image according to spatial positions to obtain a complete water body distribution layer covering the target region in the to-be-extracted month; organizing the complete water body distribution layers of each to-be-extracted month in chronological order to generate a monthly water body distribution data set of the target region.

2. The method of claim 1, wherein the U-Net and Transformer fusion architecture-based remote sensing water body extraction method is characterized by, The training process of the water body extraction hybrid model comprises: acquiring an initial sample data set, wherein the initial data set comprises multi-spectral remote sensing image sample data, water body product image sample data and knowledge product sample data of each target month in a target year of a target region; preprocessing the initial sample data set to obtain a training set; training the water body extraction hybrid model based on the training set by using a joint loss function until the water body extraction hybrid model converges, wherein the joint loss function comprises a weighted binary cross-entropy loss function and a Dice loss function.

3. The method of claim 1, wherein the U-Net and Transformer fusion architecture-based remote sensing water body extraction method is characterized by, The U-Net encoder module comprises three convolution structures, and each convolution structure comprises a convolution layer, a batch normalization layer, a ReLU activation function layer and a max-pooling layer.

4. The method of claim 1, wherein the U-Net and Transformer fusion architecture-based remote sensing water body extraction method is characterized by, The Transformer module comprises a spatial guide feature extraction layer, a spatial attention map generation layer, a sequence feature flattening layer, a spatial bias term construction layer and a Transformer encoder layer; the spatial guide feature extraction layer is used to perform three convolution operations on the spatial attention map layer block, each convolution operation is combined with max-pooling for down-sampling to generate spatial guide features; the spatial attention map generation layer is used to adjust channels of the spatial guide features by a convolution layer and generate a spatial attention map by combining a Sigmoid activation function layer to mark water body significant areas; The sequence feature flattening layer is configured to flatten each frame of spatial features output by the U-Net encoder into a vector to form an input sequence matrix containing multi-frame temporal information; The spatial bias term construction layer is configured to generate a spatial bias term based on the spatial attention map to adjust the self-attention weight; The Transformer encoder layer is configured to adopt a standard Transformer encoder structure to model the input sequence matrix in time by introducing a self-attention mechanism of the spatial bias term, and output a fusion feature sequence that fuses time dependence and spatial attention.

5. The method of claim 1, wherein the U-Net and Transformer fusion architecture-based remote sensing water body extraction method is characterized by, The U-Net decoder module includes a feature reconstruction layer and a 3-layer upsampling structure, each layer of the upsampling structure being composed of a transposed convolution, a batch normalization and a ReLU activation function; The feature reconstruction layer is configured to reshape the fusion feature sequence into a spatial feature map frame by frame, and restore the spatial feature to have the same dimension as the output of the U-Net encoder; The 3-layer upsampling structure is configured to expand the spatial size of the spatial feature map by layer-by-layer upsampling, and output a water distribution probability map consistent with the spatial size of the multi-spectral remote sensing image.

6. The method of claim 2, wherein the U-Net and Transformer fusion architecture-based remote sensing water body extraction method is characterized by, The expression of the joint loss function is: Loss total = λ1·WBCE + λ2·DicLoss; wherein, Loss total is a joint loss function value; λ1 and λ2 are adjustable hyperparameters, satisfying λ1, λ2≥0; WBCE is a weighted binary cross-entropy loss value; DicLoss is a Dice loss value; The expression of the weighted binary cross-entropy loss function is: wherein, is a binary water mask map; M e {0, 1} is the corresponding ground truth label layer in the initial sample data set; β is a class weight factor; The expression of the Dice loss function is: where ∈ is a constant smoothing term.

7. The method of claim 2, wherein the U-Net and Transformer fusion architecture-based remote sensing water body extraction method is characterized by, The preprocessing includes converting river network product vector sample data into river network product raster data, and using bilinear interpolation to resample water body product image sample data, digital elevation model sample data and river network product raster data until the spatial resolution of the water body product image sample data, the digital elevation model sample data and the river network product raster data is consistent with the spatial resolution of the multi-spectral remote sensing image sample data.

8. A computer device comprising: A memory, a processor and a computer program stored on the memory and executable on the processor, characterized in that the processor executes the computer program to implement the remote sensing water body extraction method based on the U-Net and Transformer fusion architecture according to any one of claims 1-7.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the remote sensing water body extraction method based on the U-Net and Transformer fusion architecture according to any one of claims 1-7.

10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the remote sensing water body extraction method based on the U-Net and Transformer fusion architecture according to any one of claims 1-7.

Citation Information

Patent Citations

  • CNN-Transform-based remote sensing image water body extraction method and device, electronic equipment and medium

    CN117036805A

  • SAR (Synthetic Aperture Radar) water body extraction method based on deep learning

    CN118865141A

  • Water body extraction method and device based on multi-source multi-scale optical remote sensing data

    CN120411810A

  • Method and system of extraction of impervious surface of remote sensing image

    US20200026953A1

  • Fine water body extraction method based on u-net neural network

    WO2022083202A1

Cited By

  • Low-altitude remote sensing ground feature element extraction method fused with DSM side adapter

    CN121661547A

  • Low-altitude remote sensing feature extraction method with DSM side edge adapter

    CN121661547B

  • Urban flood mapping and driving factor analysis method based on multi-source remote sensing data

    CN121861507A