An intelligent identification method for black and odorous water bodies based on multi-source remote sensing images

By combining multi-source remote sensing images with deep learning methods, high-resolution images are used to generate water body masks, multispectral data is used to calculate the black and odorous water body index, and high-frequency remote sensing images are used for confirmation. This solves the problems of accuracy and false alarm rate in identifying black and odorous water bodies in complex urban environments, and enables accurate identification and efficient supervision of small water bodies.

CN121459185BActive Publication Date: 2026-05-01SHANGHAI WEIXING DATA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI WEIXING DATA TECH CO LTD
Filing Date
2025-12-26
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing remote sensing identification technologies for black and odorous water bodies rely on a single data source, resulting in limited spatiotemporal resolution. This makes it difficult to accurately identify black and odorous water bodies in complex urban environments and easily misidentifies building shadows, dark road surfaces, etc., as black and odorous water bodies, leading to low identification accuracy and a high false alarm rate.

Method used

This study employs multi-source remote sensing imagery combined with deep learning methods. It leverages the spatial geometric advantages of high-resolution imagery to generate a global water body mask, calculates the black and odorous water body index using multispectral data, and conducts detailed verification using high-frequency remote sensing imagery. It utilizes a hybrid network structure combining the Transformer mechanism and the U-Net architecture to improve recognition accuracy, and combines adaptive thresholds and geometric filtering rules to eliminate interference.

Benefits of technology

It has achieved accurate location and identification of small and micro black and odorous water bodies in cities, reduced the false alarm rate, maintained the topological connectivity and spatial detail integrity of water body extraction, and improved the accuracy and efficiency of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121459185B_ABST
    Figure CN121459185B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of remote sensing data processing and water environment monitoring, and discloses a kind of black and odorous water intelligent identification method based on multi-source remote sensing image, the method first obtains first high-resolution remote sensing image, extracts global water mask using TransUNet model;Second, based on geometric features and basic spectral rules, non-target patches are removed;Then, the black and odorous water index is calculated using the second multispectral remote sensing image, and the suspected black and odorous water is screened by combining adaptive threshold segmentation;Further, the suspected target is finely classified and confirmed using the third high-frequency remote sensing image and Efficient Net model;Finally, the identification results are fused and output.The present application overcomes the limitations of insufficient spatial and temporal resolution of a single data source through multi-source data collaboration and coarse screening-fine classification multi-level strategy, improving the positioning accuracy and anti-interference ability of urban black and odorous water identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing data processing and water environment monitoring technology, specifically to a method for intelligent identification of black and odorous water bodies based on multi-source remote sensing images. Background Technology

[0002] Urban black and odorous water bodies are an extreme form of water pollution that arises alongside urbanization. They not only severely damage the urban living environment and harm residents' health, but also hinder sustainable urban development. With increased efforts in water environment management, the monitoring of black and odorous water bodies has shifted from fixed-point monitoring to comprehensive, grid-based monitoring. While traditional ground-based manual sampling and monitoring methods offer high accuracy, they are time-consuming, labor-intensive, and have limited coverage, making them unsuitable for large-scale, high-frequency dynamic monitoring. Therefore, utilizing remote sensing technology for the identification of black and odorous water bodies has become the mainstream approach.

[0003] Existing remote sensing technologies for identifying black and odorous water bodies primarily rely on spectral feature analysis of optical remote sensing images. The common approach is to use multispectral satellite imagery and construct specific ratio indices to distinguish between normal and black / odorous water bodies. However, this traditional method, dependent on a single data source, faces numerous challenges in complex urban environments. Firstly, limited by sensor hardware performance, achieving both spatial and spectral resolution in remote sensing images is often difficult. While low- to medium-resolution multispectral images offer rich bands and can reflect water quality characteristics, severe pixel mixing makes it difficult to accurately extract narrow tributaries, culverts, or broken ponds that are widespread in cities, easily leading to missed detections and blurred boundaries of small water bodies. Secondly, while high-resolution images can clearly present the geometry of ground features, they typically lack key bands for retrieving water quality parameters, making them unsuitable for direct determination of black and odorous water body attributes.

[0004] Furthermore, urban surface environments are highly fragmented and complex, containing numerous interfering features with spectral characteristics similar to those of polluted water bodies. For example, the shadows of high-rise buildings, dark asphalt pavements, and dark areas formed by tree canopies often exhibit extremely high similarity in spectral response to polluted water bodies—a phenomenon known as "heterogeneous objects sharing the same spectrum." Existing technologies primarily focus on threshold segmentation or shallow machine learning classification along the spectral dimension, lacking in-depth analysis of texture, shape, and spatial context information. This makes it highly susceptible to misclassification of these non-water bodies as polluted water bodies, resulting in a persistently high false alarm rate. Simultaneously, when extracting basic water bodies, conventional convolutional neural network models, due to their limited receptive field, struggle to capture global dependencies. When waterways are obstructed by bridges or roadside trees, they often fail to maintain topological connectivity, leading to fragmented water network structures that are unsuitable for refined monitoring data requirements.

[0005] Therefore, there is an urgent need for an intelligent identification method for black and odorous water bodies that can synergistically utilize the advantages of multi-source remote sensing data, take into account both spatial geometric accuracy and water quality spectral characteristics, and effectively eliminate interference from complex environments. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides an intelligent identification method for black and odorous water bodies based on multi-source remote sensing imagery. This method solves the problems of existing black and odorous water body identification methods relying on a single remote sensing data source, which limits the spatiotemporal resolution and makes it easy to misidentify objects with similar spectral characteristics, such as building shadows and dark road surfaces, as black and odorous water bodies in complex urban environments, resulting in low identification accuracy and high false alarm rate.

[0007] To achieve the above objectives, this invention employs the following technical solution: This invention proposes an intelligent identification method for black and odorous water bodies based on multi-source remote sensing imagery. This method first leverages the spatial geometric advantages of high-resolution remote sensing data, using the first high-resolution remote sensing image as a spatial positioning reference. A pre-trained semantic segmentation model performs global inference on the image, obtaining pixel-level water body probability distributions and generating a global vector water body mask. This process aims to utilize the texture details of sub-meter-level imagery to address the problems of river channel discontinuities and the omission of small water bodies caused by mixed pixels in low- and medium-resolution imagery.

[0008] Based on the acquisition of the spatial distribution of water bodies across the entire region, this method introduces a cascaded filtering mechanism based on prior knowledge. Filtering rules are constructed using the geometric and basic spectral features of the global vector water body mask. Before performing complex black and odor identification, large lakes and reservoirs, as well as obviously clean water bodies, are eliminated. This aims to compress the search space for subsequent processing, focusing on key areas suspected of pollution, thereby improving the overall efficiency of the algorithm.

[0009] To address the limitations of high-resolution images, which have limited bands and make it difficult to directly retrieve water quality parameters, this method further introduces multispectral data to supplement the spectral dimensions. A second multispectral remote sensing image, acquired synchronously with the first high-resolution image, is resampled and registered to the water body region to be identified. The rich band information of this image is used to calculate the black and odorous water index, and an adaptive threshold segmentation algorithm is used to automatically determine the segmentation threshold, thereby filtering out suspected targets with black and odorous spectral characteristics. This step utilizes the principle of heterogeneous data complementarity: high-resolution images are used to identify outlines, while multispectral images are used to distinguish colors.

[0010] To further eliminate spectral confusion caused by shadows, dark sediment, or road surfaces, this method incorporates high-revisit-cycle third-frequency remote sensing imagery and a deep learning classification model for fine-tuning. Centering on the identified suspected patches, image blocks are cropped from the third-frequency remote sensing imagery and input into the classification model. Utilizing the deep neural network's ability to extract deep water texture and contextual features, the final probability of blackening and odor is output. When the probability exceeds a set threshold, the water body is confirmed as black and odorous and output accordingly.

[0011] Regarding the specific construction of the semantic segmentation model: To improve the segmentation accuracy of narrow river channels and fragmented water bodies, this method adopts a hybrid network structure that deeply integrates the Transformer mechanism and the U-Net architecture. This structure introduces a self-attention mechanism module in the encoding path, using the Transformer to capture long-distance dependencies in the image and solve the problem of discontinuity in the river channel under shadow occlusion; at the same time, the classic skip connection structure of U-Net is retained in the decoding path, fusing the shallow high-resolution features with the deep semantic features to accurately restore the shoreline edge details of the water body.

[0012] Regarding the optimization of the loss function during model training: To address the severe imbalance in the ratio of foreground to background pixels in water segmentation, this method employs a hybrid loss function strategy during the model training phase. This strategy combines a region consistency metric and a pixel classification probability metric. The region consistency metric focuses on the spatial overlap between the predicted mask and the ground truth label, is insensitive to target size, and can prevent small water bodies from being ignored during training. The pixel classification probability metric, based on the assumption of independent and identically distributed pixels, constrains the classification accuracy of each pixel, ensuring the precision of edge segmentation.

[0013] Regarding the specific implementation of geometric feature filtering: To accurately eliminate interfering targets, this method designs specific geometric fingerprint features. On the one hand, large-scale water bodies (such as reservoirs) are directly filtered out based on area attributes; on the other hand, a density index based on the ratio of area to the square of perimeter is constructed. This index can quantify the shape complexity and elongation of water body patches, thereby effectively identifying and eliminating main waterways, allowing the algorithm to focus on urban tributaries and culverts, areas prone to black and odorous water pollution.

[0014] Regarding the construction principle of the black and odorous water body index: The black and odorous water body index used in this method is constructed based on a specific band differential response principle. Specifically, the first normalized difference calculation is performed using the green and red-edge bands of the second multispectral image to reflect the characteristics of chlorophyll and suspended matter in the water body; the second normalized difference calculation is performed using the shortwave infrared and blue bands to reflect the absorption characteristics of organic pollutants in the water body. Multiplying the normalized results of the two methods amplifies the spectral differences between black and odorous water bodies and ordinary water bodies in specific bands, thereby improving the sensitivity of identification.

[0015] Regarding adaptive thresholding and classification models: For threshold selection, this method abandons manual empirical setting and adopts a statistical strategy based on maximum inter-class variance (Otsu) to automatically find the optimal gray level that maximizes the difference between the foreground and background, enhancing the algorithm's robustness under different lighting conditions and seasons. In the final classification stage, a convolutional neural network architecture based on a composite scaling strategy is employed. By uniformly adjusting the network's depth, width, and resolution, more discriminative deep features are extracted while controlling computational load, ensuring the accuracy of the recognition results.

[0016] This invention provides a method for intelligent identification of black and odorous water bodies based on multi-source remote sensing imagery. It has the following beneficial effects:

[0017] 1. This invention utilizes the sub-meter spatial resolution advantage of a first high-resolution remote sensing image (such as GF-2), combined with the TransUNet model to extract a global water body mask, and combines the red edge and shortwave infrared bands of a second multispectral remote sensing image (such as Sentinel-2) to calculate the black and odorous water body index. This overcomes the boundary blurring problem caused by mixed pixels in medium and low resolution images. While preserving the geometric boundaries of narrow river channels and broken pits with a width of less than 5 meters, it uses multispectral features to complete the water quality qualitative analysis, achieving accurate location and identification of small black and odorous water bodies in cities.

[0018] 2. In the semantic segmentation stage, this invention adopts a hybrid network structure that deeply integrates the Transformer mechanism and the U-Net architecture. By introducing a self-attention mechanism into the encoding path to capture global long-distance dependencies, it can effectively deal with the disconnection of river channels caused by tree occlusion or shadows and maintain the topological connectivity of water body extraction. At the same time, it uses the skip connections of the decoding path to restore local spatial details, and with the constraint of pixel classification distribution by the hybrid loss function, it improves the segmentation completeness of the model for irregular water body targets in complex urban backgrounds.

[0019] 3. This invention utilizes geometric compactness index and Otsu adaptive threshold to pre-remove large clean water bodies and obviously non-black and odorous targets, which greatly reduces the amount of data in subsequent processing. Then, it introduces third high-frequency remote sensing imagery (such as Planet) and EfficientNet classification model, and uses high-resolution texture features to perform secondary confirmation of suspected targets, effectively eliminating interference objects with similar spectral features such as cloud shadows, building shadows and dark road surfaces, and reducing the false alarm rate of the identification results. Attached Figure Description

[0020] Figure 1 This is a flowchart of the method of the present invention;

[0021] Figure 2 This is a structural block diagram of the intelligent identification system for black and odorous water bodies according to the present invention. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] See attached document Figure 1 and attached Figure 2 This invention provides an intelligent identification method for black and odorous water bodies based on multi-source remote sensing imagery. This method can be executed by computer equipment or a black and odorous water body identification system deployed on a server. Logically, the system may include a data acquisition module, an image preprocessing module, a semantic segmentation module, a rule filtering module, a spectral analysis module, a deep classification module, and a result fusion module. The method specifically includes the following steps:

[0024] Step S1: The data acquisition module acquires the first high-resolution remote sensing image of the target monitoring area. After the image preprocessing module preprocesses the image, the semantic segmentation module inputs the image into the pre-trained semantic segmentation model for inference, obtains the global water body probability map, and generates a global vector water body mask.

[0025] Step S2: The rule-based filtering module constructs filtering rules based on the geometric and spectral features of the global vector water body mask, removes large lakes and reservoirs and non-black and odorous water body patches, and obtains the set of water bodies to be identified.

[0026] Step S3: The data acquisition module acquires a second multispectral remote sensing image that matches the imaging time window of the first high-resolution remote sensing image. The image preprocessing module resamples and registers the spatial resolution of the second multispectral remote sensing image to the spatial range of the water body set to be identified. The spectral analysis module calculates the black and odorous water body index based on the second multispectral remote sensing image and uses an adaptive threshold segmentation method to screen out a candidate set of suspected black and odorous water bodies.

[0027] Step S4: The data acquisition module acquires the third high-frequency remote sensing image that matches the time window of the candidate set of suspected black and odorous water bodies. The image preprocessing module crops the image blocks with the patches in the candidate set of suspected black and odorous water bodies as the center. The deep classification module inputs the image blocks into the pre-trained deep learning classification model for inference and outputs the confidence of black and odorous water bodies.

[0028] Step S5: When the confidence level of the black and odorous water body is greater than the set threshold, the result fusion module confirms that the corresponding patch is a black and odorous water body and outputs the identification results by spatial superposition.

[0029] The following section will elaborate on the acquisition and preprocessing of multi-source remote sensing data, based on the above procedures.

[0030] In this embodiment, in order to take into account the spatial details, spectral features and temporal resolution required for the identification of black and odorous water bodies, the data acquisition module is configured to collect three different types of satellite remote sensing data sources, namely, a first high-resolution remote sensing image, a second multispectral remote sensing image and a third high-frequency remote sensing image.

[0031] For the first high-resolution remote sensing image, image data from the Gaofen-2 (GF-2) satellite was selected. The GF-2 satellite carries two high-resolution cameras, capable of providing panchromatic bands with a spatial resolution better than 1 meter and multispectral bands with a spatial resolution of 3.2 meters. The band composition of the first high-resolution remote sensing image includes at least the blue, green, red, and near-infrared bands. The data acquisition module acquires GF-2 L1A or L2A standard products for the target monitoring area.

[0032] For the second multispectral remote sensing image, Sentinel-2 satellite imagery data was selected. The Sentinel-2 satellite carries a Multispectral Imager (MSI), covering 13 spectral bands with a swath width of 290 km. In this embodiment, the second multispectral remote sensing imagery includes visible light bands (blue, green, and red), red-edge bands, and shortwave infrared bands used for calculating spectral indices. The data acquisition module acquires Sentinel-2L2A-level atmospheric sub-atmosphere reflectance products matching the imaging time window of the first high-resolution remote sensing imagery, with the time window deviation controlled within ±7 days to ensure the consistency of ground feature spectral characteristics.

[0033] For the third high-frequency remote sensing imagery, PlanetScope satellite constellation imagery data was selected. The PlanetScope constellation consists of a large number of nanosatellites, possessing an extremely high revisit period. The spatial resolution of the third high-frequency remote sensing imagery is 3 to 4 meters, and after fusion processing, it can be better than or equal to the spatial resolution of the first high-resolution remote sensing imagery, and includes blue, green, red, and near-infrared bands. The data acquisition module acquires PlanetScope orthophoto products (Level 3B) that match the time window of the candidate set of suspected black and odorous water bodies, with the time window deviation controlled within ±30 days.

[0034] The image preprocessing module performs standardized preprocessing operations on the acquired multi-source remote sensing data to eliminate the influence of atmospheric, lighting, and geometric distortions on the recognition results.

[0035] For the first high-resolution remote sensing image, the image preprocessing module first performs radiometric calibration, converting the image's raw digital quantization values ​​into radiance; then atmospheric correction is performed to eliminate atmospheric scattering and absorption effects and obtain the true surface reflectance; next, orthorectification is performed using the RPC model and ground control points to eliminate geometric distortion caused by topographic relief; finally, the panchromatic band and multispectral band are fused to generate a standard orthorectified image with a spatial resolution of 0.8 meters and containing 4 spectral bands.

[0036] For the second multispectral remote sensing image, the image preprocessing module resamples each band of the L2A-level product. Specifically, the red-edge band and shortwave infrared band are resampled to a spatial resolution of 10 meters to maintain consistency with the spatial resolution of the visible and near-infrared bands. Subsequently, based on the geographic coordinate system of the first high-resolution remote sensing image, spatial registration is performed on the resampled second multispectral remote sensing image, with the registration error controlled within 0.5 pixels.

[0037] For the third high-frequency remote sensing image, since it is a Level 3B orthophoto product, the image preprocessing module mainly performs image mosaicking and cropping operations to ensure that the image covers the spatial range of the candidate set of suspected black and odorous water bodies.

[0038] To meet the training and inference requirements of subsequent deep learning models, the image preprocessing module is also configured to construct a standardized dataset. Specifically, a sliding window technique is used to crop large-format remote sensing imagery with a fixed step size, generating image tiles with fixed pixel dimensions. During training, data augmentation processing is performed on the image tiles, including random flipping, rotation, and noise addition, to improve the model's generalization ability. During inference, the geographic coordinate information of the image tiles is preserved to facilitate the subsequent stitching of the segmentation results back into a global vector map.

[0039] See attached document Figure 1 and attached Figure 2 In this embodiment, the data acquisition module and the image preprocessing module work together to build a high-standard, multi-dimensional remote sensing dataset, providing a data foundation for subsequent water body extraction and black and odor identification.

[0040] For the first high-resolution remote sensing image, this embodiment specifically selects Gaofen-2 satellite imagery as the data source. The GF-2 satellite carries two high-resolution cameras, capable of providing panchromatic imagery with a spatial resolution better than 1 meter and multispectral imagery with a spatial resolution of 3.2 meters. The multispectral bands specifically cover the blue, green, red, and near-infrared bands. The data acquisition module acquires GF-2L1A-level raw image data of the target monitoring area. This raw image data includes the satellite's attitude and orbital parameters during imaging.

[0041] The image preprocessing module performs a standardized preprocessing procedure on the acquired GF-2 images.

[0042] Radiometric calibration is performed by using absolute radiometric calibration coefficients measured before satellite launch and during in-orbit operation to convert the raw digital quantization values ​​of the image into entrance pupil radiance, thereby eliminating the sensor's own response error.

[0043] Atmospheric correction processing is performed using an atmospheric correction algorithm based on a radiative transfer model. Atmospheric scattering and absorption effects are eliminated based on atmospheric parameters at the time of imaging, and the true surface reflectance is retrieved.

[0044] Orthorectification was performed, incorporating a digital elevation model and the image's built-in rational polynomial coefficient model to perform geometrical fine correction, eliminating geometric distortions caused by terrain undulations and camera imaging angles, and ensuring the spatial accuracy of the image. Subsequently, image fusion processing was performed, employing the Gram-Schmidt spectral sharpening fusion algorithm to fuse the high spatial resolution panchromatic band with the high spectral resolution multispectral band, generating a standard orthorectified image with a spatial resolution of 0.8 meters and containing four spectral bands.

[0045] A sliding window technique was used to construct the dataset, with the window size set to 512×512 pixels and the step size to 256 pixels. The large-format image was cropped into image tiles and divided into training and validation sets.

[0046] For the second type of multispectral remote sensing image, this embodiment specifically selects Sentinel-2 satellite imagery as the data source. The multispectral imager on the Sentinel-2 satellite covers 13 bands, including visible light, near-infrared, and short-wave infrared, and has a wide swath width and high radiometric resolution.

[0047] This embodiment focuses on utilizing its blue band (Band 2), green band (Band 3), red-edge band (Band 8A), and shortwave infrared band (Band 11). The data acquisition module acquires Sentinel-2L2A level products, which have undergone atmospheric correction and represent the reflectivity of the lower atmosphere. To ensure temporal consistency of the spectral characteristics of ground features, the selected image imaging time is controlled to deviate from the imaging time of the GF-2 image within ±7 days.

[0048] The image preprocessing module performs resampling and registration on the Sentinel-2 imagery. Due to the spatial resolution differences between different bands of Sentinel-2, Band 8A and Band 11 need to be resampled to a 10-meter resolution using bilinear interpolation to unify their spatial scale with Band 2 and Band 3. Subsequently, using orthorectified GF-2 imagery as a reference, an automatic feature point matching algorithm is used to spatially register the Sentinel-2 imagery, ensuring strict geospatial alignment between the two data sources. The root mean square error (RMSE) after registration is controlled within 0.5 pixels.

[0049] For the third high-frequency remote sensing imagery, this embodiment specifically selects PlanetScope satellite constellation imagery as the data source. The PlanetScope satellite constellation has daily global revisit capability, providing high-frequency surface observation data. The selected PlanetImagery in this embodiment has a spatial resolution of 3 meters and includes four bands: blue, green, red, and near-infrared. The data acquisition module uses the time window of the suspected black and odorous water body candidate set selected in step S3 as a benchmark to acquire PlanetScope Level 3B orthophoto products with a deviation of no more than 30 days.

[0050] The image preprocessing module primarily performs mosaicking and cropping operations on Planet images. Since Planet images are typically distributed as narrow strips or tiles, multiple images must first be mosaicked into a complete image map covering the target monitoring area based on geographic coordinates. Subsequently, instead of processing the entire image, image sub-blocks with fixed pixel sizes are cropped from the Planet image using the geometric centers of each patch in the candidate set of suspected black and odorous water bodies output in step S3 as positioning references. These sub-blocks serve as input data for the subsequent deep learning classification model. This on-demand cropping strategy significantly reduces the amount of data processing, allowing focus on the detailed texture feature analysis of suspected targets.

[0051] See attached document Figure 1 and attached Figure 2 In this embodiment, the semantic segmentation module is responsible for performing the core task of step S1, which is to accurately extract water targets from the first high-resolution remote sensing image using deep learning technology. This process specifically includes three stages: model building, model training under loss function constraints, and vectorization processing after inference.

[0052] Network architecture of semantic segmentation model:

[0053] The semantic segmentation module employs a TransUNet hybrid network structure, which deeply integrates the Transformer mechanism with the U-Net architecture. This hybrid network structure is designed to combine the advantages of convolutional neural networks (CNNs) in extracting local texture features with the ability of Transformers to establish global long-range dependencies.

[0054] Specifically, the model's encoder consists of a convolutional neural network feature extraction unit and a Transformer sequence processing unit cascaded together. First, the preprocessed high-resolution remote sensing image is input into the convolutional neural network feature extraction unit, which generates a series of feature maps at different scales through successive convolutional layers and downsampling operations. These feature maps can characterize the shallow spatial structure and texture information of the image.

[0055] Subsequently, the deep feature maps output by the convolutional neural network are converted into serialized input vectors (Tokens) and fed into a Transformer sequence processing unit. This Transformer sequence processing unit contains multiple layers of self-attention mechanism modules. Under the action of the self-attention mechanism, the model calculates the correlation strength between each pixel in the input sequence and all other pixels, thereby capturing contextual information across the entire image. This global receptive field enables the model to understand the continuous topological structure of the river, effectively solving the problem of river channel discontinuity caused by bridge occlusion, tree cover, or shadow interference.

[0056] The model's decoding path employs a cascaded upsampler structure. To restore the fineness of the water body edges, the decoding path fuses features with the encoding path through a skip connection structure. Specifically:

[0057] The shallow, high-resolution feature map generated by the convolutional neural network in the encoding path is concatenated with the deep semantic feature map in the decoding path, which has been upsampled to restore its spatial dimensions, along the channel dimension. The fused features are further refined through convolution operations, ultimately restoring the local edges and texture details of the image, and outputting a global water probability map with the same size as the input image.

[0058] Construction of the hybrid loss function:

[0059] During the training of the semantic segmentation model, in order to address the severe imbalance in pixel count between water objects (foreground) and non-water backgrounds, the semantic segmentation module employs a hybrid loss function for parameter updates and constraints. This hybrid loss function is a weighted combination of a region consistency metric (Dice Loss) and a pixel classification probability metric (binary cross-entropy loss).

[0060] The region consistency metric employs the Dice loss function, which focuses on measuring the degree of spatial geometric overlap between the predicted mask and the ground truth label. It is insensitive to the size of the foreground region and effectively prevents small water bodies from being ignored by the model during training. (Dice loss function) The calculation formula is as follows:

[0061] ;

[0062] in, Represents pixels in an image; Indicates the model predicts pixels This represents the probability value of the water body's foreground. Represents pixels The actual label value (1 represents water, 0 represents background); To prevent numerical stability constants with a denominator of zero (e.g., 1e-5).

[0063] The pixel classification probability metric employs the binary cross-entropy loss function, which is based on the assumption of independent and identically distributed pixels. It focuses on constraining the probability distribution of a single pixel being correctly classified, thus helping to improve the classification accuracy of pixels at water body boundaries. Binary cross-entropy loss function The calculation formula is as follows:

[0064] ;

[0065] in, Represents the total number of pixels in an image block.

[0066] The total loss function is a weighted sum of the two above. The model parameters converge by minimizing the total loss function through the backpropagation algorithm.

[0067] Generation of global vector water mask:

[0068] The semantic segmentation module uses a trained TransUNet model to perform sliding inference on the first high-resolution remote sensing image of the target monitoring area, outputting a global water body probability map. Each pixel value in this global water body probability map represents its confidence level (range 0-1) in belonging to a water body.

[0069] Next, the global vector water mask generation step is performed. First, the global water probability map is binarized, a probability threshold is set, and pixels with a probability greater than the threshold are marked as water, while pixels with a probability less than the threshold are marked as background, resulting in a binarized image.

[0070] Next, morphological opening and closing operations are performed on the binarized image. Morphological opening operations are used to break up narrow, contiguous regions and smooth the boundaries, while morphological closing operations are used to fill in small cavities inside the water body. Simultaneously, the area of ​​connected components is calculated, and tiny patches with an area smaller than a preset threshold are removed to eliminate noise interference.

[0071] Finally, the processed binarized image is converted into vector polygon data using a raster-to-vector algorithm. This vector polygon data is then used as the global vector water mask, serving as the basis for rule-based filtering in subsequent step S2.

[0072] See attached document Figure 1 and attached Figure 2 In this embodiment, the rule filtering module is responsible for executing step S2, which involves using prior knowledge to clean and initially screen the global vector water mask generated in step S1. This process aims to utilize the relatively low computational cost of geometric and basic spectral features to pre-remove patches that clearly do not conform to the physical properties of black and odorous water bodies, thereby reducing the computational load of subsequent high-complexity spectral calculations and deep learning inference.

[0073] The rule-based filtering module first traverses each independent connected polygon (i.e., water patch) in the global vector water mask, calculates its geometric attribute data, and performs filtering according to the preset geometric feature filtering rules.

[0074] The first level of geometric filtering is based on area attributes. The rule-based filtering module calculates the projected area of ​​each water body patch. Since urban black and odorous water bodies are usually concentrated in small and medium-sized tributaries, ditches, or ponds, while large lakes, reservoirs, and wide main rivers have a lower probability of becoming black and odorous due to their strong self-purification capacity, an upper limit threshold for area is set. Patches exceeding this upper limit threshold are marked as non-monitoring objects and removed.

[0075] The second-level geometric filtering is based on shape compactness properties. The regularity filtering module constructs a compactness index to characterize the shape complexity and elongation of water body patches. This compactness index is used to distinguish between naturally occurring irregular water bodies, artificially planned regular water bodies, and linear main channels. For each water body patch, its compactness index is calculated. .

[0076] Density index The calculation formula is as follows:

[0077] ;

[0078] in, Indicates the area of ​​a water patch; Indicates the perimeter of a water patch; Pi is a constant.

[0079] According to the formula definition, the density index of a circular shape is close to 1, while the density index of a thin, elongated strip or extremely complex water body is close to 0. The rule-based filtering module sets the shape threshold range based on the water system distribution characteristics of the target monitoring area.

[0080] For example, for patches with extremely regular shapes and large areas (such as artificial landscape lakes), or for main navigable waterways that are too long and narrow and have a width exceeding a certain threshold, patches that are more in line with the characteristics of urban capillary tributaries and culverts can be identified and removed by setting corresponding density index ranges.

[0081] Building upon geometric filtering, the rule-based filtering module further incorporates spectral features for a third level of filtering. The data source used at this stage is the first high-resolution remote sensing image (GF-2). Although the GF-2 image has fewer bands, its near-infrared bands are indicative of water body identification.

[0082] The rule-based filtering module calculates the spectral mean of all pixels within each water body patch to be identified on the first high-resolution remote sensing image and calculates the mean of the Normalized Difference Water Index (NDWI).

[0083] Normalized Difference Water Index The calculation formula is as follows:

[0084] ;

[0085] in, This represents the reflectance value of the green band in the first high-resolution remote sensing image; This represents the reflectance value in the near-infrared band of the first high-resolution remote sensing image.

[0086] set up Validity threshold (e.g., 0.1), for Patches with a mean value below the validity threshold are identified as false positives in semantic segmentation of non-water bodies (such as building shadows, asphalt pavements, etc.) and removed from the set.

[0087] After the cascaded filtering based on area, shape density, and basic spectral indices, the rule filtering module outputs a set of water bodies to be identified. This set of water bodies retains only patches with moderate spatial scale, morphology consistent with tributary characteristics, and confirmed as water body attributes, which serve as the input for multispectral inversion in step S3.

[0088] See attached document Figure 1 and attached Figure 2In this embodiment, the spectral analysis module is responsible for executing step S3, which utilizes the rich spectral band information of the second multispectral remote sensing image to perform a qualitative preliminary screening of the water quality attributes of the set of water bodies to be identified retained after geometric filtering in step S2. The core of this process lies in constructing a spectral index that has sensitive response characteristics to black and odorous water bodies, and using statistical methods to adaptively determine the judgment threshold, thereby screening out a candidate set of suspected black and odorous water bodies.

[0089] The spectral analysis module first calls the second multispectral remote sensing image processed by the image preprocessing module. Based on the spatial range of the water bodies to be identified, the multispectral pixel values ​​of the corresponding locations are extracted. To quantify the degree of blackness and odor of the water bodies, this embodiment constructs a specific Black and Odorous Water Index (BOI). This Black and Odorous Water Index utilizes the specific reflectance characteristics of black and odorous water bodies in the visible green band, red-edge band, blue band, and shortwave infrared band: black and odorous water bodies usually contain a large amount of suspended particulate matter and organic pollutants, exhibiting abnormal reflectance in the red-edge band, and the differences in the shortwave infrared and blue bands are also different from those of general clean water bodies.

[0090] The spectral analysis module calculates the black and odorous water body index based on the principle of normalized difference calculation. Specifically, it performs normalized difference processing on the green band and red edge band, and the shortwave infrared band and blue band, respectively, and multiplies the calculation results of the two to amplify the spectral contrast between the black and odorous water body and the background water body.

[0091] Black and odorous water index The calculation formula is as follows:

[0092] ;

[0093] in, This represents the pixel reflectance value of the green band in the second multispectral remote sensing image; This represents the pixel reflectance value of the red-edge band in the second multispectral remote sensing image; This represents the pixel reflectance value in the shortwave infrared band of the second multispectral remote sensing image; This represents the pixel reflectance value of the blue band in the second multispectral remote sensing image.

[0094] After calculating the BOI index for the entire region or a specified area, the spectral analysis module generates a grayscale index map. To achieve automated target extraction and avoid the problem of poor spatiotemporal adaptability caused by manually setting fixed thresholds, this embodiment uses an adaptive threshold segmentation method to determine the optimal segmentation point for distinguishing between black and odorous areas and non-black and odorous areas.

[0095] Specifically, a statistical threshold selection strategy based on maximum inter-class variance is adopted. This strategy divides the pixels in the image into two groups according to gray level: background (non-black and non-saturated) and foreground (black and non-saturated). It iterates through all possible gray level thresholds and calculates the variance between the two classes of pixels. When the inter-class variance reaches its maximum, it means that the two classes of pixels have the highest distinguishability, and the corresponding gray level is the optimal segmentation threshold.

[0096] Between-class variance The calculation formula is as follows:

[0097] ;

[0098] in, Indicates less than the threshold The percentage of background pixels in the total pixels; Indicates greater than the threshold The proportion of foreground pixels in the total pixels; This represents the average grayscale value of the background pixels; This represents the average grayscale value of the foreground pixels.

[0099] After the spectral analysis module calculates the optimal segmentation threshold, it performs the final filtering operation. The filtering criteria are set as a logical AND relationship:

[0100] First, the mean BOI index within the water body patch to be identified must be greater than the optimal segmentation threshold.

[0101] To prevent misjudgment by shoreline vegetation or highly reflective artificial structures, the normalized difference water index of the patch must simultaneously meet the preset water body characteristic conditions.

[0102] Patches that meet the above dual conditions are marked as highly suspicious targets, and the spatial location and vector boundary of these patches are output to form a candidate set of suspected black and odorous water bodies, which serves as the input object for the deep learning fine classification in step S4.

[0103] See attached document Figure 1 and attached Figure 2 In this embodiment, the deep classification module is responsible for executing step S4. Leveraging the high revisit rate of the third high-frequency remote sensing imagery and the high-dimensional feature extraction capability of the deep learning model, it performs the final verification of the candidate set of suspected black and odorous water bodies selected in step S3. This step aims to address the problem that relying solely on spectral indices can easily misclassify features with similar spectral characteristics, such as cloud shadows, building shadows, and dark road surfaces, as black and odorous water bodies.

[0104] The deep classification module first performs precise cropping of image patches. To focus on the texture details of the target and reduce background noise interference, the deep classification module does not directly input the entire image into the network, but instead adopts a target-centered cropping strategy. Specifically, the data acquisition module acquires third-frequency remote sensing images that match the time window of the candidate set of suspected black and odorous water bodies.

[0105] Subsequently, the deep classification module traverses each vector patch in the suspected candidate set, calculating the geometric center coordinates of each patch. Using these geometric center coordinates as the localization reference, a sub-block with a fixed pixel size (e.g., 64×64 pixels or 128×128 pixels, depending on the dimension setting of the network input layer) is extracted from the third high-frequency remote sensing image. This sub-block not only contains the target water body but also retains some contextual information about the target's surrounding environment, helping the model determine the spatial adjacency relationship between the water body and its surrounding environment.

[0106] The deep classification module constructs and loads a pre-trained deep learning classification model to perform inference on the cropped image sub-blocks. In this embodiment, the deep learning classification model adopts a convolutional neural network architecture based on a compound scaling strategy.

[0107] Unlike traditional convolutional networks that scale only in a single dimension (such as depth, width, or resolution), this architecture employs a unified composite scaling factor to jointly adjust and optimize the network's depth (number of layers), width (number of channels), and the resolution of the input image. Through this composite scaling strategy, the model can maximize parameter utilization and feature extraction efficiency within limited computational resource constraints, adapting to the large scale variations of black and odorous water bodies in remote sensing imagery.

[0108] The fundamental building blocks of the convolutional neural network architecture integrate a moving inverted bottleneck convolutional module with compression and activation attention mechanisms. In the moving inverted bottleneck convolutional module, the network first uses a 1×1 convolutional layer to upscale the input feature map, mapping the features to a higher-dimensional space to enrich the feature representation. Then, 3×3 or 5×5 depthwise separable convolutions are performed in the depth direction to extract spatial features. Finally, a 1×1 convolutional layer is used for dimensionality reduction projection to output the feature map. This inverted bottleneck structure, small at both ends and large in the middle, effectively reduces computational cost while preserving rich information.

[0109] Meanwhile, a compression and activation attention mechanism is embedded in each moving inverted bottleneck convolutional module. This mechanism first compresses the feature map spatially using global average pooling to obtain channel-level global feature descriptors. Then, it learns the non-linear interdependencies between channels through two fully connected layers, generating channel weight coefficients. Finally, it normalizes the weights to between 0 and 1 using the sigmoid activation function and reconstructs the weights for each channel of the original feature map. By introducing this attention mechanism, the model can adaptively suppress background channel responses that are useless for distinguishing between black and odorous conditions, enhancing its sensitivity to key features such as water color, turbidity, and texture.

[0110] The deep classification module inputs the cropped image sub-blocks into the aforementioned network model. After multi-layer feature extraction and nonlinear transformation, the fully connected layer outputs a value between 0 and 1, representing the confidence score for black and odorous water bodies. This confidence score characterizes the confidence that the input patch belongs to the black and odorous water body category. If the confidence score is greater than a pre-set confirmation threshold (e.g., 0.8), the input patch is determined to be a confirmed black and odorous water body; otherwise, it is determined to be a false alarm target and removed. The final confirmed result will proceed to step S5 for spatial overlay and mapping output.

[0111] See attached document Figure 1 and attached Figure 2 In this embodiment, the result fusion module is responsible for executing step S5, which involves performing a final interpretation of the confidence level of the black and odorous water body obtained through inference calculation by the deep classification module, and fusing the interpretation result with the spatial vector data to generate result data that conforms to the geographic information system standard. This step marks the end of the multi-source remote sensing data collaborative processing flow, realizing the conversion from raw imagery to thematic vector patches.

[0112] The results fusion module first receives the confidence score for each suspected black and odorous water patch from the deep classification module. A binary classification threshold is set (e.g., 0.5 or 0.8, adjusted according to the application's preference for recall or precision). The results fusion module iterates through all candidate patches, marking patches with confidence scores greater than the set threshold as confirmed black and odorous, and patches with confidence scores less than or equal to the set threshold as non-black and odorous.

[0113] Subsequently, the results fusion module performs a spatial attribute association operation. This operation uses the vector polygon generated in step S2 and retained after filtering in steps S3 and S4 as the spatial reference. Since this vector polygon is generated based on the first high-resolution remote sensing image (GF-2), it retains sub-meter level geometric boundary accuracy. The results fusion module writes the classification labels confirmed as black and smelly, the confidence level of the deep learning model prediction, and the identification timestamp, among other attribute information, into the attribute table of the corresponding vector polygon. For patches marked as non-black and smelly, they can be removed from the results set or retained and marked as negative samples for subsequent model iteration training.

[0114] The results fusion module encapsulates vector polygon data containing attribute information into a common geospatial data format. This output file contains complete geographic coordinate system information, ensuring accurate overlay of the identification results onto the underlying geographic map. The output spatial distribution map visually displays the specific location, morphological distribution, and pollution level of black and odorous water bodies within the monitoring area.

[0115] To verify the effectiveness of the intelligent identification method for black and odorous water bodies provided in this embodiment, a specific experimental verification scenario was conducted, in which a built-up area of ​​a city was selected as the test area, with an area of ​​approximately 100 square kilometers.

[0116] The data source specifications used in the experiment are as follows:

[0117] The first high-resolution remote sensing image: GF-2 panchromatic / multispectral fusion image, spatial resolution 0.8 meters, imaging time T days. The second multispectral remote sensing image: Sentinel-2L2A image, resampling resolution 10 meters, imaging time T+2 days. The third high-frequency remote sensing image: Planet Scope orthophoto image, resolution 3 meters, imaging time T days.

[0118] The experimental verification process and results are as follows:

[0119] The semantic segmentation module extracts a global water body mask based on GF-2 imagery. Through inference using the TransUNet model, a total of 150 water body patches were extracted, including main rivers, tributaries, landscape lakes, and several small ditches.

[0120] The rule-based filtering module uses geometric features to remove large lakes with an area greater than 50,000 square meters and main river channels with a density index of less than 0.05, leaving 120 water patches to be identified.

[0121] The spectral analysis module calculated the BOI index based on Sentinel-2 imagery and combined it with Otsu threshold segmentation. Forty-five suspected black and odorous water patches with abnormal BOI indices were identified.

[0122] The deep classification module confirmed the texture features of the 45 suspected patches based on Planet imagery. The EfficientNet model determined that 38 of the patches were black and odorous water bodies (confidence > 0.8), and excluded 7 false alarm targets caused by building shadows or dark sediment.

[0123] The system output data for 38 black and odorous water patches was compared with ground-based measured data to calculate evaluation metrics. These metrics included precision, recall, and F1 score.

[0124] Accuracy The calculation formula is as follows:

[0125] ;

[0126] in, This indicates the number of patches that the system correctly identified as black and odorous water bodies; This indicates the number of patches that the system incorrectly identified as black and odorous water bodies.

[0127] Recall rate The calculation formula is as follows:

[0128] ;

[0129] in, This indicates the number of actual black and odorous water patches that the system missed detecting.

[0130] The formula for calculating fractions is as follows:

[0131] ;

[0132] in, Indicates accuracy; This indicates the recall rate.

[0133] In this experimental verification, statistics showed that... =36, =2, =4. The calculated precision is 94.7%, and the recall rate is... =90.0%, =92.3%. Experimental results show that the intelligent identification method for black and odorous water bodies based on multi-source remote sensing images provided by this invention, while ensuring a high recall rate, effectively controls the false alarm rate through multi-level screening and multi-source data verification, and can meet the technical requirements of urban black and odorous water body supervision for spatial positioning accuracy and identification accuracy.

Claims

1. A method for intelligent identification of black and odorous water bodies based on multi-source remote sensing imagery, characterized in that, Includes the following steps: Step S1: After acquiring the first high-resolution remote sensing image of the target monitoring area and performing preprocessing, input it into the pre-trained semantic segmentation model for inference, obtain the global water body probability map, and generate a global vector water body mask based on the threshold; wherein, the spatial resolution of the first high-resolution remote sensing image is better than 1 meter, and the band composition of the first high-resolution remote sensing image includes at least the blue band, green band, red band and near-infrared band. Step S2: Based on the geometric and spectral features of the global vector water body mask, construct filtering rules to remove large lakes and reservoirs and default non-black and odorous water body patches, and obtain the set of water bodies to be identified; Step S3: Obtain a second multispectral remote sensing image that matches the imaging time window of the first high-resolution remote sensing image; resample and register the spatial resolution of the second multispectral remote sensing image to the spatial range of the set of water bodies to be identified; calculate the black and odorous water body index based on the second multispectral remote sensing image, and use an adaptive threshold segmentation method to screen out a candidate set of suspected black and odorous water bodies; wherein, the second multispectral remote sensing image includes visible light band, red edge band, and shortwave infrared band for calculating the spectral index; the black and odorous water body index is constructed based on normalized difference operation, the normalized difference operation specifically is: using the green band and red edge band of the second multispectral remote sensing image to perform a first normalized difference operation, using the shortwave infrared band and blue band to perform a second normalized difference operation, and using the product of the result of the first normalized difference operation and the result of the second normalized difference operation as the value of the black and odorous water body index; Step S4: Obtain a third high-frequency remote sensing image that matches the time window of the candidate set of suspected black and odorous water bodies, select the patches in the candidate set of suspected black and odorous water bodies as the center cropped image blocks, input them into a pre-trained deep learning classification model for inference, and output the confidence score of black and odorous water bodies; wherein, the spatial resolution of the third high-frequency remote sensing image is better than or equal to the spatial resolution of the first high-resolution remote sensing image, and includes blue band, green band, red band and near-infrared band; Step S5: When the confidence level of the black and odorous water body is greater than the set threshold, the corresponding patch is confirmed as a black and odorous water body, and the identification results are spatially superimposed and output.

2. The intelligent identification method for black and odorous water bodies based on multi-source remote sensing imagery according to claim 1, characterized in that, In step S1, the semantic segmentation model is a deeply integrated hybrid network structure, Trans U-Net, which combines the Transformer mechanism with the U-Net architecture. The hybrid network structure is used to introduce a self-attention mechanism module in the encoding path, while retaining a skip connection structure in the decoding path; The self-attention mechanism module is used to capture the global context information and long-distance dependencies of the input image, and the skip connection structure is used to fuse the shallow feature map of the encoding path with the deep feature map of the decoding path and restore the local edges and texture details of the image.

3. The intelligent identification method for black and odorous water bodies based on multi-source remote sensing imagery according to claim 2, characterized in that, During the training process, the semantic segmentation model is constrained by a hybrid loss function that includes a region consistency metric and a pixel classification probability metric. The region consistency metric is used to measure the similarity between the predicted mask and the real label in spatial overlap, and is insensitive to the size of the foreground region. The pixel classification probability metric is based on the assumption of independent and identically distributed pixels, and is used to constrain the probability distribution of a single pixel being correctly classified.

4. The intelligent identification method for black and odorous water bodies based on multi-source remote sensing imagery according to claim 1, characterized in that, In step S2, the filtering rules include geometric feature filtering rules, which are specifically as follows: Large water bodies exceeding a preset size are removed based on the area attribute of water patches; A density index based on the ratio between area and the square of perimeter is constructed. The density index is used to characterize the elongated shape of water patches, thereby identifying and eliminating main river channels.

5. The intelligent identification method for black and odorous water bodies based on multi-source remote sensing imagery according to claim 1, characterized in that, In step S3, the adaptive threshold segmentation method specifically includes: A statistical threshold selection strategy based on the maximum inter-class variance is adopted; Iterate through the gray levels to calculate the variance between the background class and the foreground class, and select the gray level that maximizes the variance as the optimal segmentation threshold. The screening criteria are set as follows: the black and odorous water index of the patch is greater than the optimal segmentation threshold, and the normalized difference water index of the patch meets the preset water characteristic conditions.

6. The intelligent identification method for black and odorous water bodies based on multi-source remote sensing imagery according to claim 1, characterized in that, In step S4, the deep learning classification model adopts a convolutional neural network architecture based on a composite scaling strategy. The convolutional neural network architecture adjusts the depth, width and resolution of the network through a unified scaling factor, and integrates a moving inverted bottleneck convolution module and a compression and attention activation mechanism to extract the deep texture and spectral features of the water target.

7. The intelligent identification method for black and odorous water bodies based on multi-source remote sensing imagery according to claim 1, characterized in that, In step S4, the operation of cropping the image block is as follows: the geometric center of each patch in the suspected black and odorous water body candidate set is selected as the positioning reference, and image sub-blocks with fixed pixel size are cropped from the third high-frequency remote sensing image.

8. The intelligent identification method for black and odorous water bodies based on multi-source remote sensing imagery according to claim 1, characterized in that, In step S1, the step of generating the global vector water mask includes: The global water probability map is binarized, and morphological opening and closing operations are performed to remove tiny noise patches with an area smaller than a preset threshold. The processed binary image is converted into vector polygon data, and the vector polygon data is determined as the global vector water body mask.

Citation Information

Patent Citations

  • Black and odorous water body identification method

    CN114913437A

  • Rural black and odorous water body remote sensing recognition method

    CN118135412A