Drainage basin runoff pollution characteristic dynamic integration clustering zoning method fusing multi-source spatio-temporal data and landscape ecological characteristics

By integrating multi-source spatiotemporal data and landscape ecological characteristics into a clustering and zoning method, and combining K-Means, hierarchical clustering and Gaussian mixture model, the problem of spatial scale definition and pollution feature zoning of watershed rainfall-runoff pollution was solved, achieving more accurate pollution feature zoning.

CN121256404APending Publication Date: 2026-01-02CHINA THREE GORGES CORPORATION +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511246757.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Traditional methods are insufficient to effectively define the spatial scale of rainfall-runoff pollution within a watershed and to identify pollution patterns under different spatiotemporal characteristics, resulting in inaccurate pollution feature zoning.

Method used

A dynamic integrated clustering and zoning method for watershed runoff pollution characteristics is adopted, which integrates multi-source spatiotemporal data and landscape ecological features. The method integrates K-Means algorithm, hierarchical clustering algorithm and Gaussian mixture model, and combines self-attention mechanism and density clustering algorithm for data preprocessing. The co-coherence matrix is ​​used to evaluate the clustering results and perform stability analysis.

Benefits of technology

It enables accurate zoning of rainfall-runoff pollution within the watershed, solves the problem of spatial scale definition, and improves the accuracy and stability of pollution characteristic zoning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121256404A_ABST
    Figure CN121256404A_ABST
Patent Text Reader

Abstract

The invention relates to a watershed runoff pollution characteristic dynamic integration clustering zoning method fusing multi-source spatio-temporal data and landscape ecological characteristics, and the method comprises the steps: carrying out the data collection and preprocessing of related impact factors, such as urban rainfall runoff pollution data, meteorological element data, urban humanity attribute data, land utilization landscape pattern distribution data, and the like; establishing a multi-source heterogeneous data set; constructing a runoff pollution characteristic clustering system of different spatial scales of the watershed; and constructing a rainfall runoff pollution zoning system in combination with a clustering result. According to the method, runoff pollution data, meteorological element data, urban humanity attribute data and landscape pattern distribution data are combined, the influence of climate, urban humanity and land utilization on space-time distribution of surface runoff pollution is comprehensively considered, and runoff pollution characteristics of different regions are evaluated; and further support is provided for urban runoff pollution treatment and sponge city construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a dynamic integrated clustering and zoning method for watershed runoff pollution characteristics that integrates multi-source spatiotemporal data and landscape ecological features, belonging to the technical fields of artificial intelligence and landscape ecology. Background Technology

[0002] In recent years, rapid urban development has led to a year-on-year increase in the proportion of hardened surfaces. Hardened surfaces accelerate the flow of rainwater runoff pollutants on the surface, making it difficult for pollutants in water bodies to be effectively removed through sedimentation or soil absorption, thus significantly impacting the water quality of receiving water bodies. Due to the complexity of urban hydrology, pollutant formation processes, and precipitation patterns, precipitation water quality exhibits high spatiotemporal variability. The biogeochemistry of precipitation is influenced by clear-day sediments on urban surfaces, human activities and drainage systems, as well as local temperature, humidity, and photodegradation conditions. The quality of pollutants accumulating on urban surfaces depends on the deposition rate, the number of dry days prior to precipitation, resuspension, aggregation, and re-deposition processes. These processes are influenced by watershed characteristics such as surface impermeability, land use type, and land cover, which have effects on patch size and landscape distribution structure within the watershed. Climate, weather, land use, and impermeability characteristics within a watershed determine the complex process of runoff quality generated by rainfall; however, studies on the characteristics of runoff quality at shorter time periods or smaller spatial scales are less comprehensive and struggle to identify pollution patterns under different spatiotemporal characteristics.

[0003] Unsupervised learning in machine learning, especially clustering algorithms, can intuitively analyze the structure and distribution of datasets, understand the characteristics of the data, reveal potential patterns and trends, and analyze the similarity between different data samples. K-Means is a distance-based hard clustering method that is computationally efficient and its results are easy to interpret, but it is sensitive to initial centers and prone to getting trapped in local optima. Hierarchical clustering algorithms merge or split clusters by constructing tree-like structures; they do not require pre-specifying the number of clusters and can discover clustering structures at different levels, but this method is sensitive to noise and outliers. Gaussian mixture models are soft clustering methods based on probability distribution models, capable of capturing complex data distributions, but with high computational complexity.

[0004] By integrating the K-Means algorithm, hierarchical clustering algorithm, and Gaussian mixture model, this study comprehensively addresses the challenge of parameter determination and optimization in the use of a single algorithm due to the heterogeneity of multi-source data in rainfall-runoff pollution research. Through efficient joint optimization of the three algorithms, it effectively solves the problems of traditional analysis methods in defining the spatial scale of pollution zoning within a watershed, especially in the field of rainfall-runoff pollution, and the limitation of pollution characteristics to a single rainfall event. This makes it impossible to effectively zon the rainfall-runoff pollution characteristics at different spatial scales within a watershed. Summary of the Invention

[0005] To address the aforementioned problems, this invention discloses a dynamic integrated clustering and zoning method for watershed runoff pollution characteristics that fuses multi-source spatiotemporal data and landscape ecological features. The specific technical solution is as follows:

[0006] A dynamic integrated clustering and zoning method for watershed runoff pollution characteristics that integrates multi-source spatiotemporal data and landscape ecological features includes the following steps:

[0007] Step 1: Determine the spatial scale of rainfall runoff pollution characteristics;

[0008] Step 2: Collect basic data to form the basic dataset M;

[0009] Step 3: Preprocess the base dataset M by imputing missing values, handling outliers, and normalizing it to obtain dataset M*;

[0010] Step 4: Calculate the required total landscape area, the proportion of landscape area occupied by patches, and the landscape shape index. Add the total landscape area, the proportion of landscape area occupied by patches, the landscape shape index, and the geographic coordinate data of the study area to the dataset M* to obtain the dataset M^.

[0011] Step 5: Compare and select different clustering methods. At the underlying surface scale, use dataset M* as the input matrix for the ensemble clustering algorithm; at the functional zone scale and the city scale, use dataset M^ as the input matrix for the ensemble clustering algorithm. Entire clustering is used to calculate the characteristic zoning of rainfall runoff pollution, yielding the clustering results.

[0012] Step 6: Evaluation of Clustering Results: For the clustering results... The number of clusters X is used to measure the compactness within clusters and the separation between clusters using the silhouette coefficient and the sum of squared errors. A closer silhouette coefficient to 1 indicates a more reasonable number of clusters. A graph of the sum of squared errors versus the number of clusters is plotted; when the rate of decrease in the sum of squared errors for X as the number of clusters decreases begins to slow, X represents the optimal number of clusters. This represents the optimal clustering result.

[0013] Step 7: After the basic data of the new region is added to the basic dataset M, a new dataset N is formed. Steps 2-6 are repeated at the same spatial scale to obtain a new round of clustering results. Comparing clustering results and To determine whether the output results are stable, if the clustering results are stable under sufficient data conditions, then the new region is considered to have a pollution characteristic standard partition at that spatial scale.

[0014] Furthermore, step 1 specifically involves dividing the study area into urban scale, functional zone scale, and underlying surface scale within the watershed based on the target range of pollution characteristic zoning.

[0015] Furthermore, the basic data in step 2 includes a subset A of rainfall-runoff pollution, a subset B of meteorological elements, a subset C of urban human attributes, a subset D of land use, and a subset E of underlying surface characteristics of the study area. The basic dataset M is formed by the subsets A, B, C, D, and E.

[0016] The rainfall-runoff pollution subset A includes the average concentrations of COD, TN, TP, and SS from each rainfall event.

[0017] The meteorological element subset B includes rainfall amount, rainfall duration, pre-rain dry period, and average rainfall intensity for each rainfall event, as well as the annual average rainfall and annual average temperature of the city in the study area.

[0018] The urban humanistic attribute subset C includes urban per capita GDP, urban GDP output, the proportion of primary industry, secondary industry, and tertiary industry, as well as the green coverage rate of urban built-up areas;

[0019] Land use subset D consists of high-resolution remote sensing image data of the study area;

[0020] The underlying surface feature subset E includes material and type data of the underlying surface within the study area.

[0021] Furthermore, the data preprocessing in step 3 is as follows: missing data is imputed using the self-attention time-series interpolation model SAITS; outliers are eliminated using density-space clustering; and the data is mapped to the range of 0-1 using max-min normalization. The specific steps are as follows:

[0022] S3.1 For handling missing values, the SAITS time series imputation model based on the self-attention mechanism is used for imputation. In this model, 10% of the true values ​​are masked by an artificial mask, and the diagonal masking self-attention mechanism combined with the weighted combination module is used to effectively imput the missing spatiotemporal series data.

[0023] S3.2 For outlier processing, the density-based clustering algorithm DBSCAN is used to identify outliers;

[0024] S3.3 uses the max-min normalization method to normalize the data after imputation of missing values, mapping the data to the range of 0-1. This eliminates the influence of different units and dimensions between different feature indicators on model training, while accelerating model convergence and reducing training time. The specific formula is as follows:

[0025]

[0026] In the above formula, x * The data is normalized; x is a data point in the original dataset that needs to be normalized. min The minimum value in the dataset; x max This represents the maximum value in the dataset.

[0027] Furthermore, the specific steps for data processing of the land use subset D, which includes the city-scale and functional zone-scale data, in step 4 are as follows:

[0028] Step 4.1 Use Aowei Map to download high-resolution remote sensing data images of the study area at 30m resolution. Import the downloaded vector map into ArcGIS software, select the range corresponding to the study area, divide the land use type in the study area into three types: road surface, roof, and green space, and cut and divide them according to their shape to form a vector grid dataset D* after preliminary processing.

[0029] Step 4.2 Import the pre-processed vector grid dataset D* into Fragstats software. Use Fragstats software to calculate the total landscape area, the proportion of landscape area occupied by patches, and the landscape shape index at the landscape scale. These correspond to the total area of ​​the study area, the proportion of the three underlying surfaces of the study area (road surface, roof surface, and green space), and the degree of irregularity of the shape of road surface, roof surface, and green space in the functional area, respectively.

[0030] Furthermore, in step 5, the clustering method selection method integrates three methods: K-Means clustering, hierarchical clustering, and Gaussian mixture model based on distribution clustering algorithm, and uses co-matrix for re-clustering.

[0031] Furthermore, step 5, ensemble clustering, specifically includes the following steps:

[0032] Step 5.1 For the K-Means algorithm and hierarchical clustering, calculate the silhouette coefficient and sum of squared errors for different numbers of clusters, and plot the silhouette coefficient and sum of squared errors as a function of the number of clusters; for the Gaussian mixture model, use the Bayesian Information Criterion (BIC) to plot the distribution of different numbers of groups and BIC indices.

[0033] Step 5.2 Run the K-Means algorithm, hierarchical clustering algorithm, and Gaussian mixture model once each, and compare the running results of the three methods with different numbers of clusters, as well as the relationship between the silhouette coefficient and BIC index in the corresponding algorithms and the number of clusters. When the silhouette coefficient corresponding to the number of clusters X is closer to 1 and the rate of decrease of the sum of squared errors begins to decrease, X is determined to be the optimal number of clusters, and the corresponding clustering result shows the optimal state.

[0034] Step 5.3 Select H times to run the K-Means algorithm, hierarchical clustering algorithm, and Gaussian mixture model respectively, and compare whether the results of multiple runs are stable. If the results still change after H runs with the optimal number of cluster groups, increase the number of runs and compare the clustering results until the clustering results of the K-Means algorithm, hierarchical clustering algorithm, and Gaussian mixture model are all stable. This solves the randomness problem caused by different initial states and different datasets.

[0035] Step 5.4 uses the optimal number of clusters from Step 5.2 to perform clustering using the K-Means algorithm, hierarchical clustering algorithm, and Gaussian mixture model, generating cluster labels for each algorithm. A co-coordinate matrix is ​​constructed to record the frequency with which samples are assigned to the same cluster in multiple clustering operations, and to record the similarity between samples. Based on the advantage of hierarchical clustering in discovering hierarchical relationships between classes, hierarchical clustering is used to finally partition the co-coordinate matrix. The usage conditions of different distance metrics such as Euclidean distance, Manhattan distance, Canberra distance, and Chebyshev distance in hierarchical clustering algorithms are compared, and a suitable distance metric is selected based on the data characteristics of the input matrix.

[0036] Furthermore, in step 7, as rainfall runoff pollution data from the new area are continuously added to the rainfall runoff pollution subset A in step 2, meteorological element data, urban human attribute data, and land use data corresponding to the new area are collected and added to the corresponding subsets to form a new dataset N.

[0037] Furthermore, based on clustering results and combined with different pollution characteristics and geographical zoning at different spatial scales, the distribution characteristics of different rainfall-runoff pollution within the watershed are determined; the runoff pollution characteristics in each cluster are analyzed, and the key characteristics and causes of runoff pollution differences in different regions are given, thus realizing the zoning of pollution characteristics under different spatial conditions within the watershed.

[0038] The beneficial effects of this invention are:

[0039] This invention utilizes data on influencing factors and runoff pollution results in urban rainfall-runoff pollution processes, combined with landscape ecology techniques, to define the study area as a landscape. Different underlying surface distributions represent different landscape patterns. Through high-resolution remote sensing images, land use types in different areas are extracted. By combining machine learning and landscape ecology techniques, this invention effectively solves the problems of traditional analysis methods, such as the difficulty in defining the spatial scale of pollution zoning within a watershed, especially in the field of rainfall-runoff pollution, and the limitation of pollution characteristics to a single rainfall event. Attached Figure Description

[0040] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0041] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.

[0042] Combined with appendix Figure 1 As can be seen, this dynamic integrated clustering and zoning method for watershed runoff pollution characteristics, which combines multi-source spatiotemporal data and landscape ecological features, includes the following steps:

[0043] Step 1: Spatial scale division of rainfall-runoff pollution characteristics: Within the watershed, the study area is divided into urban scale, functional zone scale, and underlying surface scale according to the target range of pollution characteristic zoning.

[0044] Step 2: Basic Data Collection: The basic data includes subsets A (rainfall-runoff pollution), B (meteorological elements), C (urban human attributes), D (land use), and E (underlying surface characteristics) at different spatial scales of the study area. These subsets A, B, C, D, and E constitute the basic dataset M. Within the spatial scales defined in Step 1, the basic datasets required for subsequent studies at the urban and functional zone scales consist of the aforementioned subsets A, B, C, and D; the basic datasets required for subsequent studies at the underlying surface scale consist of subsets A, B, C, and E.

[0045] Step 3: Data preprocessing: Preprocessing includes imputing missing values, handling outliers, and normalizing the data in the subsets A, B, C, and E of the base dataset M to obtain dataset M*;

[0046] Step 4: Data Processing: For the land use subsets D contained at the city and functional zone scales, ArcGIS software is used to extract the underlying surface distribution data corresponding to the vector maps of different study areas. The study areas include roads, roofs, and green spaces, resulting in a pre-processed vector grid dataset D*. The pre-processed vector grid datasets D* from different study areas are imported into Fragstats software to calculate the required total landscape area (CA), the proportion of landscape area occupied by patches (PLAND), and the landscape shape index (LSI). The above three landscape pattern indices and the geographic coordinate data of the study areas are added to dataset M* to obtain dataset M^.

[0047] Step 5: A comprehensive comparison of different clustering methods is conducted. K-Means clustering, hierarchical clustering, and the distribution-based Gaussian Mixture Model (GMMM) clustering algorithm are integrated. Using the dataset M^ as the input matrix for the ensemble clustering algorithm at the same spatial scale, the characteristic partitions of rainfall-runoff pollution are explored through ensemble clustering, yielding the clustering results.

[0048] Step 6: Clustering result evaluation: Use the silhouette coefficient to measure the compactness within clusters and the separation between clusters, and verify whether the selected number of clusters conforms to the elbow rule;

[0049] Step 7: As rainfall runoff pollution data from the new region is continuously added to the rainfall runoff pollution subset A from Step 2, meteorological element data, urban human attribute data, and land use data corresponding to the new region are collected and added to their respective subsets to form a new dataset N. Steps 2-6 are repeated at the same spatial scale to obtain a new round of clustering results. Comparing clustering results and To determine whether the output results are stable, if the clustering results are stable under sufficient data conditions, then it is considered that a pollution characteristic standard zoning at this spatial scale has appeared in the field of surface rainfall runoff pollution in the watershed.

[0050] Step 8: Based on the clustering results, combined with different pollution characteristics and geographical zoning, determine the zoning of different rainfall-runoff pollution distribution characteristics within the watershed; analyze the runoff pollution characteristics in each cluster, give the key characteristics and reasons for the differences in runoff pollution in different regions, and realize the zoning of pollution characteristics under different spatial conditions within the watershed.

[0051] Furthermore, the basic dataset in step 2 includes the following indicators: Rainfall runoff pollution subset A includes the average concentrations of COD, TN, TP, and SS for each rainfall event; meteorological element subset B includes rainfall amount, rainfall duration, average rainfall intensity, and pre-rain dry period for each rainfall event, as well as the annual average rainfall and annual average temperature of the city where the study area is located; urban human attributes subset C includes the average slope of the city where the study area is located, the city's per capita GDP, and the added value of the city's primary, secondary, and tertiary industries; land use subset D consists of high-resolution remote sensing image data of the study area; landscape pattern indices include total landscape area (CA), the proportion of landscape area occupied by patches (PLAND), and the geographic coordinates of the study area (LSI); and underlying surface feature subset E includes underlying surface features and texture. Of the above data, the rainfall runoff pollution data and meteorological element data are from existing research literature; the urban human attributes data are from the official statistical websites of the cities where the study area is located, provided by statistical yearbooks; and the land use landscape pattern data are from high-resolution remote sensing image data, extracting the distribution types of the underlying surface and recording the corresponding geographic coordinates.

[0052] The data preprocessing in step 3 includes the following steps:

[0053] S2.1 For handling missing values, the SAITS time series imputation model based on the self-attention mechanism is used for imputation. In this model, 10% of the true values ​​are masked by an artificial mask, and the diagonal masking self-attention mechanism combined with the weighted combination module is used to effectively imput the missing spatiotemporal series data.

[0054] S2.2 For outlier processing, the density-based clustering algorithm DBSCAN is used to identify outliers.

[0055] S2.3 uses the max-min normalization method to normalize the data after imputation of missing values, mapping the data to the range of 0-1. This eliminates the influence of different units and dimensions between different feature indicators on model training, while accelerating model convergence and reducing training time. The specific formula is as follows:

[0056]

[0057] In the above formula, x * The data is normalized; x is a data point in the original dataset that needs to be normalized. min The minimum value in the dataset; x max This represents the maximum value in the dataset.

[0058] Furthermore, regarding the data acquisition in the land use subset D in step 4, this invention combines the areas where rainfall runoff pollution occurs with landscape ecology, utilizing landscape pattern distribution to obtain the corresponding land use distribution types. In this invention, a certain study area is considered as a complete landscape. The total landscape area (CA) represents the total area corresponding to the study area; the proportion of landscape area occupied by patches (PLAND) represents the proportion of the total area occupied by the three underlying surfaces of roads, roofs, and green spaces in the study area; and the landscape shape index (LSI) represents the degree of irregularity in the shape of roads, roofs, and green spaces in the study area. These three indicators are selected to depict the static characteristics of the study area and demonstrate the generation and flow state of rainfall runoff pollution on the land surface.

[0059] In this invention, the high-resolution remote sensing data imagery is sourced from Ovi Maps, and the specific operation steps are as follows:

[0060] Step S4.1 Based on the Ovi Map, locate the study area using high-resolution remote sensing map data at a depth of 30m. Select grid data to be analyzed with the same map parameters, including value grids and various environmental element grids. Divide the environmental element grids into three types: road surface, roof, and green space. Manually divide the underlying surface distribution type of the area and cut according to different underlying surface boundary rules.

[0061] S4.2 Load the remote sensing data vector map that has been divided in step S3.1, select Gaussian projection coordinates, and use ArcToolbox, Spatial Analyst Tools, Extraction and Extract by Mask tools in ArcMap software to input land use raster and division boundaries to generate a cropped land use map;

[0062] S4.3 Format Conversion: Convert the data that has been classified into underlying surface distribution types into vector data and grid data, classify and name the different underlying surface distribution types, and save tif. and shp. type files to obtain the pre-processed vector grid dataset D*;

[0063] Import the vector grid dataset D*, which has undergone preliminary processing in ArcGIS, into Fragstats software to calculate the landscape pattern index. The specific steps are as follows:

[0064] S4.4 Import Data: Enter the Fragstats software, import the data that has been preprocessed in the above steps, create a new project, select the grid data to be analyzed, set the Cell size to be consistent with ArcGIS, and usually set the Background value to NoData.

[0065] S4.5 Select the landscape pattern index to be calculated, which mainly includes the total landscape area, the proportion of landscape area occupied by patches, and the landscape shape index. Click Run to obtain the three landscape pattern indices required for the study area: total landscape area, proportion of landscape area occupied by patches, and landscape shape index.

[0066] Furthermore, in step 5, the K-Means algorithm, hierarchical clustering algorithm, and Gaussian mixture model are integrated to further improve the robustness and stability of the clustering results. The K-Means algorithm groups data points into k clusters represented by centroids, partitioning the data by minimizing the intra-cluster squared error, thus minimizing the sum of the squared distances from each data point to the centroid of its assigned cluster. The hierarchical clustering algorithm treats each sample as a cluster and progressively merges the most similar clusters. The Gaussian mixture model, as a soft clustering method based on probability distribution, estimates the Gaussian distribution function of each cluster by maximizing the likelihood function of the data. By integrating the K-Means algorithm, hierarchical clustering algorithm, and Gaussian mixture model, constructing a co-matrix, and applying consensus clustering, the advantages of hard and soft clustering can be effectively combined to solve the problem of algorithm application in complex data distribution scenarios such as multimodal data. The specific operation steps are as follows:

[0067] Step 5.1 Run the K-Means algorithm, hierarchical clustering algorithm, and Gaussian mixture model algorithm in a single run. Calculate the silhouette coefficient and the quantitative relationship between clusters in the K-Means and hierarchical clustering algorithms, as well as the BIC index distribution in the Gaussian mixture algorithm. Based on the silhouette coefficient distribution and BIC index in the K-Means, hierarchical, and Gaussian mixture model algorithms, when the number of clusters is X, the clustering results of the three algorithms produce similar characteristics. Take X as the number of clusters, and retain the preliminary result characteristics of the three algorithms in a single run.

[0068] Step 5.2 Run the K-Means algorithm H times under the condition of a random subset of the dataset. Each random run has a different dataset distribution and yields different results. Run the hierarchical clustering algorithm H times under random initial states and obtain different results. Run the Gaussian mixture model H times under random initial states. Based on the above random initial states and dataset distributions, select the most representative result and combine it with the number of clusters X determined in the initial results to obtain the K-Means algorithm clustering result. Hierarchical clustering results Clustering results with Gaussian mixture model

[0069] Step 5.3 Constructs a co-co-matrix based on the results obtained in steps 5.1 and 5.2, and counts the frequency of samples in the data set being assigned to the same cluster in different clustering methods to quantify their similarity.

[0070] ① Define indicator functions for the results of K-Means clustering and hierarchical clustering as follows:

[0071]

[0072] in, This represents the cluster label of sample i for the m-th clustering method. Let represent the cluster label of sample j for the m-th clustering method.

[0073] ② Define the probability consistency function for the clustering results of the Gaussian mixture model as follows:

[0074]

[0075] Where, p ik p represents the probability that sample i belongs to cluster k. jk Let J represent the probability that sample j belongs to cluster k, where K is the number of clusters, satisfying the condition...

[0076] ③ For K-Means clustering, hierarchical clustering, and Gaussian mixture models, the elements of the co-coordinate matrix are defined as follows:

[0077]

[0078] Where δ KM (i,j) represents the indicator function used by the K-Means clustering algorithm; δ HC (i,j) represents the indicator function for using the hierarchical clustering algorithm; Let be the probability consistency function of the Gaussian mixture model.

[0079] Step 5.4 Based on the co-coherence matrix, a hierarchical clustering algorithm is used to cluster the co-coherence matrix. The specific process is as follows:

[0080] ① Distance matrix transformation: Convert the co-coherence matrix into a distance matrix:

[0081] D ij =1-C ij

[0082] Among them, D ij It is the transformed distance matrix, C ij It represents the similarity between samples i and j in the co-co-matrix.

[0083] ② Use hierarchical clustering algorithm to analyze the distance matrix D ij Cluster analysis is performed, a hierarchical tree is constructed, and a bottom-up agglomerative algorithm is selected. The most similar clusters are merged step by step based on the distance between clusters. The distance between two clusters is defined as the average distance between all sample pairs of the two clusters.

[0084]

[0085] Where P and Q represent two clusters respectively; i represents a sample data in P; and j represents a sample data in Q.

[0086] Based on steps 5.1-5.4 above, the K-Means algorithm, hierarchical clustering algorithm, and Gaussian mixture model algorithm are integrated. By constructing a co-coherence matrix, and leveraging the advantage of hierarchical clustering in discovering hierarchical relationships between each cluster, hierarchical clustering is used again to perform cluster analysis on the co-coherence matrix. Considering the differences in distance metric calculation methods (i.e., different distance calculation methods) and applicable scenarios in the hierarchical clustering algorithms, and combining the type of data distribution in the obtained co-coherence matrix, a suitable distance metric is selected. Based on the above selection of clustering algorithms and distance metric, cluster analysis is performed to obtain the clustering results based on the dataset M^.

[0087] Table 1 Application scenarios of different distance metrics

[0088] Distance metric Applicable Scenarios Features Euclidean distance Continuous numerical data, commonly used in K-means, etc. Sensitive to outliers, calculate straight-line distance Maximum distance Suitable for scenarios that focus on the greatest difference Calculate the maximum value of two points Manhattan distance Grid space, urban grid distance Insensitive to outliers, calculated along the coordinate axis path Canberra Distance Sparse vectors emphasize small differences. Sensitive to zero values, suitable for sparse distances Binary distance binary features Considering only the difference between 0 and 1 Minkowski distance A universal distance metric applicable to various scenarios Flexible and diverse depending on the value of p.

[0089] Note: This is not a specific conclusion, but rather an appropriate choice based on the characteristics of each dataset when applying the method. This does not refer to clustering M^ according to any one of the methods in steps 5.1-5.4, but rather to the results after steps 5.1-5.4 (co-coupling matrix C). ij Cluster analysis was performed.

[0090] The evaluation of the clustering results in step 6 involves the following steps: based on the hierarchical clustering algorithm, using the co-coherence matrix C... ij The corresponding distance matrix D ij As the input matrix, the silhouette coefficient distribution map corresponding to different numbers of clusters is calculated to determine whether the results obtained based on the above ensemble clustering algorithm are reasonable. The closer the silhouette coefficient is to 1, the more reasonable the clustering result and the better the effect.

[0091] The aforementioned ensemble algorithm considers different spatial scales within the watershed and is applicable to urban scales, functional zone scales, and underlying surface scales, but is not limited to these. Taking urban scales, functional zone scales, and underlying surface scales as examples, the required data is shown in Table 2. As the types of data acquired increase, the indicator features selected for different scales can increase with data richness, and are not limited to those listed in Table 2.

[0092] Table 2 Data types at different spatial scales

[0093]

[0094] As new monitoring data is continuously added, enriching the basic dataset M, the pollution feature partitions obtained based on this method tend to stabilize. The newly acquired data are then combined with the basic dataset M to form a new dataset N for a new round of clustering, yielding new clustering results based on the new dataset N. The results of the new round of clustering Clustering results based on the base dataset M The results are compared to determine if the output is stable. When the clustering results are stable under sufficient data conditions, it is considered that a pollution characteristic standard zoning has emerged in the field of surface rainfall runoff pollution in the watershed. The specific process steps are shown in the figure below.

[0095] For step 8, which involves determining the different rainfall-runoff pollution distribution characteristics within the watershed based on clustering results, combined with different pollution features and geographical zoning, the specific operational steps are as follows:

[0096] Based on the differences in runoff pollution levels within each cluster and the spatial distribution of rainfall events within that cluster, the runoff pollution characteristics of each cluster are analyzed. The key characteristics and causes of runoff pollution differences in the study areas corresponding to different clusters are given, thus enabling pollution characteristic zoning under different spatial conditions within the watershed.

[0097] Those skilled in the art will understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the meaning consistent with their meaning in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as herein.

[0098] Based on the above-described preferred embodiments of the present invention, and through the foregoing description, those skilled in the art can make various changes and modifications without departing from the inventive concept. The technical scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.

Claims

1. A watershed runoff pollution characteristic dynamic integrated clustering partitioning method fusing multi-source spatio-temporal data and landscape ecological characteristics, characterized in that, It comprises the following steps: Step 1: dividing the spatial scale of rainfall runoff pollution characteristics; Step 2: collecting basic data to form a basic data set M; Step 3: preprocessing the basic data set M, filling in missing values, processing outliers and normalizing to obtain a data set M*; Step 4: calculating the required total landscape area, patch landscape area proportion and landscape shape index, adding the total landscape area, patch landscape area proportion and landscape shape index and geographic coordinate data of the study area to the data set M* to obtain a data set M^; Step 5: Comprehensive comparison and selection of different clustering methods, in the underlying scale, the data set M* is taken as the input matrix of the integrated clustering algorithm; in the functional area scale and the city scale, the data set M^ is taken as the input matrix of the integrated clustering algorithm, the feature partition of rainfall runoff pollution is calculated by the integrated clustering method, and the clustering result is obtained Step 6: Cluster result evaluation: for the cluster number X in the cluster result , the compactness within the cluster and the separation between clusters are measured by the silhouette coefficient and the sum of squared errors: the closer the silhouette coefficient corresponding to X is to 1, the more reasonable the cluster number is; a curve of the sum of squared errors changing with the cluster number is drawn, and when the sum of squared errors corresponding to X starts to slow down in the speed of decline with the cluster number, it means that X is the best cluster number, , the best cluster result. Step 7: After the basic data of the new area is added to the basic data set M, a new data set N is formed, and steps 2-6 are repeated at the same spatial scale to obtain a new round of clustering results Comparison of clustering results and Determine whether the output is stable. When the clustering results are stable under sufficient data conditions, it is considered that the new area has the pollution characteristic standard partition at the spatial scale.

2. The dynamic integrated clustering and zoning method for watershed runoff pollution characteristics that integrates multi-source spatiotemporal data and landscape ecological features according to claim 1, characterized in that: The step 1 is specifically: in the range of the watershed, according to the target range of pollution characteristic zoning, the study area is divided into urban scale, functional area scale and underlying surface scale.

3. The dynamic integrated clustering and zoning method for watershed runoff pollution characteristics that integrates multi-source spatiotemporal data and landscape ecological features as described in claim 1, characterized in that: The basic data in step 2 includes rainfall runoff pollution sub-data set A, meteorological element sub-data set B, urban human attribute sub-data set C, land use sub-data set D and underlying surface feature sub-data set E of the study area, which are combined to form the basic data set M; The rainfall runoff pollution sub-data set A includes the average concentration of COD, TN, TP and SS of each rainfall event; The meteorological element sub-data set B includes rainfall amount, rainfall duration, dry period before rain and average rainfall intensity of each rainfall event, and annual average rainfall and annual average temperature of the city where the study area is located; The urban human attribute sub-data set C includes urban per capita GDP, urban GDP output value, first industry proportion, second industry proportion and third industry proportion, and urban built-up area green coverage rate; The land use sub-data set D is high-resolution remote sensing image data of the study area; The underlying surface feature sub-data set E includes the material and type data of the underlying surface in the study area.

4. The method of claim 1, wherein the method further comprises: The data preprocessing process in step 3 is: using a time series interpolation model SAITS based on self-attention mechanism to fill in missing data; using a density-based clustering algorithm DBSCAN to eliminate outliers; using a maximum-minimum normalization method to map the data to 0-1; the specific steps are: S3.1 for missing value processing, using a time series interpolation model SAITS based on self-attention mechanism to fill in, in which 10% of the real values are masked using artificial mask, and the diagonal masking self-attention mechanism combined with the weighted combination module is used to effectively interpolate the missing data of the time-space sequence; S3.2 for outlier processing, using a density-based clustering algorithm DBSCAN to determine outliers; S3.3 using a maximum-minimum normalization method to normalize the data after missing value interpolation, mapping the data to 0-1, eliminating the influence of different dimensions and units on model training, speeding up model convergence, reducing model training time, and the specific formula is as follows: x * is the normalized data; x is a data to be normalized in the original data set, x min is the minimum value in the data set; x max is the maximum value in the data set.

5. The method of claim 1, wherein the method further comprises: The specific steps of data processing of the land use sub-data set D contained in the urban scale and functional area scale in step 4 are as follows: Step 4.1, 30m resolution high-definition remote sensing data images in the study area are downloaded using the Ovid map, the downloaded vector map is imported into ArcGIS software, the corresponding range of the study area is selected, the land use types in the study area are divided into road surface, roof and green land, and the shape is cut and divided to form a preliminary processed vector grid data set D*; Step 4.2, the preliminary processed vector grid data set D* is imported into Fragstats software, and the total area of the landscape, the proportion of the landscape area occupied by the patch and the landscape shape index in the landscape scale are calculated using Fragstats software, which respectively correspond to the total area of the study area, the proportion of road surface, roof and green land in the study area, and the irregularity of the shape of road surface, roof and green land in the functional area.

6. The method of claim 1, wherein the method further comprises: Step 5, in the selection of clustering method, three methods of K-Means clustering, hierarchical clustering and Gaussian mixture model based on distribution clustering algorithm are selected for integration, and the co-correlation matrix is used for re-clustering.

7. The method of claim 6, wherein the method further comprises: Step 5, the integrated clustering specifically includes the following steps: Step 5.1, for K-Means algorithm and hierarchical clustering, the silhouette coefficient and error sum of squares under different cluster numbers are calculated, and the curve graph of the change of the silhouette coefficient and the error sum of squares with the cluster number is drawn; for Gaussian mixture model, the Bayesian information criterion BIC is used to draw the distribution graph of different grouping numbers and BIC index; Step 5.2, respectively single run K-Means algorithm, hierarchical clustering algorithm and Gaussian mixture model, compare the running results of K-Means algorithm, hierarchical clustering algorithm and Gaussian mixture model when the cluster number is different, and the change relationship of the silhouette coefficient and BIC index with the cluster number in the corresponding algorithm, when the silhouette coefficient corresponding to the cluster number X is closer to 1, the error sum of squares begins to decrease, X is determined as the best cluster number, and the corresponding clustering result is optimal; Step 5.3, select H times to run K-Means algorithm, hierarchical clustering algorithm and Gaussian mixture model respectively, compare whether the running results are stable, if the optimal clustering group number changes after H times running, increase the running times, compare the clustering results, until the clustering results of K-Means algorithm, hierarchical clustering algorithm and Gaussian mixture model are stable, so as to solve the randomness problem caused by different initial states and different data sets; Step 5.4, using the optimal cluster number of step 5.2, respectively using K-Means algorithm, hierarchical clustering algorithm and Gaussian mixture model for clustering, generating corresponding clustering labels under different algorithms; construct the co-correlation matrix, record the frequency of samples being divided into the same cluster in multiple clustering, and record the similarity between samples; based on the advantage of hierarchical clustering in discovering the hierarchical relationship between each class, the hierarchical clustering is used for final division of the co-correlation matrix; compare the use conditions of different distance measurements such as Euclidean distance, Manhattan distance, Canberra distance and Chebyshev distance in hierarchical clustering algorithm, and select the appropriate distance measurement method combined with the data characteristics of the input matrix.

8. The method of claim 1, wherein the method further comprises: The meteorological element data, the urban human attribute data and the land use data corresponding to the new area are collected and added to the corresponding sub-data set to form a new data set N when the rainfall runoff pollution data of the new area are continuously added to the rainfall runoff pollution sub-data set A in step 2.

9. The method of claim 1, wherein the method further comprises: Based on the clustering results, different pollution characteristics and geographical divisions, the distribution characteristics of different rainfall runoff pollution in the basin range are determined under different spatial scales; the runoff pollution characteristics in each cluster are analyzed, the key characteristics and reasons for the difference in runoff pollution in different regions are given, and the pollution characteristic division under different spatial conditions in the basin range is realized.