Urban lake pollution source identification method and electronic device

CN122530798APending Publication Date: 2026-08-07GUANGZHOU INST OF GEOGRAPHY GUANGDONG ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-11
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

现有技术普遍缺少将上述城市语境信息与水体形态、水文连通关系进行建模的技术方案,导致城市湖泊对象无法被识别和监测

Benefits of technology

[0017]与现有技术相比,本发明具有以下有益效果:本发明结合水体对象的描述符、湖泊性评分模型与城市语境评分模型,可在河道、沟渠、湿地等复杂水体中稳定区分城市湖泊对象,解决传统仅依靠遥感光谱提取水体的缺陷,识别准确性与适应性大幅提升。将城市湖泊划分为若干湖泊单元,基于状态向量与期望状态向量精准提取异常斑块,解决了传统像元级监测不精准的问题,实现污染异常区域的精细化定位。构建湖泊-源区异构图,通过源区相似性评分与后验概率计算,实现排口、入湖河道、岸线片区等污染来源的精准判别,解决现有技术无法定位污染源头的问题。综合异常风险等级、执行成本与响应时间,为候选污染源区自动匹配核验方案与干预方案,实现监测—溯源—处置一体化,大幅提升城市湖泊污染治理的响应速度与处置效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122530798A_ABST
    Figure CN122530798A_ABST
Patent Text Reader

Abstract

The application discloses a kind of urban lake pollution source discrimination method, electronic equipment.The method is first obtained and preprocessed to obtain water body probability graph by multiple source data, according to this, candidate water body object is constructed and descriptor is calculated, combined with lake nature score model and city context score model accurately identify urban lake object;Urban lake is divided into several lake units, based on state vector and expected state vector, abnormal patch is extracted, and then candidate pollution area is constructed and lake-source area heterostructure graph is built;The source area similarity score of each abnormal patch is calculated by the heterostructure graph, the posterior probability of the candidate pollution source area is obtained, finally, combined with abnormal risk level, scheme execution cost and response time, for candidate pollution source area Matching verification scheme and intervention scheme.The present application can stably identify lake in complex urban water body, improve water body recognition accuracy;Can accurately distinguish pollution source, greatly improve the response speed and control efficiency of urban lake pollution treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of urban governance technology, and in particular to a method for identifying pollution sources in urban lakes and an electronic device. Background Technology

[0002] Urban lakes are important water bodies within cities, with clearly defined boundaries and frequent and complex interactions with humans, making them more susceptible to pollution and ecological stress, thus requiring focused monitoring. However, urban areas exhibit a complex mix of water body types, including not only natural or artificial lakes but also rivers, streams, canals, landscape waterways, reservoirs, wetlands, and seasonally flooded areas. Current technologies can extract water body boundaries using multispectral remote sensing imagery, but they cannot reliably distinguish urban lakes from similarly shaped objects such as rivers, canals, and landscape waterways in complex urban environments.

[0003] In addition to remote sensing spectroscopy, identifying urban lakes also requires consideration of urban contextual information such as surrounding land use, the proportion of impervious surfaces, road density, distribution of points of interest, population or travel activity, and administrative boundaries. Current technologies generally lack technical solutions for modeling the aforementioned urban contextual information in relation to water body morphology and hydrological connectivity, resulting in the inability to identify and monitor urban lakes.

[0004] Existing technologies, after identifying large water bodies, can usually only identify areas of water color or area change, or perform inversion and early warning of indicators such as chlorophyll, turbidity, suspended solids, and cyanobacteria risk. However, most technologies are still at the level of pixel-level water body extraction or single water quality monitoring, and cannot accurately identify pollution sources, including discharge outlets, rivers flowing into lakes, shoreline activity areas, construction areas, or catchment areas. Summary of the Invention

[0005] To address the problems existing in the prior art, the present invention aims to provide a method for identifying pollution sources in urban lakes and an electronic device capable of identifying urban lake objects from complex urban water bodies, performing status monitoring, extracting abnormal patches, identifying pollution sources, and providing inspection and intervention suggestions.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: A method for identifying pollution sources in urban lakes includes: Acquire multi-source data; Preprocess the multi-source data to obtain a water body probability map; Candidate water bodies are constructed based on the water body probability map, and the descriptors of the candidate water bodies are calculated. Based on descriptors, lake-related scoring models, and urban context scoring models, urban lake objects are identified. The urban lake object is divided into several lake units, and the state vector of each lake unit is obtained based on multi-source data; Extracting abnormal patches based on state vectors and expected state vectors; Based on multi-source data, several candidate pollution areas are selected for each abnormal patch to form a set of candidate pollution source areas. Lake units, lake units, candidate pollution areas, and abnormal patches are used as nodes, and spatial adjacency, drainage connectivity, confluence, wind direction consistency, and historical similarity relationships determined based on multi-source data are used as edges to construct a lake-source area heterogeneous map. The source region similarity score for each anomalous patch is calculated based on the lake-source region heterogeneity map and multi-source data. For each abnormal patch, the posterior probability of the candidate pollution source area is calculated based on the source area similarity score; Based on posterior probability, anomaly risk level, implementation cost and response time, verification and intervention plans are matched for candidate pollution source areas.

[0007] As a further improvement of the present invention, the multi-source data includes: remote sensing image data, hydrological and meteorological data, and fixed monitoring data; The steps of preprocessing multi-source data to obtain a water body probability map include: Generate a median composite image based on remote sensing image data within a 30-day time window; Low cloud composite images are generated based on remote sensing image data within a 10-day time window. Construct a seasonal baseline based on historical remote sensing image data from the same period; The reflectance, exponential features, texture features, and local image patch features of each band are calculated based on the median composite image and then input into the neural network to obtain the water body probability map.

[0008] As a further improvement of the present invention, the step of constructing candidate water bodies based on the water body probability map and calculating the descriptors of the candidate water body objects includes: Water body probability Figure 2 Values ​​are then processed, and morphological opening operations are used to remove isolated noise, closing operations are used to fill small holes, and connected component analysis is performed to obtain a set of candidate water bodies. The candidate water bodies are obtained by deleting objects with an area smaller than a set first threshold from the candidate water body object set and retaining small-area objects that match the existing lake management list. Descriptors for candidate water bodies are calculated based on remote sensing image data, urban contextual data, and hydrological and meteorological data.

[0009] As a further improvement of the present invention, the multi-source data includes: urban context data, and the step of determining the urban lake object based on descriptors, a lake-specific scoring model, and an urban context scoring model includes: Input the descriptor into the lake characterization scoring model and calculate the lake characterization score; Calculate urban context scores based on urban context data; Based on the lake characteristics score and the urban context score, the urban lake objects are identified.

[0010] As a further improvement of the present invention, the step of extracting abnormal patches based on state vectors and expected state vectors includes: Calculate the anomaly score based on the state vector and the expected state vector; When the abnormal score exceeds the second set threshold and the lake unit forms a connected region in space, it is extracted as an abnormal patch.

[0011] As a further improvement of the present invention, the present invention also includes: a step for calculating the desired state vector: A convolutional neural network model was used to extract spatial texture and local contextual information from 12 phases of low cloud composite images to obtain a time series. The time series is input into the Transformer time series model to obtain the prediction results; The expected state vector is obtained by weighting and fusing the prediction results with the seasonal baseline.

[0012] As a further improvement of the present invention, the step of calculating the source region similarity score of each anomalous patch based on the lake-source region heterogeneity map and multi-source data includes: Based on the lake-source region heterogeneity map, the source region similarity score is calculated according to hydrological distance, arrival time, wind direction consistency, rainfall confluence, and historical similarity.

[0013] As a further improvement of the present invention, the step of calculating the posterior probability of the candidate pollution source region based on the source region similarity score for each abnormal patch includes: Information is updated on the lake-source region heterogeneity map to obtain the embeddings of abnormal patch nodes and candidate pollution source region nodes; Based on the embedding of abnormal patch nodes, the embedding of candidate pollution source area nodes, and the source area similarity score, the posterior probability of each candidate pollution source area is calculated using a normalized exponential function. The formula for calculating the posterior probability is as follows:

[0014]

[0015] in, Indicates at time abnormal plaques Corresponding candidate pollution source areas The posterior probability of each anomalous patch; The corresponding set of candidate pollution source areas is denoted as ; Indicates abnormal plaques Node embedding; Indicates candidate pollution source areas Node embedding; Indicates candidate pollution source areas Pointing to abnormal plaques Source region similarity score; The weighting parameters represent the posterior probability calculation. Indicates feature splicing; Represents the set of candidate pollution source areas Any candidate pollution source area in the [context missing]. As a further improvement of the present invention, the method further includes the step of: After on-site verification, the posterior probability is corrected using the on-site verification results through a Bayesian update method.

[0016] The present invention also provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory is used to store computer programs; and the processor is used to implement the above-mentioned method for identifying urban lake pollution sources when executing the program stored in the memory.

[0017] Compared with existing technologies, this invention has the following advantages: It combines water body object descriptors, lake-specific scoring models, and urban context scoring models to stably distinguish urban lake objects in complex water bodies such as rivers, ditches, and wetlands, overcoming the shortcomings of traditional methods that rely solely on remote sensing spectroscopy for water body extraction, thus significantly improving identification accuracy and adaptability. By dividing urban lakes into several lake units and accurately extracting abnormal patches based on state vectors and desired state vectors, it solves the problem of inaccurate traditional pixel-level monitoring, achieving refined positioning of pollution anomaly areas. A lake-source region heterogeneity map is constructed, and through source region similarity scoring and posterior probability calculation, it achieves accurate identification of pollution sources such as discharge outlets, inflowing rivers, and shoreline areas, solving the problem that existing technologies cannot locate pollution sources. Considering anomaly risk levels, execution costs, and response time, it automatically matches verification and intervention schemes for candidate pollution source areas, achieving integrated monitoring-source tracing-treatment, significantly improving the response speed and treatment efficiency of urban lake pollution control. Attached Figure Description

[0018] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.

[0019] Figure 1 This is a flowchart of the urban lake pollution source identification method described in this invention; Figure 2 This is a schematic diagram showing the relationship between the urban lake object and the candidate pollution source map described in this invention. Detailed Implementation

[0020] To make the technical problems solved by the present invention, the technical solutions adopted, and the technical effects achieved clearer, the technical solutions of the embodiments of the present invention will be further described in detail below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] This invention provides a method for identifying pollution sources in urban lakes, such as... Figure 1 As shown, it includes: S1. Acquire multi-source data, including: remote sensing image data, urban contextual data, hydrological and meteorological data, and fixed monitoring data.

[0022] The remote sensing imagery data includes: Sentinel-2 L2A surface reflectance products as the primary data source, with a time span of no less than 12 months, used to construct urban lake objects. The spatial resolution of bands B2, B3, B4, and B8 is 10m, and the spatial resolution of bands B5, B6, B7, B8A, B11, and B12 is 20m, resampled to 10m. Landsat 8 / 9 Level-2 surface reflectance products can also be acquired to supplement the time series data under cloudy and rainy weather conditions. When it is necessary to improve local accuracy, high-resolution satellite, aerial imagery, or UAV multispectral / hyperspectral imagery can also be acquired as augmentation data. In key lake areas, UAV multispectral or hyperspectral data can also be overlaid as local fine-grained verification data.

[0023] Urban contextual data includes: WorldCover or GlobeLand30 land use products, impervious surface ratio data, road network data from OpenStreetMap or urban basic geographic information systems, point of interest (POI) data, administrative boundaries, existing lake directory data from urban parks or water or landscaping departments, and population activity or mobility data. Mobility data can be obtained using at least one of the following: public transport card swipes, mobile phone signaling, or rasterized travel intensity data. Hydrological and meteorological data includes: 30m resolution DEM, drainage pipe network or natural drainage network, catchment sub-area boundaries, lake inlet and outlet locations, hourly or daily rainfall data, and wind direction and speed data. Fixed monitoring data includes: buoy monitoring data (such as chlorophyll content), manual sampling data, or shoreline inspection records.

[0024] The aforementioned multi-source data are used in different steps: remote sensing imagery data is used to generate water body probability maps, construct candidate water bodies, calculate water body indices, and construct lake unit state vectors; urban context data is used for preserving small-area candidate water bodies, calculating urban context scores, screening management units, and identifying shoreline activity zones; hydrological and meteorological data is used to calculate hydrological connectivity, construct confluence relationship edges, wind direction consistency relationship edges, arrival time, and source area similarity scores; fixed monitoring data is used to correct lake unit state vectors, assist in the confirmation of anomalous patches, form historical similarity relationships, and correct the posterior probability of candidate pollution source areas after on-site verification. In key lake areas, high-resolution satellite, aerial imagery, or UAV multispectral / hyperspectral data can also be overlaid as local fine-grained identification or on-site verification data.

[0025] S2. Preprocess the multi-source data to obtain a water body probability map.

[0026] Specifically, the remote sensing image data is processed as follows: for Sentinel-2 L2A remote sensing images, clouds, shadows, cirrus clouds, and anomalous pixels are removed using the SCL scene classification layer and QA60 quality layer; for Landsat 8 / 9 Level-2 data, clouds and cloud shadows are removed using the QA_PIXEL quality layer; all images are uniformly projected to the CGCS2000 coordinate system, and the 20m or 30m bands are uniformly converted to a 10m grid using bilinear resampling or cubic convolution resampling.

[0027] To balance object identification stability and temporal monitoring sensitivity, two temporal composite strategies can be used to process remote sensing image data: for object identification, a 30-day time window is used to fuse reflectance images after cloud removal based on the median, in order to reduce the impact of short-term noise and occasional floating objects; for temporal monitoring, a 10-day time window is used to select low cloud cover images and combine them with historical images from the same period to generate a seasonal baseline.

[0028] Based on median composite images, reflectance, exponential features, texture features, and local image patch features for each band are calculated. Exponential features include Normalized Difference Water Index (NDWI), Modified Normalized Difference Water Index (MNDWI), Automatic Water Body Extraction Index (AWEI), Normalized Difference Vegetation Index (NDVI), Normalized Difference Building Index (NDBI), and Phytoplankton Index (FAI). These reflectance, exponential features, texture features, and local image patch features are then input into a convolutional neural network to obtain a water body probability map. , represented as:

[0029] in, For pixels A local image patch centered on the image, with a selectable size of 128×128 pixels; The sigmoid function is used; CNN(·) is a convolutional neural network feature extractor. Represents a pixel or superpixel unit The probability of water bodies; represents the weight matrix of the output features of the convolutional neural network; β1 to β6 represent the weight coefficients corresponding to each index feature; b represents the bias term; NDWI, MNDWI, AWEI, FAI, NDVu and NDBI represent the normalized difference water index, modified normalized difference water index, automatic water extraction index, phytoplankton index, normalized difference vegetation index and normalized difference building index at pixel u, respectively.

[0030] As one implementation method, convolutional neural networks employ U-Net, SegFormer, or pixel-level segmentation networks with equivalent functionality. Represents a pixel or superpixel unit.

[0031] S3. Construct candidate water bodies based on the water body probability map and calculate the descriptors of the candidate water bodies. Candidate water bodies refer to objects that are initially determined to potentially contain water surfaces, generated through remote sensing image segmentation or map features.

[0032] In this process, the water body probability map is first generated. Greater than the threshold The pixels are binarized, and then isolated noise is removed by morphological opening operation, small holes are filled by closing operation, and connected component analysis is performed to obtain a set of candidate water bodies. From the set of candidate water bodies, objects with an area smaller than a set first threshold are deleted, and small-area objects that have spatial overlap with at least one stable water body management unit such as the existing lake management list are retained to obtain candidate water body objects. For example, you can first delete objects with an area smaller than 2000m².

[0033] When calculating candidate water body object descriptors, remote sensing image data is used to determine object boundaries and calculate area, perimeter, compactness, elongation, water body permanence, and shoreline complexity; urban context data is used to calculate the proportion of impervious surfaces around the object, road density, POI density, mobility intensity, major land use types, and management directory relationships; hydrological and meteorological data are used to calculate hydrological connectivity and the connectivity between the object and rivers or channels based on the DEM, drainage network, inlet and outlet locations.

[0034] Candidate water bodies were calculated based on remote sensing imagery data, urban contextual data, and hydrological and meteorological data. The descriptor. Object permanence is optional and determined based on at least 12 months of time-series data. If an object only appears briefly after a rainstorm in terms of time sequence, then its permanence is determined. The lower level distinguishes it from stable lakes.

[0035] The descriptor includes: area ,perimeter Compactness Elongation Skeleton length Branching rate Permanent water bodies shoreline complexity and hydrological connectivity :

[0036]

[0037]

[0038] ; in, and Candidate water bodies The maximum and minimum eigenvalues ​​of the shape covariance matrix, Indicates the number of skeleton branch points. This represents the probability threshold for water bodies, and a value of 0.5 can be selected. Indicates the length of the statistical time series. Indicates candidate water bodies The set of cells contained; |Ωo| represents the number of cells; 1[·] represents an indicator function, which takes the value 1 when the condition in the parentheses is true, and takes the value 0 otherwise; ε represents a constant to avoid the denominator being zero.

[0039] S4. Based on descriptors, lake-like scoring models, and urban context scoring models, it is determined that urban lake objects, river and canal objects usually have higher elongation and branching rates, stronger directional flow connectivity, and lower compactness, while lake-like objects usually have higher closure and more stable object permanence.

[0040] Input the descriptor into the lake quality scoring model and calculate the lake quality score. Through lake-based scoring Lake-like objects are distinguished from river, stream, and canal objects. Lake-like objects are those that, based on a comprehensive assessment of their morphology and hydrological connectivity, more closely resemble lake characteristics than river or canal characteristics; lake-like scoring. The calculation formula is:

[0041] in, For the Sigmoid function, to These represent the weight parameters of the lake-specific scoring model; Indicates candidate water bodies The closure index; Indicates candidate water bodies Hydrological connectivity with waterways, drainage networks, or canals; This represents the graph embedding features constructed based on the relationship between candidate water bodies and surrounding land features and hydrology.

[0042] Calculating urban context score based on urban context data Score based on urban context To determine whether a lake is located within a city, an urban lake is defined as either meeting urban contextual requirements or listed in the city's management directory. Further, an urban contextual score is calculated based on factors such as the proportion of impervious surfaces, road density, POI density, mobility intensity, administrative boundary relationships, and management directory relationships within a 300m and 500m buffer zone around the candidate lake. Based on this, it is possible to distinguish between lakes located in urban built-up areas and lakes located in suburban, woodland, or agricultural areas, thus enabling urban context scoring. The formula for calculating the judgment is:

[0043] in, For the Sigmoid function, to The weight parameters represent the urban context scoring model. Indicates a closed-loop indicator. This indicates the proportion of impermeable surfaces within a 300m buffer zone extending from the candidate water body. This indicates the road density within the 500m buffer zone. This represents the density of points of interest within a 500m buffer zone. This indicates the mobility intensity within the extended 500m buffer zone. Indicates the main land use type codes around the lake. This is a management constraint indicating whether the object exists within an existing lake directory database. This represents the embedding features of the object graph neural network.

[0044] Based on lake characteristics score and urban context score, urban lakes are identified. Specifically, when the lake characteristics score... ≥0.60 and urban context rating When the value is ≥0.50, the candidate water bodies will be... Identified as an urban lake object; when candidate water body object If a lake already exists in the existing lake list data, the threshold conditions can be relaxed to maintain consistency.

[0045] After obtaining the urban lake objects, a list of urban lakes is constructed, and the urban lake objects are written into the object database. The object boundaries, object types, main surrounding land use types, and the data time range and model version used for object identification are recorded.

[0046] S5. Divide the urban lake object into several lake units, and obtain the state vector of each lake unit based on multi-source data.

[0047] As one implementation method, the lake unit is divided into 10m×10m regular grid units. Alternatively, when the lake area is small or the shoreline is complex, the SLIC superpixel segmentation method can be used to form superpixel units with an area of ​​100m² to 400m². At the same time, an inner shoreline zone with a width of 30m is generated inward along the lake shoreline, and an outer shoreline zone with a width of 30m to 60m is generated outward. A buffer zone with a radius of 50m to 100m is generated at the lake inlet and outlet.

[0048] Before condition monitoring, urban lakes often have adjacent buildings, hard shorelines, and tree canopy obstruction, making pixels near the shoreline susceptible to land reflectance contamination. To unify the reflectance characteristics acquired by different sensors at different times, overlapping observation samples or samples from stable areas of the lake can be used to estimate the uniformity parameters for each band. and In one alternative embodiment, Landsat and Sentinel-2 data are normalized. Water reflectance can be corrected using the following formula.

[0049]

[0050]

[0051] Where b represents the spectral band of the remote sensing image, u represents the lake unit or the pixel within the lake unit; R(b,u) represents the apparent water reflectance of band b before correction at unit u. This represents the water reflectance of band b at cell u after shoreline mixing correction; This represents the average reflectance of land pixels in the neighborhood of cell u in band b; The land mixing ratio coefficient is determined by the point diffusion function, shoreline distance, and land-water mixing ratio within cell u, and its value ranges from 0 to 1. This represents a positive constant that avoids a denominator of zero.

[0052] This represents the reflectance after the auxiliary sensor data has been normalized to the reflectance scale of the main sensor; This represents the original reflectivity of the auxiliary sensor at band b and unit u; and The linearly normalized slope parameter and intercept parameter corresponding to band b are respectively obtained by fitting stable water samples, stable land samples or pseudo-invariant target samples that overlap Landsat data and Sentinel-2 data in time and space.

[0053] Based on remote sensing data and fixed monitoring data, chlorophyll content (Chla), total suspended solids (TSM), phytoplankton index (FAI), turbidity index (NDTI), cyanobacteria index (NDCI), surface temperature, and lake area change were obtained. A state vector for the urban lake unit was constructed using at least four of these parameters. Chlorophyll content, total suspended matter, and algae index can be obtained from inversion formulas or machine learning regression models. Inversion formulas include:

[0054]

[0055]

[0056] in, , and These represent the reflectance in the red light, near-infrared, and short-wave infrared bands, respectively. , and These represent the center wavelengths of the corresponding wavebands; to , to Represents the inversion parameters obtained through sample fitting. S6. Extracting abnormal patches based on state vectors and expected state vectors. In this process, firstly based on state vectors... Expected state vector Calculate abnormal scores When the abnormal score exceeds the second set threshold and the lake unit forms a connected region in space, it is extracted as an abnormal patch.

[0057] As one implementation method, the desired state vector can be calculated through the following steps: A convolutional neural network (CNN) model was used to extract spatial texture and local contextual information from 12 periods of low-cloud composite images to obtain a time series. The low-cloud composite images underwent shoreline correction before being input into the CNN. The time series was then fed into a Transformer time series model to obtain prediction results. These prediction results were then fused with a seasonal baseline weighting to obtain the desired state vector. Convolutional neural network models can use 4 to 6 layers of convolutional encoders, while Transformer temporal models can use 2 to 4 layers of encoder structures.

[0058] Prediction results The expression is:

[0059] in, This represents the unit-level features extracted by the convolutional neural network.

[0060] Abnormal scores The expression is:

[0061] This represents the expected state vector obtained by weighted fusion of the seasonal baseline and the Transformer prediction results. Representation unit Historical state covariance matrix, Indicates an abnormal score. Optionally, when When a region exceeds a preset threshold and forms a connected region with an area of ​​not less than 400 m² in space, it is extracted as an abnormal patch.

[0062] Furthermore, abnormal patches can be further classified into chlorophyll-increased type, turbidity-increased type, algal bloom-accumulation type, nearshore black and odorous type, or area shrinkage type. Optionally, based on the state vector... The direction and magnitude of the increase or decrease of each component are labeled with a type. For example, when chlorophyll and phytoplankton increase at the same time, it can be labeled as an algal bloom-related anomaly. When turbidity and total suspended matter increase at the same time and the abnormal area is close to the lake inlet, it can be labeled as a turbidity inflow-related anomaly.

[0063] S7. Based on multi-source data, select several candidate pollution areas for each abnormal patch to form a set of candidate pollution source areas. Use lake units, lake units, candidate pollution areas, and abnormal patches as nodes, and use spatial adjacency relationships, drainage connectivity relationships, confluence relationships, wind direction consistency relationships, and historical similarity relationships determined based on multi-source data as edges to construct a lake-source area heterogeneous map.

[0064] For each abnormal plaque It is possible to extract the upstream direction, the adjacent shoreline zone, or the management unit. The candidate pollution source areas form a set of candidate pollution source areas, among which The options are 3 to 10.

[0065] As one implementation method, such as Figure 2 As shown, candidate pollution areas can be generated from the list of urban pipe network outlets, river channels flowing into the lake, commercial and service areas along the shoreline, areas with multiple construction sites, and the boundaries of upstream catchment sub-areas of the lake. The nodes of the lake-source region heterogeneity map include: lake units, lake clusters, candidate pollution areas, and anomalous patches. The edges of the lake-source region heterogeneity map include: spatial adjacency edges, drainage connectivity edges, confluence relationship edges, wind direction consistency edges, and historical similarity edges. Through the lake-source region heterogeneity map, the spatial, hydrological, meteorological, and historical relationships between anomalous patches and candidate pollution source areas can be uniformly encoded.

[0066] Among these criteria, spatial adjacency relationships are determined based on the spatial location relationships of lake units, anomalous patches, and candidate pollution source areas; drainage connectivity relationships are determined based on drainage pipe networks, natural drainage networks, lake inlets, and outlet locations; runoff relationships are determined based on DEM, catchment sub-region boundaries, and rainfall data; wind direction alignment relationships are determined based on wind direction and speed data, as well as the azimuth angle of candidate pollution source areas pointing towards anomalous patches; and historical similarity relationships are determined based on historical anomaly records, fixed monitoring data, and on-site verification results. S8. Calculate the source region similarity score for each anomalous patch based on the lake-source region heterogeneity map and multi-source data. Source region similarity score Simultaneously considering hydrological distance, time of arrival, wind direction consistency, rainfall in the past 24 hours, and historical similarity, the similarity score is increased if a candidate pollution source area has a direct drainage connection with the anomalous patch; conversely, the similarity score is decreased if the source area and the anomalous patch are oriented significantly away from the current wind direction and confluence direction. Source Area Similarity Score The calculation formula is as follows:

[0067] in, Indicates time From candidate pollution source areas Pointing to abnormal plaques Similarity score, Indicates hydrological distance, Indicates arrival time. Indicates from candidate pollution source areas Pointing to abnormal plaques The direction angle, Indicates the wind direction angle. This indicates the cumulative rainfall over the past 24 hours. Indicates historical similarity.

[0068] The source area similarity score between candidate pollution source area m and abnormal patch p is obtained. Subsequently, the lake-source region heterogeneous graph is input into a heterogeneous graph neural network for node information updates, incorporating information from spatial adjacency, drainage connectivity, runoff relationships, wind direction consistency, and historical similarity. The node update formula can be expressed as:

[0069] Represents the nodes in the graph In the The feature vector of the layer, This represents the posterior probability of a candidate pollution source area.

[0070] Where v represents any node in the lake-source region heterogeneous graph, and the node types include lake units, candidate pollution source regions, and anomalous patches; u represents the node adjacent to node v; and l represents the layer index of the graph neural network. This represents the feature vector of node v at layer l. This represents the feature vector of node v after the update at layer l+1; This represents a set of edge relationship types, which include at least one of the following: spatial adjacency, drainage connectivity, confluence, wind direction alignment, and historical similarity. This represents the set of nodes that are adjacent to node v through relation type r; This represents the attention weight or normalized weight when node u transmits information to node v through relation type r. This represents the transformation matrix of a node's own characteristics. This represents the feature transformation matrix corresponding to relation type r; This represents a non-linear activation function.

[0071] After updating through one or more heterogeneous graph neural networks, the final layer feature vector of the abnormal patch node p is taken as hp, and the final layer feature vector of the candidate pollution source area node m is taken as hm. hp and hm are then used for subsequent posterior probability calculation of the candidate pollution source area.

[0072] S9. For each abnormal patch, calculate the posterior probability of the candidate pollution source area based on the source area similarity score.

[0073] The formula for calculating the posterior probability is:

[0074] in, This represents the posterior probability of the candidate pollution source region m corresponding to the abnormal patch p at time t; This represents the set of candidate pollution source regions corresponding to the abnormal patch p; Represents the set of candidate pollution source areas Any candidate pollution source area in the list; , , This represents the unnormalized association score of the candidate pollution source area m relative to the abnormal patch p; This represents the node embedding features of the abnormal patch node p after being updated by the heterogeneous graph neural network; This represents the node embedding features of candidate pollution source area node m after being updated by a heterogeneous graph neural network; The similarity score of the source regions pointing from candidate pollution source region m to abnormal plaque p at time t is represented. Represents the weight vector or weight matrix used to calculate the posterior probability; Indicates the bias term; Indicates feature splicing; This represents an exponential function.

[0075] Through the above normalization process, the sum of the posterior probabilities of all candidate pollution source areas corresponding to the same abnormal patch p is 1. Further processing can be performed according to... Candidate pollution source areas are ranked from highest to lowest probability, and the candidate pollution source area with the highest posterior probability is determined as the priority verification source area. As one implementation method, a heterogeneous graph neural network can be used to update the lake-source region heterogeneous map to obtain the posterior probability of the candidate pollution source areas. In one optional embodiment, the graph neural network may have 3 layers, the hidden dimension may be 64 or 128, and the relationship types may include at least 4 categories: spatial adjacency, drainage connectivity, consistent wind direction, and historical similarity.

[0076] The information update of the lake-source region heterogeneous map specifically includes: using the initial node features of lake units, anomalous patches, and candidate pollution source regions as node features, and using spatial adjacency, drainage connectivity, confluence, wind direction consistency, and historical similarity as edge relationship types, and combining the source region similarity score, message passing, weighted aggregation, and feature updating are performed on the adjacent node features under different edge relationship types to obtain the node embeddings of anomalous patches and candidate pollution source regions.

[0077] S10. Based on the posterior probability, the level of abnormal risk, the implementation cost and response time of the plan, match the verification plan and the intervention plan for the candidate pollution source area.

[0078] The verification plan includes: drone aerial survey, dense deployment of buoys, portable sampling and outlet inspection; the intervention plan includes: temporary interception, aeration and oxygenation, pollution interception or emergency containment.

[0079]

[0080]

[0081] in, This represents the utility value of implementing plan a. This indicates the information gain that can be achieved. This indicates a risk of reduced expectations. Indicates execution cost, Indicates response time. to The weighting coefficients representing information gain, risk reduction, execution cost, and response time in the utility function can be preset based on governance objectives or historical experience data. This represents a candidate verification or intervention plan, where a* indicates the plan with the highest utility value.

[0082] If the posterior probabilities of candidate pollution source areas are relatively dispersed, verification schemes with greater information gain can be selected, such as near-real-time aerial surveys by UAVs, outlet inspections, or increased buoy density. If a candidate pollution source area has a high posterior probability and a high level of abnormal risk, interventions such as temporary interception, aeration, or pollution containment can be initiated simultaneously with verification. For outlet-type candidate pollution source areas, options include outlet inspections, portable water sampling, and upstream pipeline investigation. For river-type candidate pollution source areas flowing into the lake, options include UAV aerial surveys along the river, river cross-section sampling, and temporary interception. For shoreline activity areas or construction areas, options include on-site inspections, sediment containment inspections, and verification of shoreline interception measures.

[0083] S11. After on-site verification, the posterior probability is corrected using the on-site verification results through a Bayesian update method, as shown in the following formula:

[0084] in, Indicates the candidate pollution source area before verification The posterior probability, This represents the posterior probability obtained after verification and correction. This indicates the results of the on-site verification. Indicating in candidate pollution source areas Verification results are generated under actual pollution source conditions. The likelihood probability; This represents any candidate pollution source area in the set of candidate pollution source areas.

[0085] Finally, the lake objects, anomalous patches, candidate pollution source areas, and intervention plans are written into the object and source area time-series databases. Optionally, the model version, threshold parameters, input data time range, and field verification results used in each step are also recorded for subsequent queries.

[0086] The invention will now be described in detail with reference to specific implementation processes: Sentinel-2 L2A and Landsat 8 / 9 Level-2 imagery of a certain city were selected, covering a continuous time range of 18 months. After sampling all remote sensing image data to a 10m grid, a stable reflectance layer was generated using a 30-day median composite image. Candidate water bodies were then constructed by combining land use, impervious surfaces, road networks, POIs, and mobility data.

[0087] A candidate With an area of ​​approximately 0.82 km², it is compact. =0.61, elongation =2.30, branch rate =0.08, water body permanence =0.93, =0.12; its 300m buffer zone impermeable surface ratio is 0.46, its 500m buffer zone road density is 8.7km / km², its POI density is 245 / km², and its mobility intensity normalized value is 0.58. Substituting these values ​​into the lake-related scoring and urban context scoring models, we obtain... =0.84, =0.79, therefore... It was identified as an urban lake.

[0088] In contrast, a certain candidate Although it meets the requirements for water in terms of spectrum, its elongation... =9.40, branching rate =0.52, =0.81, exhibiting clear river channel characteristics, ultimately yielding The results indicate that this invention does not identify objects as water bodies solely based on spectral density, but further combines object morphology, hydrological connectivity, and urban contextual information to distinguish urban lake objects.

[0089] Continuous monitoring was conducted on the identified urban lakes, which were then divided into 10m × 10m lake units, and shoreline zones within 30m were generated. After mixed pixel correction of the nearshore units, chlorophyll content (Chla), total suspended matter (TSM), phytoplankton index (FAI), turbidity index (NDTI), cyanobacteria index (NDCI), surface temperature, and lake area changes were calculated. The current desired state was then predicted using an 18-period Transformer model.

[0090] After a 24-hour cumulative rainfall of 42 mm, an anomaly appeared near the lake's eastern inlet. Calculations showed that the chlorophyll content in this area increased from 18.2 to 41.6, and the NDTI increased from 0.12 to 0.31. The anomaly score... =3.4, and formed a connected region with an area of ​​approximately 2600 m² in space, thus being identified as an anomalous patch related to turbidity inflow.

[0091] Furthermore, this invention extracts candidate pollution source areas from the discharge outlet M1, the river channel flowing into the lake M2, the construction area M3, and the shoreline commercial area M4, and calculates the posterior probability using a lake-source heterogeneity map. The results show that... =0.18, =0.56, =0.17, =0.09, with the M2 channel flowing into the lake ranking highest.

[0092] Based on the utility function, a combined scheme of drone verification along the inflow river and portable sampling at the inflow point is output. Subsequent field verification results show that inflow river M2 carries high-turbidity runoff into the lake after rainfall. After Bayesian update... =0.81, the system further generates suggestions for temporary interception of the river flowing into the lake and aeration at the lake inlet.

[0093] In summary, the method for identifying urban lake pollution sources of the present invention has the following beneficial effects: 1. This invention does not only extract water bodies based on pixel spectral analysis, but also uses water body morphology, hydrological connectivity and urban context information together to identify urban lakes, which can effectively distinguish urban lakes from rivers, ditches, landscape waterways and seasonal waterlogging areas.

[0094] 2. This invention directly uses the results of urban lake object identification as the basis for subsequent status monitoring, abnormal patch extraction, and pollution source identification, enabling the identification, monitoring, and identification processes to run under the same system, thereby improving the stability of urban lake monitoring and the traceability of monitoring results.

[0095] 3. This invention models remote sensing anomaly patches with outfalls, rivers flowing into the lake, shoreline activity areas, construction areas, and catchment sub-areas using a lake and source area map model. It can provide a ranking of candidate pollution sources when anomalies occur and automatically select inspection and intervention suggestions, thereby improving the efficiency of urban lake management.

[0096] 4. This invention uses an object-source area time series database to continuously track urban lake objects, monitoring results, source area ranking, and intervention scheme results, which can be used for subsequent model iteration, risk analysis, and governance assessment.

[0097] Based on the same inventive concept, the present invention also provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory is used to store computer programs; and the processor is used to implement the above-described method for identifying urban lake pollution sources when executing the program stored in the memory.

[0098] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc.

[0099] The communication interface is used for communication between the aforementioned terminal and other devices.

[0100] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0101] The aforementioned processors can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. Field-Programmable Gate Array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0102] In this description, references to terms such as "an embodiment," "example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example.

[0103] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style of the specification is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

[0104] The technical principles of the present invention have been described above with reference to specific embodiments. These descriptions are merely for explaining the principles of the invention and should not be construed as limiting the scope of protection of the invention in any way. Based on this explanation, those skilled in the art can readily conceive of other specific embodiments of the invention without inventive effort, and these embodiments will all fall within the scope of protection of the present invention.

Claims

1. A method for identifying pollution sources in urban lakes, characterized in that, include: Acquire multi-source data; Preprocess the multi-source data to obtain a water body probability map; Candidate water bodies are constructed based on the water body probability map, and the descriptors of the candidate water bodies are calculated. Based on descriptors, lake-related scoring models, and urban context scoring models, urban lake objects are identified. The urban lake object is divided into several lake units, and the state vector of each lake unit is obtained based on multi-source data; Extracting abnormal patches based on state vectors and expected state vectors; Based on multi-source data, several candidate pollution areas are selected for each abnormal patch to form a set of candidate pollution source areas. Lake units, candidate pollution areas, and abnormal patches are used as nodes, and spatial adjacency, drainage connectivity, confluence, wind direction consistency, and historical similarity relationships determined based on multi-source data are used as edges to construct a lake-source area heterogeneous map. The source region similarity score for each anomalous patch is calculated based on the lake-source region heterogeneity map and multi-source data. For each abnormal patch, the posterior probability of the candidate contaminated area is calculated based on the source region similarity score; Based on posterior probability, anomaly risk level, implementation cost and response time, verification and intervention plans are matched for candidate pollution source areas.

2. The method for identifying pollution sources in urban lakes according to claim 1, characterized in that, The multi-source data includes: remote sensing image data, hydrological and meteorological data, and fixed monitoring data; The steps of preprocessing multi-source data to obtain a water body probability map include: Generate a median composite image based on remote sensing image data within a 30-day time window; Low cloud composite images are generated based on remote sensing image data within a 10-day time window. Construct a seasonal baseline based on historical remote sensing image data from the same period; The reflectance, exponential features, texture features, and local image patch features of each band are calculated based on the median composite image and then input into the neural network to obtain the water body probability map.

3. The method for identifying pollution sources in urban lakes according to claim 1, characterized in that, The steps of constructing candidate water bodies based on the water body probability map and calculating the descriptors of the candidate water body objects include: The water body probability map is binarized, and then isolated noise is removed by morphological opening operation, small holes are filled by closing operation, and connected component analysis is performed to obtain a set of candidate water bodies. The candidate water bodies are obtained by deleting objects with an area smaller than a set first threshold from the candidate water body object set and retaining small-area objects that match the existing lake management list. Descriptors for candidate water bodies are calculated based on remote sensing image data, urban contextual data, and hydrological and meteorological data.

4. The method for identifying pollution sources in urban lakes according to claim 1, characterized in that, The multi-source data includes: urban context data. The step of determining urban lake objects based on descriptors, lake-specific scoring models, and urban context scoring models includes: Input the descriptor into the lake characterization scoring model and calculate the lake characterization score; Calculate urban context scores based on urban context data; Based on the lake characteristics score and the urban context score, the urban lake objects are identified.

5. The method for identifying pollution sources in urban lakes according to claim 2, characterized in that, The step of extracting abnormal patches based on state vectors and expected state vectors includes: Calculate the anomaly score based on the state vector and the expected state vector; When the abnormal score exceeds the second set threshold and the lake unit forms a connected region in space, it is extracted as an abnormal patch.

6. The method for identifying pollution sources in urban lakes according to claim 5, characterized in that, It also includes: the steps for calculating the desired state vector: A convolutional neural network model was used to extract spatial texture and local contextual information from 12 phases of low cloud composite images to obtain a time series. The time series is input into the Transformer time series model to obtain the prediction results; The expected state vector is obtained by weighting and fusing the prediction results with the seasonal baseline.

7. The method for identifying pollution sources in urban lakes according to claim 1, characterized in that, The step of calculating the source region similarity score for each anomalous patch based on the lake-source region heterogeneity map and multi-source data includes: Based on the lake-source region heterogeneity map, the source region similarity score is calculated according to hydrological distance, arrival time, wind direction consistency, rainfall confluence, and historical similarity.

8. The method for identifying pollution sources in urban lakes according to claim 1, characterized in that, The step of calculating the posterior probability of candidate pollution source regions based on source region similarity scores for each abnormal patch includes: Information is updated on the lake-source region heterogeneity map to obtain the embeddings of abnormal patch nodes and candidate pollution source region nodes; Based on the embedding of abnormal patch nodes, the embedding of candidate pollution source area nodes, and the source area similarity score, the posterior probability of each candidate pollution source area is calculated using a normalized exponential function. The formula for calculating the posterior probability is as follows: in, Indicates at time abnormal plaques Corresponding candidate pollution source areas The posterior probability of each anomalous patch; The corresponding set of candidate pollution source areas is denoted as ; Indicates abnormal plaques Node embedding; Indicates candidate pollution source areas Node embedding; Indicates candidate pollution source areas Pointing to abnormal plaques Source region similarity score; The weighting parameters represent the posterior probability calculation. Indicates feature splicing; Represents the set of candidate pollution source areas Any candidate pollution source area in the list.

9. The method for identifying pollution sources in urban lakes according to claim 1, characterized in that, It also includes the following steps: After on-site verification, the posterior probability is corrected using the on-site verification results through a Bayesian update method.

10. An electronic device, characterized in that, The system includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus; the memory is used to store computer programs; and the processor is used to execute the program stored in the memory to implement the urban lake pollution source identification method as described in any one of claims 1 to 9.