Water body pollution space-time distribution dynamic monitoring system based on multi-source remote sensing data fusion and deep learning

The dynamic monitoring system for the spatiotemporal distribution of water pollution, which integrates multi-source remote sensing data fusion and deep learning, has solved the problem of insufficient accuracy in monitoring water pollution in the tributaries and lakeside areas of the middle and lower reaches of the Yangtze River. It has achieved accurate identification of small farmland drainage outlets and decentralized aquaculture sewage discharge points, as well as accurate prediction of pollution diffusion trajectories, thus improving the timeliness and accuracy of the monitoring system.

CN121527640APending Publication Date: 2026-02-13QUJING NORMAL UNIV

Patent Information

Application Number
CN202511695205.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing remote sensing monitoring technologies have insufficient accuracy in monitoring water pollution in the tributaries and lakeside areas of the middle and lower reaches of the Yangtze River. In particular, it is difficult to identify small farmland drainage outlets and decentralized aquaculture discharge points. Furthermore, ground monitoring points are difficult to adjust dynamically, resulting in deviations in the accuracy of pollution source location and the description of pollution diffusion trajectories.

Method used

A dynamic monitoring system for the spatiotemporal distribution of water pollution, employing multi-source remote sensing data fusion and deep learning, collects data in real time from satellite hyperspectral imagery, UAV multispectral data, and ground-based IoT devices. By combining spatial discretization and quadtree index recursive partitioning, dynamic spatial correction parameters are generated to optimize the pollutant diffusion model. Furthermore, a spatiotemporal convolutional neural network is used to extract the spatiotemporal co-evolutionary features of pollutants and predict their diffusion time series.

Benefits of technology

It improves the accuracy of monitoring the spatiotemporal distribution of water pollution and the reliability of pollution early warning, adapts to the spatial heterogeneity of different water areas, reduces data-level errors, shortens the feature extraction and prediction cycle, and enhances the timeliness and accuracy of pollution situation prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121527640A_ABST
    Figure CN121527640A_ABST
Patent Text Reader

Abstract

The invention provides a water body pollution spatial-temporal distribution dynamic monitoring system based on multi-source remote sensing data fusion and deep learning, and relates to the technical field of data processing, and the system comprises the steps: obtaining a satellite hyperspectral image, unmanned aerial vehicle multispectral data, and water quality parameter data collected in real time through ground Internet of Things equipment; the inversion module is used for performing pollutant concentration inversion on the satellite hyperspectral image to obtain pollutant concentration data; performing local pollution source positioning on the unmanned aerial vehicle multispectral data to obtain pollution source position data; acquiring real-time water quality parameter data from ground Internet of Things equipment; the spatial discretization module is used for carrying out spatial discretization processing on pollutant concentration data, pollution source position data and real-time water quality parameter data, establishing a quadtree index and carrying out recursion division on the quadtree index into sub-regions; and generating a dynamic space correction parameter according to the statistical characteristics of the data points in each sub-region. The method can effectively improve the water body pollution time-space distribution monitoring precision and the pollution early warning reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a dynamic monitoring system for the spatiotemporal distribution of water pollution based on multi-source remote sensing data fusion and deep learning. Background Technology

[0002] Precise and dynamic monitoring of water pollution is a crucial foundation for the prevention and control of agricultural non-point source pollution in the tributaries and lakeside areas of the middle and lower reaches of the Yangtze River. Pollution from farmland runoff and aquaculture discharge in these areas is characterized by its dispersed and concealed nature, requiring high spatiotemporal adaptability of monitoring technologies. Although related monitoring has integrated remote sensing and ground-based monitoring methods, limitations remain in practical applications. In terms of remote sensing monitoring, satellite and UAV remote sensing have become common tools for pollution investigation in this region, but their synergistic effectiveness has not been fully realized. While satellite remote sensing can cover the entire basin, its conventional field of view (10 to 30 meters) is limited. Spatial resolution presents challenges in accurately identifying hidden pollution sources such as small farmland drainage outlets and dispersed aquaculture sewage discharge points. While drones can achieve high-resolution imaging within 1 meter and capture details of sewage discharge in polder ditches, their limited battery life makes it difficult to continuously track the process of pollutants flowing into the main stream after heavy rain via tributaries. In existing technologies, the two types of remote sensing data are mostly simple superpositions of large-scale scans and local snapshots, lacking collaborative calibration for spatial heterogeneity such as tributary bends and polder shallows, resulting in deviations in the accuracy of pollution source location and the description of pollution diffusion trajectories.

[0003] In terms of the linkage between remote sensing and ground monitoring, the coordination mechanism between the two is still imperfect. Ground IoT monitoring points are mostly statically deployed in the early stage, making it difficult to dynamically adjust them according to the pollution anomaly areas identified by satellites or drones. For example, in the monitoring of the Poyang Lake reclamation area, satellite remote sensing can quickly locate areas with abnormal total phosphorus concentrations, but ground monitoring points often cannot keep up in time due to their fixed layout, resulting in insufficient collection of water quality parameters in the core pollution area, making it difficult to verify the accuracy of remote sensing inversion results. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a dynamic monitoring system for the spatiotemporal distribution of water pollution based on multi-source remote sensing data fusion and deep learning, which can effectively improve the monitoring accuracy of the spatiotemporal distribution of water pollution and the reliability of pollution early warning.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: Firstly, a dynamic monitoring system for the spatiotemporal distribution of water pollution based on multi-source remote sensing data fusion and deep learning includes: The acquisition module is used to acquire satellite hyperspectral imagery, UAV multispectral data, and water quality parameter data collected in real time through ground-based IoT devices; The execution module is used to invert pollutant concentration from satellite hyperspectral imagery to obtain pollutant concentration data; locate local pollution sources from UAV multispectral data to obtain pollution source location data; and acquire real-time water quality parameter data from ground-based IoT devices. The processing module is used to spatially discretize pollutant concentration data, pollution source location data, and real-time water quality parameter data, establish a quadtree index, and recursively divide it into sub-regions; and generate dynamic spatial correction parameters based on the statistical characteristics of data points in each sub-region. The optimization module is used to optimize the parameters of the pollutant diffusion numerical model constructed by fusing multi-source data through data assimilation technology based on the dynamic spatial correction parameters, so as to obtain the optimized pollutant diffusion model. The extraction module is used to input the spatiotemporal sequence data output by the optimized pollutant diffusion model into the spatiotemporal convolutional neural network to extract the spatiotemporal co-evolution features of pollutants and obtain the identification results. The evaluation module is used to input the identification results and the real-time status of the optimized pollutant diffusion model into a spatiotemporal convolutional neural network to predict the pollutant diffusion time series, and obtain risk level distribution data through risk level quantification and spatial interpolation.

[0006] In a second aspect, a computing device includes: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, enable the one or more processors to implement the system.

[0007] Thirdly, a computer-readable storage medium storing a program that, when executed by a processor, implements the system.

[0008] The above-described solution of the present invention has at least the following beneficial effects: By recursively dividing sub-regions through spatial discretization and quadtree indexing, and generating dynamic spatial correction parameters based on the statistical characteristics of sub-region data, it can adapt to the spatial heterogeneity of different water areas, such as tributary bends and shoals, avoiding the spatial bias problem of direct superposition of multi-source data and reducing error propagation at the data level. Through deep processing of the spatiotemporal sequence data output by the spatiotemporal convolutional neural network, it can more accurately capture the spatiotemporal co-evolution law of pollutants. Furthermore, by combining the real-time status of the model, it can realize diffusion time series prediction, which can shorten the feature extraction and prediction cycle and improve the timeliness and accuracy of pollution situation prediction. Attached Figure Description

[0009] Figure 1 This is a schematic diagram of a dynamic monitoring system for the spatiotemporal distribution of water pollution based on multi-source remote sensing data fusion and deep learning, provided by an embodiment of the present invention.

[0010] Figure 2 This is a flowchart illustrating the process of spatially discretizing pollutant concentration data, pollution source location data, and real-time water quality parameter data according to an embodiment of the present invention, establishing a quadtree index and recursively dividing it into sub-regions; and generating dynamic spatial correction parameters based on the statistical characteristics of data points in each sub-region. Detailed Implementation

[0011] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0012] like Figure 1 As shown, embodiments of the present invention propose a dynamic monitoring system for the spatiotemporal distribution of water pollution based on multi-source remote sensing data fusion and deep learning, comprising: The acquisition module is used to acquire satellite hyperspectral imagery, UAV multispectral data, and water quality parameter data collected in real time through ground-based IoT devices; The execution module is used to invert pollutant concentration from satellite hyperspectral imagery to obtain pollutant concentration data; locate local pollution sources from UAV multispectral data to obtain pollution source location data; and acquire real-time water quality parameter data from ground-based IoT devices. The processing module is used to spatially discretize pollutant concentration data, pollution source location data, and real-time water quality parameter data, establish a quadtree index, and recursively divide it into sub-regions; and generate dynamic spatial correction parameters based on the statistical characteristics of data points in each sub-region. The optimization module is used to optimize the parameters of the pollutant diffusion numerical model constructed by fusing multi-source data through data assimilation technology based on the dynamic spatial correction parameters, so as to obtain the optimized pollutant diffusion model. The extraction module is used to input the spatiotemporal sequence data output by the optimized pollutant diffusion model into the spatiotemporal convolutional neural network to extract the spatiotemporal co-evolution features of pollutants and obtain the identification results. The evaluation module is used to input the identification results and the real-time status of the optimized pollutant diffusion model into a spatiotemporal convolutional neural network to predict the pollutant diffusion time series, and obtain risk level distribution data through risk level quantification and spatial interpolation.

[0013] In this embodiment of the invention, by spatial discretization and recursively dividing sub-regions using quadtree indexing, and combining the statistical characteristics of sub-region data to generate dynamic spatial correction parameters, it can adapt to the spatial heterogeneity of different water areas, such as tributary bends and shoals, avoiding the spatial bias problem of direct superposition of multi-source data and reducing error propagation at the data level. By using a spatiotemporal convolutional neural network to deeply process the spatiotemporal sequence data output by the model, it can more accurately capture the spatiotemporal co-evolution law of pollutants. Furthermore, by combining the real-time state of the model to realize diffusion time series prediction, it can shorten the feature extraction and prediction cycle and improve the timeliness and accuracy of pollution situation prediction.

[0014] In a preferred embodiment of the present invention, acquiring satellite hyperspectral imagery, UAV multispectral data, and water quality parameter data collected in real time via ground-based IoT devices specifically includes: selecting satellites carrying hyperspectral sensors, which need to be able to cover the tributaries of the middle and lower reaches of the Yangtze River and the lakeside area; common choices include the Gaofen-5 satellite or the Sentinel-2A satellite; determining the image acquisition period, avoiding rainy weather, as rainy weather leads to excessively high atmospheric water vapor content, affecting image quality; receiving the raw hyperspectral image data transmitted by the satellite through a satellite ground receiving station at a fixed time each day when the satellite passes over the area; after receiving the data, performing radiometric calibration on the raw data, using the calibration coefficients built into the satellite sensor to convert the digital quantization value of each pixel in the image into surface reflectance; specifically, the conversion method is to use each pixel... The digital quantization value is multiplied by the corresponding calibration coefficient, and then added to the sensor's preset offset to obtain the surface reflectance corresponding to that pixel. Atmospheric correction is performed by acquiring atmospheric parameters of the area at the time of image capture using meteorological monitoring equipment. These parameters include atmospheric aerosol concentration and atmospheric water vapor content. Based on the principle of atmospheric radiative transfer, the absorption and scattering coefficients of the atmosphere for different spectral bands are calculated using the acquired atmospheric parameters. These coefficients are then used to correct the radiometrically calibrated image, removing the influence of atmospheric water vapor and aerosols on the image, resulting in a surface reflectance image that reflects the true surface conditions. Based on the administrative boundary data of the Yangtze River's middle and lower reaches tributaries and lakeside areas, and existing water body distribution vector data, the corrected image is cropped, retaining only the portion of the image belonging to water bodies within that area. The cropped image is used for subsequent pollutant concentration inversion work.

[0015] The drone flight area was determined by referencing processed satellite hyperspectral imagery to identify water bodies with abnormal reflectivity. These areas are likely to have abnormal pollution concentrations and were designated as key drone flight zones. The flight path also included surrounding tributary dikes and shallow waters, as these areas are prone to agricultural non-point source pollution. The drone flight routes were planned using a grid pattern based on the shape and area of ​​the key flight zones. The spacing between flight routes was designed to ensure a 30% overlap between images captured by adjacent routes, preventing gaps during subsequent image stitching. The flight altitude was set below 300 meters, allowing the drone's multispectral sensor to acquire sub-meter resolution images. To identify small farmland drainage outlets and decentralized aquaculture wastewater discharge points; when choosing flight times, it is necessary to avoid the midday sun and the evening light, as the light intensity varies greatly during these two periods, resulting in inconsistent radiance of images taken at different times. Flights are usually selected between 9:00 and 11:00 AM or 2:00 and 4:00 PM. Before flying, the weather forecast for the next two hours should be checked to ensure that there are no severe weather conditions such as strong winds or heavy rain during the flight period, so as to prevent affecting the safety of drone flight and the quality of image capture; to carry out drone flight operations, the operator controls the drone equipped with a multispectral sensor and flies it along a planned route. During the flight, the drone takes a multispectral image every two seconds and records the geographical location information at each time of capture, including the longitude, latitude, and altitude of the capture point.

[0016] After the flight, the multispectral images and corresponding geographic location information stored by the UAV are transmitted to a computer. The images are then stitched together. Based on the geographic location information, all individual images are arranged according to the order of capture and spatial location. The feature point matching degree between adjacent images is calculated, and feature points with high matching degrees are connected to merge multiple images into a complete multispectral image of the key area. Radiometric consistency correction is then performed on the stitched image. Using the radiance of the first image in the stitched image as a benchmark, the radiometric difference value between each subsequent image and the first image is calculated. The corresponding radiometric difference value is subtracted from the radiance of each subsequent image to ensure that the radiance of the entire stitched image is consistent. The processed image is used for subsequent local pollution source location work.

[0017] The initial deployment locations of ground-based IoT devices were determined based on the watershed distribution of the middle and lower reaches of the Yangtze River and the lakeside area. The first batch of water quality sensors were deployed near the estuaries of major tributaries, regular monitoring sections in the lakeside area, and known agricultural non-point source pollution discharge points. These locations are areas where pollutants are prone to accumulate or where pollution is likely to occur. The distance between each deployment point was set between one and two kilometers to ensure coverage of the main water bodies in the area. The types of water quality parameters to be collected were determined. Referring to the main pollutants from agricultural non-point source pollution in the area, the water quality parameters to be collected included the concentrations of chemical oxygen demand, ammonia nitrogen, total phosphorus, total nitrogen, and chlorophyll a. These parameters can comprehensively reflect the degree of agricultural non-point source pollution in the water bodies. The water quality parameter collection frequency was set. Under normal monitoring conditions, each water quality sensor collected water quality parameter data every thirty minutes. When satellite hyperspectral imagery or UAV multispectral data identified an abnormal pollution area, the collection frequency of all water quality sensors within a five-kilometer radius of the abnormal pollution area was immediately adjusted to collect data every five minutes. This allows for more intensive acquisition of data on water quality changes in the polluted area.

[0018] Real-time data transmission is achieved by equipping each water quality sensor with a built-in wireless communication module. After collecting water quality parameter data, the sensor preprocesses the data to remove obvious outliers. Specifically, the collected water quality parameter data is compared with the average value of that parameter collected by the sensor over the past 24 hours. If the current collected data is greater than three times the average or less than one-third of the average, it is considered an outlier and discarded. The processed normal data is then transmitted to a remote data center via the wireless communication module. Real-time data transmission is crucial; the time from when the sensor collects the data to when the data is received by the data center must be recorded. The interval should not exceed one minute. Based on the location of the pollution source determined by the UAV multispectral data, the deployment of ground-based IoT devices should be adjusted. The number of water quality sensors should be increased within a kilometer radius of the pollution source. Specifically, two sensors should be added in the upwind and downwind directions of the pollution source. The upwind sensor is used to monitor the water quality before the pollution spreads, and the downwind sensor is used to monitor the water quality after the pollution spreads. At the same time, the acquisition frequency of the original sensors around the pollution source should be maintained at once every five minutes until the pollution in the area is brought under control. This will obtain real-time water quality parameter data of the core pollution area and the diffusion path.

[0019] In this embodiment, the acquisition and processing of satellite hyperspectral images can obtain water images covering the entire basin of the middle and lower reaches of the Yangtze River and the lakeside area. Furthermore, through radiometric calibration and atmospheric correction, radiation interference and atmospheric effects are removed, resulting in images that reflect the true water conditions. This solves the problem that satellite remote sensing is greatly affected by the atmosphere when monitoring in this area and is difficult to adapt to the needs of dispersed monitoring of agricultural non-point source pollution.

[0020] In a preferred embodiment of the present invention, pollutant concentration inversion is performed on satellite hyperspectral imagery to obtain pollutant concentration data; local pollution source location is performed on UAV multispectral data to obtain pollution source location data; and real-time water quality parameter data is obtained from ground-based IoT devices, which may include: Satellite hyperspectral imagery is processed to obtain pollutant concentration data. Spatial variability in the pollutant concentration data is analyzed to identify concentration anomaly areas. Specifically, this involves selecting bands in the satellite hyperspectral imagery that are relevant to pollutant characteristics. For agricultural non-point source pollution in the middle and lower reaches of the Yangtze River and lakeside areas, specific wavelength bands reflecting chemical oxygen demand (COD), ammonia nitrogen, total phosphorus, total nitrogen, and chlorophyll a are selected. Specifically, the green band is selected in the 530-590 nm range, the red band in the 630-690 nm range, and the near-infrared band in the 850-910 nm range. These bands are matched with the spectral absorption or reflectance characteristics of different pollutants. Measured concentration data of corresponding pollutants collected by ground-based IoT devices in the area over the past three months are also obtained. This data must cover sampling results under different weather conditions and at different times. To ensure data representativeness, a correlation is established between the surface reflectance of the selected bands and the measured concentration data. Specifically, the correlation coefficient for each band is obtained through correlation analysis between historical measured data and the reflectance of the corresponding band. For example, the correlation coefficient for chemical oxygen demand (COD) in the green band is calculated using linear regression of the measured COD value and the green band reflectance over the past year. The correlation coefficient for ammonia nitrogen in the red band is determined using the same method. The surface reflectance of each band is multiplied by the corresponding correlation coefficient, and the calculation results for all bands are summed. Finally, a preset correction value is added, determined based on the average deviation between the calculated and measured values ​​in historical data, to obtain the pollutant concentration data for each pixel in the satellite hyperspectral image; and the pollutant concentration data for the downstream tributaries and the target monitoring area of ​​the lakeside region.

[0021] Pollutant concentration data are arranged sequentially from west to east and from north to south according to latitude and longitude. The distance between two adjacent spatial locations is determined based on the pixel resolution of the satellite hyperspectral imagery; for example, if the pixel resolution is 30 meters, adjacent locations are selected at 30-meter intervals. The concentration difference between two adjacent spatial locations is calculated by subtracting the concentration of the previous location from the concentration of the latter location. The concentration difference values ​​of all adjacent locations within the entire monitoring area are statistically analyzed, and all difference values ​​are summed and then divided by the total number of adjacent locations to obtain the average concentration difference value. Considering the topographical characteristics of the tributaries and lakeside areas in the middle and lower reaches of the Yangtze River, the concentration difference values ​​within 50 meters of each dike are examined in the shallow polder areas. For the bend areas of tributaries, the concentration difference values ​​of the inner side of the bend are examined, as these areas are prone to the accumulation of agricultural non-point source pollution. When the concentration difference values ​​of three or more consecutive adjacent locations in a certain area are greater than twice the average value calculated above, and the pollutant concentration values ​​of all locations in the area are higher than the average pollutant concentration of the entire monitoring area, the area is marked as a concentration anomaly area.

[0022] UAV multispectral data collection and analysis were conducted on areas with abnormal concentrations to obtain pollution source location data; Based on pollution source location data, the spatial deployment scheme of ground-based IoT devices is optimized to obtain real-time water quality parameter data for the core pollution area and diffusion path. Specifically, this includes: determining the drone flight range based on the extent of the concentration anomaly area; if the anomaly area is irregularly shaped, using a polygon formed by extending 500 meters outward from the boundary of the anomaly area as the flight range to ensure complete coverage of the anomaly area and surrounding areas where pollution sources may exist; simultaneously including polder areas and shallow waters near farmland drainage outlets in the middle and lower reaches of the Yangtze River within the range as key collection areas; planning drone flight routes using a grid pattern, with route spacing calculated based on the field of view and flight altitude of the drone's multispectral sensor. For example, with a sensor field of view of 60 degrees and a flight altitude of 200 meters, the ground coverage width is approximately 230 meters. To ensure a 30% overlap in images captured by adjacent flight paths, the route spacing is set to 160 meters; the flight altitude is controlled below 200 meters to obtain sub-meter resolution multispectral images, meeting the needs for identifying small farmland drainage outlets and aquaculture discharge outlets; and selecting flight times between 9:00 AM and 11:00 AM or 2:00 PM and 4:00 PM for data collection. Stable light intensity is ensured to avoid overexposure due to strong midday sunlight or insufficient brightness due to weak evening light. The weather forecast for the next two hours is checked before flight to ensure no strong winds, heavy rain, or other weather conditions affecting flight safety and image quality. During acquisition, the drone captures multispectral images every two seconds, simultaneously recording the geographic location information (longitude, latitude, and altitude) for each capture via its built-in GPS. After acquisition, the multispectral images are stitched together according to geographic location information to create a complete image of the concentration anomaly area. The spectral reflectance values ​​of normal water bodies in the same wavelength band are first obtained for this area. For example, on a sunny day, the green band reflectance of unpolluted water is about 5% to 8%, and the red band reflectance is about 2% to 4%. By comparing the spectral reflectance values ​​of different locations in the image, we can find locations where the reflectance value exceeds the normal range by more than 10%. Combining the surrounding terrain features of these locations, we can check whether they are close to farmland drainage outlets (observe whether there are traces of farmland drainage) or aquaculture sewage outlets (observe whether there are pipe outlets or signs of sewage discharge). These locations with abnormal reflectance values ​​and signs of sewage discharge are marked as pollution source locations. By integrating the longitude and latitude information of all marked locations, we can obtain pollution source location data.

[0023] The pollution core area is defined by centered on each pollution source location, with a circular area of ​​500 meters in radius. If the core areas of multiple pollution sources overlap, they are merged into a single unified core area. The pollution diffusion path is analyzed, referencing regional hydrological data, such as river flow maps provided by the local water authority. If no readily available data is available, the water flow direction is determined by observing water surface textures captured by drones or by ground observation of water ripples. A 1000-meter area extending from the pollution source location along the water flow direction is defined as the pollution diffusion path, with a path width of 200 meters to cover the potential diffusion range. The optimized deployment scheme involves a 600-meter spacing between ground-based IoT devices in the conventional areas of the Yangtze River's middle and lower reaches tributaries. Within the pollution core area, this spacing is halved to 300 meters to ensure that each device... The coverage area is no more than 200 meters, enabling intensive monitoring of water quality in the core area. Along the pollution diffusion path, one device is deployed every 300 meters in the direction of water flow, and an additional device is installed 100 meters on each side of the centerline of the diffusion path, forming a three-dimensional monitoring network to capture water quality changes during the diffusion process. After adjustment, all devices are activated to collect water quality parameters at a frequency of once every five minutes. The collected parameters include the concentrations of chemical oxygen demand, ammonia nitrogen, total phosphorus, total nitrogen, and chlorophyll a. After each device collects data, it transmits the data in real time to a remote data center via a built-in 4G or 5G wireless communication module. The data center classifies and organizes the received data, storing it separately according to the pollution core area and the diffusion path. After integration, real-time water quality parameter data for the pollution core area and the diffusion path are obtained.

[0024] This embodiment optimizes the deployment scheme of ground-based IoT devices by adjusting the spacing and adding devices according to the core pollution area and diffusion path. This enables more accurate water quality parameter collection in the core pollution area and diffusion path, solving the problem of insufficient data in the core pollution area caused by static deployment of ground-based devices.

[0025] like Figure 2 As shown, in another preferred embodiment of the present invention, the pollutant concentration data, pollution source location data, and real-time water quality parameter data are spatially discretized to establish a quadtree index and recursively divide them into sub-regions; based on the statistical characteristics of data points in each sub-region, dynamic spatial correction parameters are generated, which may include: Pollutant concentration data, pollution source location data, and real-time water quality parameter data are projected onto the same spatial coordinate system to obtain a spatially registered multi-source dataset. Specifically, this involves selecting the WGS84 coordinate system, suitable for the middle and lower reaches of the Yangtze River, which accurately reflects the geographic coordinate relationships of the region. Processing the pollutant concentration data involves extracting the latitude and longitude coordinates corresponding to each pixel. If the original coordinates are image row and column numbers, they are converted using parameters from the satellite sensor. For example, the pixel resolution of the Gaofen-5 satellite is 30 meters, and the sensor data during image capture... The starting latitude and longitude of the top left corner is 116 degrees east longitude and 30 degrees north latitude. During the conversion, the row number of the image is multiplied by 30 meters to obtain the offset relative to the starting latitude. This offset is converted into a latitude difference (1 meter is approximately equal to 0.000009 degrees). Adding this to the starting latitude gives the actual latitude of the pixel. Similarly, the column number of the image is multiplied by 30 meters to obtain the offset relative to the starting longitude. This offset is converted into a longitude difference (1 meter is approximately equal to 0.000011 degrees). Adding this to the starting longitude gives the actual longitude of the pixel, thus matching the pollutant concentration data with the WGS84 coordinate system.

[0026] For processing pollution source location data, if the original coordinates are in the Beijing 54 coordinate system, they are converted to latitude and longitude in the WGS84 coordinate system using a coordinate transformation formula. Specifically, the x-coordinate of the original coordinates is multiplied by a transformation factor of 1.00001, and then an x-coordinate offset of -5 meters is added to obtain the x-coordinate in the WGS84 coordinate system; the y-coordinate of the original coordinates is multiplied by a transformation factor of 0.99998, and then an y-coordinate offset of 3 meters is added to obtain the y-coordinate in the WGS84 coordinate system. The x and y coordinates are then converted to latitude and longitude. For processing real-time water quality parameter data, extraction is performed on each ground-based IoT device. If the original latitude and longitude of the installation location is in another coordinate system, it should be converted to latitude and longitude in the WGS84 coordinate system using the same conversion method as the pollution source location data. The three types of data after conversion should be integrated according to the latitude and longitude coordinates. For example, under a certain latitude and longitude coordinate (116.0002 degrees east longitude and 30.0003 degrees north latitude), the pollutant concentration value, whether it is a pollution source location (marked as 1 if yes, 0 otherwise), and real-time water quality parameter value should be integrated to ensure that the same latitude and longitude location contains complete three types of data information, forming a spatially registered multi-source dataset.

[0027] Spatially discretize the multi-source dataset for spatial registration to obtain a discretized data point set. Based on the spatial distribution density of the discretized data point set, a quadtree spatial index is established to obtain the initial spatial partitioning. Specifically, this includes: calculating the average distribution density of data points within the coverage area of ​​the spatially registered multi-source dataset. For example, if the average distribution density of data points in this area is one data point per 1000 square meters, to ensure that each grid contains at least one data point, the grid side length is set to 30 meters (the area of ​​a single grid is 900 square meters); dividing the area covered by the dataset into 30-meter sides... The length is divided into equally sized square grids. Each grid is identified by the latitude and longitude coordinates of its upper left and lower right corners. For example, the upper left corner of a grid might be 116.0000°E, 30.0003°N, and the lower right corner might also be 116.0003°E, 30.0000°N. Based on the latitude and longitude range of the grid, it is determined whether the latitude and longitude of each data point in the spatially registered multi-source dataset falls within that grid. For example, if a data point has a longitude of 116.0002° and a latitude of 30.0002°, it falls within the aforementioned grid. All such data points within that grid are then extracted. The data points include pollutant concentration values, pollution source location identifiers, and real-time water quality parameter values. The average value of each parameter within the grid is calculated by summing the values ​​of the same parameter across all data points within the grid and then dividing by the number of data points in that grid. For example, if there are 5 data points in a grid with COD values ​​of 2 mg / L, 3 mg / L, 4 mg / L, 3 mg / L, and 2 mg / L, the sum is 14 mg / L, and dividing by 5 gives an average COD value of 2.8 mg / L. The average values ​​of other parameters are calculated using the same method. The center point of each grid is used as the reference point. To discretize the location of the data points, the latitude and longitude of the grid center point are obtained by calculating the average of the latitude and longitude of the top left and bottom right corners of the grid. For example, the longitude of the center point of the above grid is (116.0000 + 116.0003) ÷ 2 = 116.00015 degrees, and the latitude is (30.0003 + 30.0000) ÷ 2 = 30.00015 degrees. By mapping the average values ​​of each parameter of the grid to the location of the center point, a discretized data point is formed. The discretized data points corresponding to all grids are integrated and arranged in latitude and longitude order to obtain the discretized data point set.

[0028] Find the maximum and minimum longitude and latitude values ​​among all data points. Use the minimum longitude and maximum latitude as the top-left corner coordinates of the region, and the maximum longitude and minimum latitude as the bottom-right corner coordinates. Define a rectangular boundary as the root node region of the quadtree. Calculate the area of ​​the root node region. For example, if the top-left corner coordinates are 116°E, 30°N, and the bottom-right corner coordinates are 117°E, 29°N, a 1-degree difference in longitude corresponds to approximately 111 kilometers on the ground, and a 1-degree difference in latitude corresponds to approximately 111 kilometers on the ground. The area of ​​the root node region is 111 kilometers multiplied by 111 kilometers, resulting in 12321 square kilometers. Calculate the root node area. The spatial distribution density of the root node region is calculated by dividing the number of discretized data points within the region by the area of ​​the root node region. For example, if there are 600 data points, the density is approximately 0.0487 points per square kilometer (600 points divided by 12321 square kilometers). The density threshold is determined based on the typical data density of the Yangtze River's middle and lower reaches tributaries and lakeside areas. This region includes different sub-regions such as polder areas, shallow waters, and open lake areas, resulting in varying typical data densities, generally ranging from 0.3 to 0.7 points per square kilometer. The data points are denser in polder areas and concentrated areas of farmland drainage outlets, with a typical density of 0.6 to 0.7 points per square kilometer. Data points are sparser in open lake areas, with a typical density of 0.3 to 0.4 points per square kilometer. Therefore, the density threshold needs to be selected based on the main terrain type of the root node region. If the root node region is mainly composed of polder areas, the threshold should be set to 0.6 to 0.7 points per square kilometer; if it is mainly composed of open lake areas, the threshold should be set to 0.3 to 0.4 points per square kilometer; if it includes multiple terrain types, the median value of 0.4 to 0.5 points per square kilometer should be used. Comparing the spatial distribution density of the root node region with the threshold, if the density of the root node region is less than the threshold (e.g., the root node is an open lake area, and the threshold is set to 0.3 points per square kilometer), the threshold will be lower. If the calculated density is 0.2 points per square kilometer, no division is performed, and the root node region becomes the initial spatial partition. If the density of the root node region is greater than the threshold, for example, if the root node is a polder area, the threshold is set to 0.6 points per square kilometer, and the number of data points is 7500, the density is approximately 0.609 points per square kilometer (7500 divided by 12321 square kilometers), which is greater than the threshold, then the root node region is divided into two equal parts along the horizontal direction (longitude) and the vertical direction (latitude). The horizontal direction is divided according to the average longitude, and the vertical direction is divided according to the average latitude, forming four sub-regions of equal size. Each sub-region becomes a child node of the root node.

[0029] Calculate the number of discretized data points in each sub-region. Divide the number of data points in each sub-region by the area of ​​that sub-region (the area of ​​the sub-region is one-quarter of the area of ​​the root node region, i.e., 3080.25 square kilometers) to obtain the spatial distribution density of each sub-region. If the spatial distribution density of a certain sub-region is still greater than the threshold, for example, if a certain sub-region has 2000 data points and the density is 2000 divided by 3080.25 square kilometers, approximately 0.649 data points per square kilometer, which is greater than the threshold of 0.6 data points per square kilometer, then continue to divide the sub-region in the same way until the spatial distribution density of all sub-regions is less than or equal to the threshold. Integrate all the finally obtained sub-regions and number them in order from west to east and from north to south to form the initial spatial partition. At the same time, record the boundary coordinates of each sub-region and the number of discretized data points in the region. Establish a quadtree spatial index, in which each node corresponds to a sub-region and contains sub-region information and data point association information.

[0030] The initial spatial partition is recursively divided to obtain a set of sub-regions. Statistical characteristics of data points within each sub-region set are calculated to obtain dynamic spatial correction parameters. Specifically, based on the monitoring accuracy requirements of the Yangtze River's middle and lower reaches tributaries and lakeside areas, the preset number of data points is set to 20, and the preset area is set to 100 square kilometers. When the number of discretized data points in a given initial spatial partition exceeds 20, or the area of ​​the partition is greater than 100 square kilometers, the partition is recursively divided. The division method involves dividing the partition into two equal parts both horizontally and vertically, with the horizontal division... The midpoint of the longitude range of the region is used as the dividing line, and the midpoint of the latitude range of the region is used as the dividing line in the vertical direction to form four new sub-regions. Each new sub-region is checked to see if it meets the recursive division conditions. For example, if a new sub-region has 25 data points and an area of ​​120 square kilometers, the condition is met and the division continues. If a new sub-region has 15 data points and an area of ​​80 square kilometers, the condition is not met and the division stops. This process is repeated until all sub-regions no longer meet the recursive division conditions. All the final sub-regions are then organized according to the division level to form a set of sub-regions.

[0031] Calculating the statistical characteristics of data points within each sub-region requires first calculating the mean pollutant concentration. This is done by summing the pollutant concentration values ​​of all data points within the sub-region and then dividing by the number of data points in that sub-region. For example, if there are 5 data points in a sub-region with concentration values ​​of 2 mg / L, 3 mg / L, 4 mg / L, 3 mg / L, and 2 mg / L, the sum is 14 mg / L, which is then divided by 5 to obtain the mean concentration of 3 mg / L. Next, the concentration variance is calculated by subtracting the mean concentration from the pollutant concentration value of each data point within the sub-region. This gives the concentration deviation for each data point. For example, if the deviations for the above data points are -1 mg / L, 0 mg / L, 1 mg / L, 0 mg / L, and -1 mg / L, the squares of each deviation are summed to obtain 1 + 0 + 1 + 0 + 1 = 3 (mg / L), which is then divided by the number of data points in the sub-region. With a quantity of 5, a concentration variance of 0.6 (mg / L) was obtained. Simultaneously, the mean and variance of ammonia nitrogen, total phosphorus, total nitrogen, and chlorophyll a concentrations in the real-time water quality parameters were calculated, using the same methods as for the pollutant concentration mean and variance. For example, when calculating the ammonia nitrogen mean, the ammonia nitrogen values ​​of all data points within the sub-region were summed and divided by the number of data points. When calculating the ammonia nitrogen variance, the deviation was obtained by subtracting the ammonia nitrogen mean from each ammonia nitrogen value, and the sum of squared deviations was divided by the number of data points. The pollutant concentration mean and variance, as well as the mean and variance of ammonia nitrogen, total phosphorus, total nitrogen, and chlorophyll a for each sub-region, were integrated and sorted by parameter type to form the dynamic spatial correction parameters corresponding to that sub-region. The dynamic spatial correction parameters of all sub-regions were arranged sequentially according to the sub-region number, collectively constituting the dynamic spatial correction parameters for the entire monitoring area of ​​the middle and lower reaches of the Yangtze River and the lakeside area.

[0032] In this embodiment, the recursive partitioning and statistical characteristic calculation of the initial partitions, by setting clear partitioning conditions and calculating the mean and variance through specific data examples, can obtain targeted dynamic spatial correction parameters for each sub-region, solving the problem of large data deviations in different regions and improving the accuracy of overall data processing.

[0033] In a preferred embodiment of the present invention, the pollutant diffusion numerical model constructed by fusing multi-source data through data assimilation technology is optimized according to dynamic spatial correction parameters to obtain an optimized pollutant diffusion model, which may include: The dynamic spatial correction parameters are fused with the initial parameters of the pollutant diffusion numerical model to generate parameter optimization criteria. Specifically, this includes: clarifying the content of the dynamic spatial correction parameters, which include the mean and variance of pollutant concentrations in the tributaries of the middle and lower reaches of the Yangtze River and the open lake areas surrounding the drainage outlets of farmland in various sub-regions of the lakeside area, as well as the mean and variance of ammonia nitrogen, total phosphorus, total nitrogen, and chlorophyll a; determining the initial parameters of the pollutant diffusion numerical model, specifically for agricultural non-point source pollution, farmland drainage nitrogen and phosphorus pollution, and aquaculture discharge in this region. The initial parameters include the water flow diffusion coefficient (set to 0.5 m² / s based on regional hydrological data), the pollutant degradation rate (set to 0.1 m / s based on 20-25℃ climatic conditions), and the water flow velocity (set to 0.3 m / s based on measured values ​​of the tributaries); and generating parameters according to their types. The criterion, taking the water flow diffusion coefficient as an example, is calculated by multiplying the concentration variance in the dynamic correction parameters by the first weight, and then adding the initial diffusion coefficient multiplied by the second weight. The first weight is set according to the reliability of real-time water quality parameters, taking 0.55 to 0.65 when there are no abnormalities after 5 consecutive minutes of sampling, and 0.45 to 0.55 when there are a few missing parameters. The second weight is 1 minus the first weight. For example, if the concentration variance is 0.8 mg / L squared, the first weight is 0.6, and the initial diffusion coefficient is 0.5 square meters per second, then the baseline value of the diffusion coefficient is 0.8 mg / L squared multiplied by 0.6 plus 0.5 square meters per second multiplied by 0.4, resulting in 0.68 square meters per second. Similarly, the degradation rate and water flow velocity are processed to finally form a parameter optimization criterion containing the weight rules of the baseline values ​​of each parameter.

[0034] Based on parameter optimization criteria, pollutant concentration data, pollution source location data, and real-time water quality parameter data are processed to obtain a multi-source dataset fused using data assimilation techniques. Specifically, this includes: organizing three types of data: pollutant concentration data (pixel-by-pixel concentration values ​​obtained from processed satellite imagery); pollution source location data (latitude and longitude coordinates of farmland drainage outlets and aquaculture discharge outlets); and real-time water quality parameter data (COD, ammonia nitrogen, and other parameter values ​​collected every 5 minutes after optimized deployment); and assigning data weights according to optimization criteria, with real-time water quality parameters weighted at 0.55 to 0.65 and pollutant concentration data weighted at 0.25 to 0. .35 The outlier removal weight for pollution source location data was set between 0.05 and 0.15. Data assimilation processing was adopted. First, the concentration prediction value was calculated using the initial parameters of the model. The predicted value was compared with the real-time measured value to calculate the difference. The difference was multiplied by a weight of 0.6 to obtain the correction amount. Then, the pollutant concentration data was multiplied by a weight of 0.3 and the correction amount was added. At the same time, outlier data with concentrations exceeding 3 times the average value beyond 5 kilometers of the pollution source location were removed. The process of calculating the predicted value, the difference, and the correction amount was repeated until the deviation was less than 0.1 mg / L for 3 consecutive times. Finally, the three types of data were integrated to obtain a fused multi-source dataset.

[0035] The multi-source dataset is iteratively matched with the pollutant diffusion numerical model to obtain the adjusted model parameter combination. Specifically, this involves: inputting the concentration and flow velocity values ​​from the fused data into the model as observation data; the model outputs simulated concentration values ​​using the initial parameters; calculating the average error between the simulated and observed values ​​by subtracting the observed value from the simulated value at each spatial location, summing all differences, and dividing by the total number of locations; adjusting the parameters according to the error: if the error is positive and the simulated value is high, the degradation rate is increased by multiplying the error by 0.18 to 0.22 of the regional pollution degradation law coefficient, and the new degradation rate is the original rate plus the error multiplied by 0.18 to 0.22; if the error is negative and the simulated value is low, the diffusion coefficient is increased by multiplying the absolute value of the error by 0.18 to 0.22, and the new diffusion coefficient is the original coefficient plus the absolute value of the error multiplied by 0.18 to 0.22; substituting the adjusted parameters into the model, the simulated values ​​and errors are recalculated, and the adjustment process is repeated until the average error is less than the regional monitoring accuracy threshold of 0.05 mg / L. The diffusion coefficient, degradation rate, and flow velocity at this point are recorded to form the adjusted parameter combination.

[0036] The adjusted model parameters are used to replace the original parameters in the pollutant diffusion numerical model to obtain an optimized pollutant diffusion model. Specifically, this includes: streamlining the model parameter structure, which includes calculations related to water flow motion to simulate the flow trajectory of water bodies in the middle and lower reaches of the Yangtze River and lakeside areas based on water flow velocity, such as calculating the flow direction of a tributary section and the time it takes for the water to flow through each grid; calculations related to pollutant migration to simulate the spatial diffusion range of pollutants with water flow based on the diffusion coefficient, such as calculating the number of grids through which nitrogen and phosphorus pollutants discharged from farmland drainage outlets diffuse within one hour; and calculations related to pollutant transformation to calculate the amount of pollutant decomposition in the natural environment according to the degradation rate, such as calculating the daily reduction in total phosphorus concentration in a certain area due to degradation. Each parameter has an independent storage location and a clearly defined calculation step. The adjusted water flow velocity is then replaced with the parameters for each calculation part. The initial input values ​​for the water flow motion calculation section are used to ensure that the simulated water flow trajectory more closely matches the actual tributary flow velocity. The adjusted diffusion coefficient is then used to replace the initial input values ​​for the pollutant migration calculation section, making the pollutant diffusion range calculation more accurate. The adjusted degradation rate is then used to replace the initial input values ​​for the pollutant transformation calculation section, ensuring that the pollutant decomposition amount calculation conforms to the regional pollution degradation pattern. After the replacement, the storage location and corresponding calculation steps of each parameter are checked one by one to confirm that there are no missing or mismatched parameters. Then, the model is validated by randomly selecting 30% of the data from the fused data as validation data and inputting this data into the model after parameter replacement. The model will output the simulated concentration values ​​at each spatial location and time point. These simulated values ​​are compared with the measured concentration values ​​in the validation data, and the average error between the two is calculated. If the average error is still less than 0.05 mg / L, it indicates that the parameter replacement is effective, and the optimized pollutant diffusion model is obtained.

[0037] This embodiment, through parameter replacement and model verification, can ensure that the parameters of the pollutant diffusion model are accurately matched with the actual pollution situation in the tributaries of the middle and lower reaches of the Yangtze River and the lakeside area, thus solving the problem of large deviations between simulation results and reality caused by fixed model parameters.

[0038] In a preferred embodiment of the present invention, the spatiotemporal sequence data output by the optimized pollutant diffusion model is input into a spatiotemporal convolutional neural network to extract the spatiotemporal co-evolution features of pollutants and obtain the identification result, which may include: The spatiotemporal sequence data output by the optimized pollutant diffusion model is processed to obtain the spatiotemporal distribution characteristics of pollutants. The changes in pollutant concentration gradients within these characteristics are analyzed to obtain pollutant propagation direction data. Specifically, this includes: acquiring the spatiotemporal sequence data output by the model, arranged in the order of time, longitude, latitude, and concentration, containing continuous 72-hour monitoring data, recorded hourly, with the coverage area divided into 100m × 100m grids, and the latitude and longitude of each grid center point accurate to six decimal places. The data includes the pollutant concentration value corresponding to each time point and each grid center point; for example, the COD concentration value of the grid at 116.0002°E, 30.0003°N in the 5th hour is 3.2 mg / L. The data is then grouped by time dimension, grouping all grid concentration values ​​at the same time point together (e.g., grouping all grid concentration values ​​in the 12th hour), calculating the maximum, minimum, and average values ​​for each group. To calculate the maximum value, the largest value is found among all concentration values ​​in the group; to calculate the minimum value, the smallest value is found; and to calculate the average value, the largest value is found among all concentration values ​​in the group. The concentration values ​​of each group are summed and then divided by the total number of grids at that time point. For example, the maximum concentration at the 12th hour is 8 mg / L, the minimum is 1 mg / L, and the average is 3.5 mg / L. Grouping is done by spatial dimension, grouping the concentration values ​​of the same grid at all time points into one group. For example, the concentration values ​​of the grid at 116.0005 degrees east longitude and 30.0006 degrees north latitude over 72 hours are grouped together, and the trend of this group of data is analyzed. By comparing the concentration values ​​at adjacent time points, it is determined whether the concentration is rising, falling, or stable. For example, the concentration in a certain polder grid rises from 2 mg / L in 1 hour to 5 mg / L in 6 hours, and then falls to 3 mg / L in 12 hours, clearly showing the change pattern of the grid concentration over time. The statistical values ​​of the time dimension and the trend of the spatial dimension are integrated, and the maximum, minimum, and average concentration values ​​at each time point are correlated with the concentration change trends of each grid. For example, combining the overall concentration statistics at the 12th hour with the concentration changes in each polder and tributary grid, a spatiotemporal distribution characteristic reflecting the overall distribution and spatial location changes of pollutants at different times is formed.

[0039] Using a 100m × 100m grid as the calculation unit, the concentration value of each grid is the average of the concentration values ​​of all data points within that grid. For example, if a grid has 5 data points with concentration values ​​of 2.1mg / L, 2.3mg / L, 2.2mg / L, 2.4mg / L, and 2.0mg / L, the average concentration of that grid is (2.1 + 2.3 + 2.2 + 2.4 + 2.0) / 5 = 2.2mg / L. The concentration difference between each grid and its four adjacent grids in the east, south, west, and north directions is calculated. The eastward adjacent grid refers to the 100m × 100m grid whose eastern edge is completely adjacent to the current grid boundary. The same applies to other directions. The concentration difference is calculated by subtracting the average concentration of adjacent grids from the average concentration of the current grid. For example, if the current grid concentration is 2.5mg / L and the eastward adjacent grid concentration is 2.0mg / L, the eastward concentration difference is 0.5mg / L. The propagation direction is then determined. If the eastward concentration difference is positive, it indicates that the current... If the concentration in the current grid is higher than that of the adjacent grid to the east, the pollutant may spread westward from the current grid to the adjacent grid to the east. If the concentration difference to the east is negative, it means that the concentration in the current grid is lower than that of the adjacent grid to the east, and the pollutant may spread eastward from the adjacent grid to the current grid. The same method is used to determine the possible propagation directions in the south, west, and north. Compare the absolute values ​​of the concentration differences in the four directions. The larger the absolute value, the more obvious the concentration gradient in that direction, and the higher the probability that the pollutant will spread in that direction. The direction with the largest absolute value is determined as the main propagation direction of that grid. For example, if the absolute value of the northward concentration difference in a certain grid is 2.5 mg / L and the northward concentration difference is positive, it means that the concentration in the current grid is higher than that of the adjacent grid to the north, and the main propagation direction is southward. Integrate the main propagation directions of all grids and record the information of each grid in the format of longitude, latitude, and propagation direction. For example, the propagation direction of the grid at 116.0008 degrees east longitude and 30.0009 degrees north latitude is east, forming pollutant propagation direction data covering the entire monitoring area.

[0040] The pollutant propagation direction data is processed to obtain the main direction set of pollutant diffusion. Based on this main direction set, a ray tracing algorithm is executed to obtain the ray cluster distribution of pollutant diffusion. Specifically, the monitoring area is divided into three sub-regions according to topography: polder areas, tributary channels, and lakeside areas. The polder areas refer to low-lying areas surrounded by dikes and used for agricultural planting, typically distributed around tributaries, with an area between 0.5 and 5 square kilometers. Tributary channels refer to tributaries of the main stream of the middle and lower reaches of the Yangtze River, with a width between 10 and 50 meters and a relatively fixed water flow direction. The lakeside areas... This refers to areas near the lake shore with a water depth of less than 5 meters, which are significantly affected by changes in lake water levels. Each sub-region contains multiple consecutive 100m x 100m grids, and the propagation direction data of each grid belongs to the corresponding sub-region. The frequency of each propagation direction within each sub-region is counted. All grids within the sub-region are traversed, and the propagation direction of each grid is recorded. Each time a certain direction appears, its frequency is incremented by 1. For example, in a polder area with 100 grids, there are 35 grids with propagation directions heading east, 25 heading south, 20 heading west, and 20 heading north.

[0041] The percentage of each direction is calculated by dividing the frequency of each direction by the total number of grids in the sub-region. For example, the percentage of the eastward direction in the polder area is 35 divided by 100, resulting in 35%. The percentages of other directions are calculated in the same way. The direction with the largest percentage is selected as the main direction of the sub-region. If two directions have the same percentage and are both the largest, for example, the percentages of the eastward and southward directions in a tributary are both 30%, and are higher than other directions, then these two directions are listed as the main directions of the sub-region. The main directions of the three sub-regions are integrated. In the polder area, agricultural non-point source pollution spreads mainly along ditches, so the main direction is eastward. In the tributary, the water flow direction is stable, so the main direction is southward. In the lakeside area, influenced by the lake flow, the main direction is southeastward. These sub-regions and their corresponding main directions are organized to form a set of main directions of pollutant diffusion.

[0042] Using the center point of the pollution source as the emission source, the emission source for farmland drainage outlets is the center point of the drainage outlet pipe outlet, with the latitude and longitude consistent with the pipe outlet location; the emission source for livestock sewage outlets is the center point of the sewage outlet discharge point, ensuring that the rays start tracking from the origin of pollution; the polder area has complex terrain with many farmland ditches and embankments, so the ray step length is set to 50 meters to accurately capture the influence of terrain on the rays; the tributary river channel has regular terrain and stable water flow direction, so the ray step length is set to 100 meters to improve tracking efficiency; the lakeside area is open and unobstructed, so the ray step length is set to 200 meters to cover a wider range; the emission angle is extended by 15 degrees according to the main direction of each sub-region. For example, if the main direction of the polder area is east, the emission angle range is 15 degrees east of north to 15 degrees east of south. Within this range, one ray is set every 5 degrees, and a total of 7 rays are emitted for each pollution source to ensure that the rays can cover the range where pollutants may spread.

[0043] Rays are emitted from the emission source according to a set step length and angle. Each step forward checks for terrain obstacles, such as dikes in polder areas, at the current location. The latitude and longitude of the ray's current location are compared with those of the dike to determine if they intersect. If they intersect, the intersection point is used as the new starting point, and the ray angle is adjusted. The adjusted angle should not deviate from the original main direction by more than 15 degrees to avoid the ray passing through unnecessary terrain. If no obstacles are encountered, the ray continues to extend according to the original step length and angle until it reaches the boundary of the corresponding sub-region, or until the pollutant concentration at the ray's location is below 0.5 mg / L (the regional background concentration threshold). Information on all rays is recorded, including the starting latitude and longitude of each ray, the latitude and longitude of its current location after each step, the location of the adjusted angle, and the adjusted angle. All rays corresponding to each pollution source form a ray cluster. Integrating the ray clusters of all pollution sources creates a ray cluster distribution for pollutant diffusion.

[0044] Spatially matching the ray cluster distribution with real-time water quality parameter data yields a corrected ray distribution result. The corrected ray distribution result is then reconstructed using a directional field to obtain the directional field data of the pollutant diffusion path. Specifically, this includes acquiring the latitude and longitude (accurate to 6 decimal places, with sensors installed 0.5 meters below the water surface) of each sensor from an optimized network of ground-based IoT devices, along with measured concentration values ​​collected every 5 minutes. For example, the sensor at 116.0010°E, 30.0012°N collected an ammonia nitrogen concentration of 1.2 mg / L at 8 hours and 5 minutes. The ray distribution within the ray cluster is then processed. Each ray is divided into several line segments according to a set step size. The two endpoints of each line segment are the matching points. For example, for a ray with a step size of 50 meters, an endpoint is taken every 50 meters from the starting point. Each endpoint has a corresponding latitude and longitude and a predicted concentration value. The straight-line distance between each matching point and the installation positions of all sensors is calculated. The distance calculation uses the Earth's surface distance formula and converts the latitude and longitude difference into meter-level distance. If the straight-line distance between a matching point and a sensor is less than 50 meters (the sensor's monitoring range threshold), the matching point is considered to have successfully matched with the sensor. The measured concentration value of the sensor is then correlated with the predicted concentration value of the matching point.

[0045] If the measured concentration at the matching point is lower than the predicted concentration at that point, it indicates that the ray extends beyond the actual pollution range. The ray is then shortened to the matching point and no longer extended. If the measured concentration is higher than the predicted concentration, it indicates that there may be undetected pollution diffusion directions around that point. Two new rays are added at the matching point, one in the northwest (45 degrees west of north) and the other in the northeast (45 degrees east of north), with a step size of 50 meters each, and the tracking continues. If the deviation between the measured and predicted concentrations is within ±0.5 mg / L, it indicates that the prediction result is basically consistent with the actual situation, and the original state of the rays remains unchanged. All corrected rays are integrated, and the starting point, corrected ending point, adjusted position, and information of newly added rays for each ray are recorded to form the corrected ray distribution result.

[0046] Using a 100m x 100m grid as the unit of the direction field, each grid corresponds to a direction vector, which includes two attributes: propagation direction and vector length. The propagation direction of rays within each grid is extracted. Each corrected ray is traversed, and each grid through which the ray passes is determined. The propagation direction is determined based on the ray's forward direction within that grid; for example, if a ray enters from the west side of the grid and exits from the east side, the propagation direction within that grid is eastward. If the ray's angle is adjusted within the grid, the adjusted direction is used as the propagation direction for that grid. The frequency of each propagation direction within each grid is counted, and the direction with the highest frequency is determined as the dominant direction for that grid. The vector length is related to the number of rays passing through the grid. If one ray passes through the grid, the vector length is set to 1; if two rays pass through, the vector length is set to 2, and so on. If no ray passes through the grid, the propagation direction is determined by the direction of the ray passing through the adjacent grid. Interpolate the vectors; when interpolating non-radial grids, select four adjacent radial grids (east, south, west, and north) of the grid, obtain the direction and vector length of each adjacent grid, and calculate the average angle of the adjacent grid directions (east corresponds to 90 degrees, south to 180 degrees, west to 270 degrees, and north to 0 degrees). For example, if the directions of the four adjacent grids are east, south, west, and north, the average angle is (90+180+270+0) / 4=135 degrees, corresponding to the northeast direction; the vector length is taken as the average of the vector lengths of the four adjacent grids. If the length of the adjacent grids is 1, the vector length of the grid is set to 1; organize the main direction and vector length of each grid and record them in the format of longitude, latitude, direction, and vector length. For example, the direction of the grid at 116.0015 degrees east longitude and 30.0018 degrees north latitude is southeast, and the vector length is 2, forming pollutant diffusion path direction field data covering the entire monitoring area.

[0047] The directional field data and spatiotemporal sequence data are fused to obtain fused spatiotemporal feature data. This fused spatiotemporal feature data is then input into a spatiotemporal convolutional neural network for deep feature extraction to obtain the spatiotemporal co-evolutionary features of pollutants. Specifically, this involves: determining the network's purpose—to extract the spatiotemporal co-evolutionary features of pollutants in the middle and lower reaches of the Yangtze River and its lakeside areas; the input data being directional field data and spatiotemporal sequence data; and the output data being co-evolutionary features reflecting the common temporal and spatial changes of pollutants, such as the correlation between pollutant concentration changes over time and diffusion direction changes; and designing the network structure, which includes an input layer, a spatiotemporal convolutional layer, a pooling layer, and a fully connected layer. The input layer converts data from 72 time steps, 100×100 spatial grids, and two features (concentration and direction vectors) into a tensor format that the network can process. To ensure the data dimensions meet the requirements of subsequent calculations, the spatiotemporal convolutional layer uses 3D convolutional kernels with a size of 3×3×3 (3 for time dimension, 3 for spatial height, and 3 for spatial width), with 32 kernels and a stride of 1. It extracts both the temporal and spatial distribution features of the data through convolution operations, and uses the ReLU activation function after convolution. The pooling layer uses 3D max-pooling kernels with a size of 2×2×2 and a stride of 2. It downsamples the convolutional feature map to reduce data dimensionality and computational cost while retaining key features. The fully connected layer consists of two layers. The first layer has 1024 neurons, converting the pooled two-dimensional feature map into a one-dimensional feature vector for initial feature integration. The second layer has 256 neurons, further filtering and optimizing the initially integrated features to output the final 256-dimensional co-evolutionary features.

[0048] Historical data from the past three years was collected for this region. Monitoring data from March to October each year (the peak period for agricultural non-point source pollution) was selected. Ten days were selected each month, and four sets of 72-hour continuous concentration, direction, and vector samples were collected daily (with a start time every 6 hours), totaling 3×8×10×4=960 samples. Forty additional samples were collected from sudden pollution events, such as concentrated discharge of farmland runoff and excessive pollution from livestock farming, bringing the total to 1000 samples. These samples were then divided into a training set (700 samples) and a validation set (300 samples) in a 7:3 ratio to ensure that both sets included data from different pollution types and seasons. The Adam optimizer was used with a learning rate of 0.001. The optimizer was used to adjust network parameters (such as convolutional kernel weights, full...). The loss function is minimized (connection layer neuron weights). The loss function is the mean squared error loss function. The specific calculation process is as follows: obtain the 256-dimensional predicted co-evolutionary feature vector output by the network, and then obtain the 256-dimensional true co-evolutionary feature vector determined based on historical authoritative data. Subtract the corresponding elements of the two vectors respectively (subtract the true feature element value from the predicted feature element value) to obtain 256 element differences. Square each element difference (to avoid the cancellation of positive and negative differences). Add the 256 squared values ​​to obtain the total squared error. Then divide the total squared error by the feature vector dimension of 256 to obtain the mean squared error of a single sample, which is the loss value of that sample. During training, the average loss value of all training samples is used as the optimization target.

[0049] The training run consists of 50 rounds. After each round, the validation loss is calculated using the validation set (calculated in the same way as the training loss). If the validation loss stops decreasing or shows an upward trend after 5 consecutive rounds, an early stopping strategy is adopted to stop training and avoid overfitting. During training, the network parameters are saved every 5 rounds, and the training loss and validation loss for each round are recorded to facilitate subsequent tracking of optimal parameters. To validate the network performance, 200 additional samples that were not used in training are selected from historical data as a test set (including samples of regular pollution and sudden pollution, in the same proportion as the training set). The test set is input into the trained network to obtain the co-evolutionary features output by the network. The similarity is calculated with the real features determined based on authoritative historical data, i.e., using cosine similarity, which is calculated by dividing the dot product of the two feature vectors by the product of the magnitudes of the two vectors. If the similarity exceeds 0.85, the network training is considered successful and can effectively extract the spatiotemporal co-evolutionary features of pollutants. If the similarity is below 0.85, the network structure is adjusted or 50 new sudden pollution samples are added, and the network is retrained until the requirements are met.

[0050] The two types of data are aligned according to the dimensions of time, space, and features. In the time dimension, each time step (72 in total) corresponds to the same monitoring time; for example, the 5th hour time step includes both orientation field data and spatiotemporal sequence data for that time. In the spatial dimension, each 100×100 spatial grid corresponds to the same latitude and longitude range, ensuring a one-to-one correspondence between the orientation vector and concentration value for the same grid. In the feature dimension, the concentration value and orientation vector of each time step and each grid are combined into a feature vector. For example, the feature vector for the 5th hour grid at 116.0020°E, 30.0022°N is a concentration of 4 mg / L, orientation southeast, and vector length of 2. The feature vectors for all time steps and all grids are then combined. The vectors are arranged in order to form a fused spatiotemporal feature data tensor with dimensions of 72×100×100×2. Standardization is used to eliminate the influence of different feature magnitudes. The standardization calculation method is to subtract the minimum value of each feature value and then divide by the difference between the maximum and minimum values ​​of the feature. For example, the minimum value of the concentration value is 0.1 mg / L and the maximum value is 10 mg / L. After standardization, the concentration value of 4 mg / L in a certain grid is (4-0.1)÷(10-0.1)≈0.394. The minimum value of the direction vector length is 1 and the maximum value is 5. After standardization, the vector length of 2 in a certain grid is (2-1) / (5-1)=0.25, ensuring that all feature values ​​are within the range of 0-1.

[0051] Loading a properly trained spatiotemporal convolutional neural network involves inputting standardized, fused spatiotemporal feature data into the network. The data is first converted into a network-compatible format through the input layer, and then enters the spatiotemporal convolutional layer. The 3D convolutional kernel extracts the temporal concentration change trend and spatial diffusion direction distribution features, and the ReLU activation function enhances the nonlinear expression of the features. Next, it enters the pooling layer, where max pooling filters out key features and reduces data redundancy. Finally, it enters the fully connected layer. The first layer of 1024 neurons integrates the pooled features, and the second layer of 256 neurons outputs the final feature vector. This 256-dimensional feature vector is the spatiotemporal co-evolution feature that reflects the co-evolution law of pollutants in time and space.

[0052] By integrating the spatiotemporal co-evolution characteristics of pollutants with the directional field data of pollutant diffusion paths, pollutant identification results are obtained. Specifically, from 256 elements of spatiotemporal co-evolution characteristics, elements related to pollution type and evolution stage are selected, such as concentration change rate (concentration difference between adjacent time steps divided by time interval) and coverage growth rate (area difference of polluted area between adjacent time steps divided by time interval). These elements can reflect the change law of pollutants. The propagation direction and vector length of each grid are extracted from the directional field data. The propagation direction can reflect the main direction of pollution diffusion, and the vector length can reflect the intensity of pollution diffusion.

[0053] Multiplying spatially relevant elements (such as coverage growth rate) in the spatiotemporal co-evolutionary characteristics with the vector length in the directional field data yields enhanced spatial features, highlighting the influence of diffusion intensity on spatial characteristics. Multiplying temporally relevant elements (such as concentration change rate) in the spatiotemporal co-evolutionary characteristics with directional stability (standard deviation of the angle change of the propagation direction over 10 consecutive time steps in a grid; a smaller standard deviation indicates higher stability) yields enhanced temporal features, highlighting the influence of directional change on temporal characteristics. Based on historical pollution event data for the region, characteristic vector benchmarks for typical pollution types are determined. The average characteristic vectors of two main pollution types—farmland runoff pollution and livestock wastewater discharge—are calculated separately as two typical examples. The feature vector benchmarks for different pollution types are used. For example, the concentration change rate in the feature vector corresponding to farmland runoff is typically 0.2 mg / L per hour, and the coverage growth rate is typically 5% per hour. The concentration change rate in the feature vector corresponding to livestock wastewater is typically 0.5 mg / L per hour, and the coverage growth rate is typically 10% per hour. The cosine similarity between the enhanced feature vector and the benchmark feature vectors for the two typical pollution types is calculated. The similarity is calculated by multiplying the corresponding elements of the enhanced feature vector and the benchmark vector, summing the results, and then dividing by the product of the magnitudes of the two vectors. The magnitude is the square root of the sum of the squares of the elements in the vectors. The pollution type corresponding to the benchmark with the highest similarity (greater than 0.8) is selected as the identified pollution type. Finally, the evolution stage is determined by combining the concentration change trend in the time features. If the concentration rises rapidly within 0 to 6 hours, it is determined to be in the early stage of diffusion; if the concentration stabilizes or slowly decreases within 6 to 24 hours, it is determined to be in the middle stage of diffusion; if the concentration drops to near the background value after 24 hours, it is determined to be in the stable stage of diffusion. The pollution type and evolution stage are integrated to form the pollutant identification result based on pollution type and evolution stage.

[0054] This embodiment, through pollution identification, can combine historical templates and feature associations to determine the type and stage of pollution, making up for the deficiency of pollution identification relying on single data and making the identification results more reliable.

[0055] In a preferred embodiment of the present invention, the identification result and the real-time state of the optimized pollutant diffusion model are jointly input into a spatiotemporal convolutional neural network to predict the pollutant diffusion time series. Risk level distribution data is obtained through risk level quantification and spatial interpolation, which may include: The identification results are fused with the real-time status of the optimized pollutant diffusion model to obtain the dynamic evolution features of pollutants. The dynamic evolution features of pollutants are then input into a spatiotemporal convolutional neural network for spatiotemporal feature extraction to obtain deep feature data. Specifically, the identification results are converted into numerical features. The main pollution types in this area are farmland runoff and aquaculture discharge. According to historical pollution event statistics, these two account for more than 90%, so farmland runoff pollution is set as 1 and aquaculture discharge is set as 2. The diffusion stages are divided into the initial stage (0 to 6 hours, concentration rises rapidly), the middle stage (6 to 24 hours, concentration stabilizes or slowly decreases), and the stable stage (after 24 hours, concentration approaches the background value), which are set to 0.2, 0.5, and 0.8 respectively. These values ​​are arranged in the order of pollution type and evolution stage to form the identification feature vector of each grid.

[0056] The model updates its parameters hourly based on the latest real-time water quality data. The real-time state includes the pollutant concentration value of each grid at the current moment, the real-time calculated water flow diffusion coefficient (reflecting the impact of the current water flow on the diffusion of pollutants), and the real-time calculated pollutant degradation rate (reflecting the decomposition rate of pollutants under the current environmental conditions). These three parameters are arranged in the order of concentration, diffusion coefficient, and degradation rate to form the state vector of each grid. The state vectors of all grids are arranged into a state matrix according to their spatial location. Multi-dimensional feature fusion involves expanding the state vector of each grid by appending the identification feature vector element of that grid to the end of the state vector. The fused vectors of all grids are arranged in the order of time, space, and feature. Each time step corresponds to a set of fused vectors of spatial grids, forming a dynamic evolution feature that reflects the real-time changes of pollutants.

[0057] The dynamic evolution characteristics are preprocessed by selecting six time steps, including the current moment and the previous five moments, to form a six-time-step feature sequence, ensuring the capture of recent pollution trends. Each feature value is standardized to a range of 0 to 1. Standardization is achieved by subtracting the minimum value of the feature within the six time steps and then dividing by the difference between the maximum and minimum values. For example, if the minimum concentration of a certain grid cell is 3.0 mg / L and the maximum is 5.0 mg / L within the six time steps, the standardized value of the current concentration of 4.5 mg / L is (4.5 - 3.0) / (5.0 - 3.0) = 0.75, ensuring... The data format meets the network input requirements. The spatiotemporal convolutional neural network is fine-tuned. Since the original network input had 72 time steps, the current input has 6 time steps, so the input layer dimension needs to be adjusted to 6×100×100×5 (5 features: concentration, diffusion coefficient, degradation rate, pollution type, and evolution stage). Simultaneously, the temporal dimension of the spatiotemporal convolutional layer kernel is adjusted from 3 to 2 to adapt to the 6-time-step data length. Fifty recently collected dynamic evolution feature data points are selected as new samples for fine-tuning the network parameters. During fine-tuning, the learning rate is kept at 0.0001, and the network is trained for 5 epochs to ensure it can adapt to the new input dimension and feature type.

[0058] The preprocessed 6-time-step dynamic evolution features are input into the fine-tuned network. The data is first converted into tensor format through the input layer and then enters the spatiotemporal convolutional layer. The convolutional kernels simultaneously extract temporal features such as concentration changes, diffusion coefficient adjustments, and degradation rate changes within the 6 time steps, as well as spatial distribution features between different grids. The pooling layer downsamples the convolutional features, retaining key spatiotemporal correlation features. The fully connected layer integrates the pooled features into a one-dimensional vector and filters out deep information closely related to pollution diffusion through neuron weight adjustment, such as the correlation between pollution type and diffusion coefficient, and the correspondence between evolution stage and concentration change. Finally, 256-dimensional deep feature data is output.

[0059] By correlating deep feature data with the directional field data of pollutant diffusion paths, enhanced spatiotemporal evolution features are obtained, and time-series predictive analysis is performed to obtain preliminary prediction results. These preliminary prediction results are then validated and calibrated with real-time water quality parameter data to obtain pollutant diffusion prediction data. Specifically, this includes: separating spatially relevant features (such as concentration differences between different grids) and temporally relevant features (such as the rate of concentration change over time) from the deep feature data; multiplying the spatially relevant features by the vector length in the directional field data (a larger vector length indicates higher pollution diffusion intensity in that grid, thus enhancing the influence of diffusion intensity in the spatial features); and multiplying the temporally relevant features by the rate of directional change (the angular difference in the propagation direction of a grid at adjacent time points divided by the time interval), a larger rate of directional change indicates more frequent changes in diffusion direction, thus enhancing the temporal features. The influence of changes in the direction of the phenomenon is considered. The enhanced spatial and temporal features are sequentially spliced ​​together to form enhanced spatiotemporal evolution features. A sliding window method is adopted, with the window size set to 6 time steps. That is, the enhanced spatiotemporal evolution features of the past 6 time steps are used to predict the pollutant concentration of the next time step. For example, the features of the 1st to 6th hours are used to predict the concentration of the 7th hour, the features of the 2nd to 7th hours are used to predict the concentration of the 8th hour, and so on, continuously predicting the concentration value of the next 24 hours. During the prediction process, if the calculated predicted concentration value is lower than 0.5 mg / L (regional background concentration), it is adjusted to 0.5 mg / L. If it is higher than 10 mg / L (reasonable upper limit of regional pollution concentration), it is adjusted to 10 mg / L to ensure that the predicted concentration is within a reasonable range. The concentration values ​​of all prediction time points are integrated to form the preliminary prediction results of pollutant diffusion.

[0060] The first predicted time point in the preliminary prediction results, i.e., the hour after the current moment, is used as the calibration node. This node has the shortest time interval from the current moment, and the real-time water quality parameter data is closest to the actual situation at this node, resulting in better calibration. The predicted concentration values ​​of all grids at this calibration node are extracted. The corresponding real-time measured values ​​are then obtained. From the real-time water quality parameter data collected by ground-based IoT devices, the measured concentration value with the smallest time difference from the calibration node is found. For example, if the calibration node is hour t+1, and there is measured data at 5 minutes of hour t+1, then this data is selected as the measured value for the corresponding grid. If there is no data at this time point, then the measured data at 55 minutes of hour t and 10 minutes of hour t+1 are selected, and the average of the two is calculated as the measured value for the corresponding grid. For grids with measured values, the calibration coefficient is the measured concentration value of the grid divided by the predicted concentration value. For example, if the measured value of a grid is 4.0 mg / L, and the predicted value is... The value is 4.5 mg / L, and the calibration coefficient is 4.0 / 4.5≈0.889. For grids without measured values, the calibration coefficient is calculated using linear interpolation. Three grids with measured values ​​are selected around the grid. The weight of each surrounding grid is calculated based on its distance from the grid (the closer the distance, the greater the weight). The weight is the reciprocal of the distance. The calibration coefficients of the surrounding grids are multiplied by their corresponding weights, summed, and then divided by the total weights to obtain the calibration coefficient of the grid without measured values. The preliminary prediction results are adjusted using the calibration coefficient. The predicted concentration value of each grid at each time point in the next 24 hours is multiplied by the corresponding calibration coefficient to obtain the calibrated concentration value. For example, if the predicted value of a grid in the next t+2 hours is 5.0 mg / L and the calibration coefficient is 0.889, the calibrated concentration value is 5.0×0.889≈4.445 mg / L. All calibrated concentration values ​​are integrated to form the final pollutant diffusion prediction data.

[0061] The predicted data of pollutant diffusion is compared and analyzed with historical pollutant concentration data to obtain the trend of pollutant concentration changes. Based on the trend of pollutant concentration changes, risk levels are classified to obtain the quantitative results of pollutant risk levels. The quantitative results of pollutant risk levels are then spatially interpolated to obtain risk level distribution data. Specifically, this includes: collecting historical concentration data of the region for the same period of the past 5 years as the predicted data, where the same period refers to the same month and date range; for the historical data of the same period each year, calculating the monthly average concentration (summing the daily average concentrations of all days in the same period of the month and then dividing by the number of days), the monthly maximum concentration (the maximum of the daily maximum concentrations of all days in the same period of the month), and the monthly minimum concentration (the minimum of the daily minimum concentrations of all days in the same period of the month), forming historical concentration statistics.

[0062] The predicted concentration is calculated daily (summing the concentration values ​​of all grids each day and dividing by the total number of grids), daily maximum concentration (the highest concentration among all grids each day), and daily minimum concentration (the lowest concentration among all grids each day). If the predicted data spans multiple months, such as from May 30th to June 4th, the data from May 30th to 31st is included in the May statistics, and the data from June 1st to 4th is included in the June statistics, ensuring consistency with the monthly statistical standards of historical data. Comparative analysis involves subtracting the daily average concentration of the predicted data from the historical monthly average concentration for the same period. A positive difference indicates a higher predicted concentration. Compared to historical averages, pollutant concentrations show an upward trend. If the difference is negative, it indicates that the predicted concentration is lower than the historical average, and the pollutant concentration shows a downward trend. If the difference is close to zero (absolute value less than 0.1 mg / L), it indicates that the predicted concentration is basically the same as the historical average, and the pollutant concentration shows a stable trend. By comparing the predicted daily maximum concentration with the historical monthly maximum concentration for the same period, if the predicted daily maximum concentration is higher than the historical monthly maximum concentration, it indicates that the peak concentration shows an upward trend; if it is lower, it shows a downward trend; if it is close, it shows a stable trend. Integrating the overall concentration trend and the peak trend, a pollutant concentration change trend is formed.

[0063] Risk level classification standards were established, and based on the actual situation of agricultural non-point source pollution in the middle and lower reaches of the Yangtze River and lakeside areas, concentration ranges for low, medium, and high risk were determined. Low risk is defined as COD < 2 mg / L, ammonia nitrogen < 0.5 mg / L, and total phosphorus < 0.02 mg / L, corresponding to quantitative value 1; medium risk is defined as COD 2 to 5 mg / L, ammonia nitrogen 0.5 to 1.0 mg / L, and total phosphorus 0.02 to 0.05 mg / L, corresponding to quantitative value 2; high risk is defined as COD > 5 mg / L, ammonia nitrogen > 1.0 mg / L, and total phosphorus > 0.05 mg / L, corresponding to quantitative value 3. This ensures that the classification standards meet the regional pollution prevention and control needs. Risk levels are adjusted based on concentration trends. If the concentration in a certain grid is at the upper limit of medium risk (e.g., COD = 5 mg / L, ammonia nitrogen = 1.0 mg / L, and total phosphorus = 0.05 mg / L) and shows an upward trend, it indicates that the pollution may worsen further, and the risk level is adjusted accordingly. The risk level is upgraded by one level, from medium risk (quantification 2) to high risk (quantification 3). If the concentration of a certain grid is at the upper limit of low risk (e.g., COD=2mg / L, ammonia nitrogen=0.5mg / L, total phosphorus=0.02mg / L) and shows a downward trend, it indicates that the pollution is decreasing, and the low risk level (quantification 1) remains unchanged. If the concentration shows a stable trend, the quantification value is determined directly according to the risk level corresponding to the current concentration. The risk level of each grid is quantified, and the risk level corresponding to the COD, ammonia nitrogen, and total phosphorus concentrations of each grid is determined separately. The highest risk level among the three parameters is taken as the comprehensive risk level of the grid. For example, if a certain grid corresponds to medium risk (quantification 2) for COD, medium risk (quantification 2) for ammonia nitrogen, and high risk (quantification 3) for total phosphorus, then the comprehensive risk level quantification value of the grid is 3. The comprehensive risk level quantification values ​​of all grids are integrated to form the pollutant risk level quantification result.

[0064] Kriging interpolation was employed. This method can interpolate based on the spatial correlation of sample points, adapting to the uneven data distribution in the tributaries and lakeside areas of the middle and lower reaches of the Yangtze River. The grid data in the polder areas is dense, while that in the lake areas is sparse. Kriging interpolation can accurately estimate the risk level of sparse areas through the correlation between sample points. Next, interpolation sample points were prepared. Sample points were selected from the risk level quantification results. In the polder areas, one sample point was selected at the center of every 10,000 square meters of grid space to ensure coverage of high, medium, and low risk levels within the polder areas. In tributary channels, one sample point was selected at the center of every 5,000 square meters of channel cross-section, evenly distributed along the channel direction. In the lakeside areas, one sample point was selected at the center of every 20,000 square meters of grid space, focusing on covering the transition area between the nearshore and the lake center. Each sample point recorded its precise longitude and latitude (accurate to six decimal places) and the corresponding quantified risk level value, for example, 116.0030 degrees east longitude and 30.00 degrees north latitude. The quantization value of a 35-degree sample point is 2, ensuring that the sample points uniformly cover the entire monitoring area and reflect the risk differences in different areas. The spatial correlation index is used to describe the degree of correlation between sample points. All sample points are grouped according to the straight-line distance between each pair, with the grouping interval set at 500 meters, i.e., 0 to 500 meters, 500 to 1000 meters, 1000 to 1500 meters, etc. The average distance and average semivariance of all sample point pairs within each group are calculated. The semivariance is calculated by taking the difference in the quantization value of each sample point pair within the group, square the difference, divide it by 2, and then calculate the average semivariance of all sample point pairs to obtain the average semivariance of the group.

[0065] By adjusting three key parameters (the sill value reflects the maximum difference between sample points, the range reflects the effective distance of spatial correlation, and the nugget value reflects the error caused by random factors), the deviation between the semivariogram calculated based on these parameters and the actual mean semivariogram is minimized, thereby determining the specific characteristics of the spatial correlation of sample points. The range in this area is typically in the range of 500 to 1500 meters. Spatial interpolation is performed by finding 3 to 5 closest sample points within the above range for each 100m × 100m grid to be interpolated, using the grid center point as a reference. The distance between the sample points and the grid center point is typically controlled between 100 and 1000 meters (too close a distance can lead to sample duplication, while too far a distance exceeds the range of spatial correlation). Weights are calculated based on the distance between the sample points and the grid center point; the greater the distance, the higher the weight. The closer the distance, the greater the weight. The weight value is the reciprocal of the distance, ranging from 0.001 to 0.01 (the weight is 1 / 100=0.01 for a distance of 100 meters and 1 / 1000=0.001 for a distance of 1000 meters). For example, the weight of a sample point at a distance of 200 meters is 0.005, the weight of a sample point at a distance of 400 meters is 0.0025, and the weight of a sample point at a distance of 500 meters is 0.002. The risk level quantification value of each sample point is multiplied by its corresponding weight, summed, and then divided by the sum of all weights (the sum of all weights is usually in the range of 0.005 to 0.025) to obtain the risk level quantification value of the grid to be interpolated. The quantification values ​​of all grids to be interpolated and the original sample point grids are integrated to form risk level distribution data covering the entire monitoring area.

[0066] This embodiment, through risk quantification, can determine the risk level by combining concentration and trend, making up for the deficiency of risk classification that ignores the influence of trend, and making the risk level more in line with the development of pollution.

[0067] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the system as described above. All implementations in the above system embodiments are applicable to this embodiment and can achieve the same technical effects.

[0068] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the system as described above. All implementations in the above system embodiments are applicable to this embodiment and can achieve the same technical effects.

[0069] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A dynamic monitoring system for the spatiotemporal distribution of water pollution based on multi-source remote sensing data fusion and deep learning, characterized in that, include: The acquisition module is used to acquire satellite hyperspectral imagery, UAV multispectral data, and water quality parameter data collected in real time through ground-based IoT devices; The execution module is used to invert pollutant concentration from satellite hyperspectral imagery to obtain pollutant concentration data; locate local pollution sources from UAV multispectral data to obtain pollution source location data; and acquire real-time water quality parameter data from ground-based IoT devices. The processing module is used to spatially discretize pollutant concentration data, pollution source location data, and real-time water quality parameter data, establish a quadtree index, and recursively divide it into sub-regions; and generate dynamic spatial correction parameters based on the statistical characteristics of data points in each sub-region. The optimization module is used to optimize the parameters of the pollutant diffusion numerical model constructed by fusing multi-source data through data assimilation technology based on the dynamic spatial correction parameters, so as to obtain the optimized pollutant diffusion model. The extraction module is used to input the spatiotemporal sequence data output by the optimized pollutant diffusion model into the spatiotemporal convolutional neural network to extract the spatiotemporal co-evolution features of pollutants and obtain the identification results. The evaluation module is used to input the identification results and the real-time status of the optimized pollutant diffusion model into a spatiotemporal convolutional neural network to predict the pollutant diffusion time series, and obtain risk level distribution data through risk level quantification and spatial interpolation.

2. The water pollution spatiotemporal distribution dynamic monitoring system based on multi-source remote sensing data fusion and deep learning according to claim 1, characterized in that, Pollutant concentration inversion was performed on satellite hyperspectral imagery to obtain pollutant concentration data; local pollution source location data was obtained from UAV multispectral data; real-time water quality parameter data was acquired from ground-based IoT devices, including: Satellite hyperspectral images are processed to obtain pollutant concentration data; spatial variation characteristics in the pollutant concentration data are analyzed to identify concentration anomaly regions. UAV multispectral data collection and analysis were conducted on areas with abnormal concentrations to obtain pollution source location data; Optimize the spatial deployment scheme of ground-based IoT devices based on pollution source location data to obtain real-time water quality parameter data for the core pollution area and diffusion path.

3. The water pollution spatiotemporal distribution dynamic monitoring system based on multi-source remote sensing data fusion and deep learning according to claim 2, characterized in that, Pollutant concentration data, pollution source location data, and real-time water quality parameter data are spatially discretized, a quadtree index is established, and the data is recursively divided into sub-regions. Based on the statistical characteristics of data points within each sub-region, dynamic spatial correction parameters are generated, including: By projecting pollutant concentration data, pollution source location data, and real-time water quality parameter data onto the same spatial coordinate system, a spatially registered multi-source dataset is obtained. Spatial discretization is performed on the multi-source dataset for spatial registration to obtain a discretized data point set; based on the spatial distribution density of the discretized data point set, a quadtree spatial index is established to obtain the initial spatial partition; The initial spatial partition is recursively divided to obtain a set of sub-regions, and the statistical characteristics of the data points in each sub-region set are calculated to obtain the dynamic spatial correction parameters.

4. The water pollution spatiotemporal distribution dynamic monitoring system based on multi-source remote sensing data fusion and deep learning according to claim 3, characterized in that, Based on the dynamic spatial correction parameters, the pollutant diffusion numerical model constructed by fusing multi-source data through data assimilation technology is optimized to obtain the optimized pollutant diffusion model, including: The dynamic spatial correction parameters are fused with the initial parameters of the pollutant diffusion numerical model to generate parameter optimization criteria. Based on parameter optimization criteria, pollutant concentration data, pollution source location data, and real-time water quality parameter data are processed to obtain a multi-source dataset fused by data assimilation technology. The multi-source dataset is iteratively matched with the pollutant diffusion numerical model to obtain the adjusted model parameter combination; The adjusted model parameter combination replaces the original parameters in the pollutant diffusion numerical model, resulting in a parameter-optimized pollutant diffusion model.

5. The water pollution spatiotemporal distribution dynamic monitoring system based on multi-source remote sensing data fusion and deep learning according to claim 4, characterized in that, The spatiotemporal sequence data output from the optimized pollutant diffusion model is input into a spatiotemporal convolutional neural network to extract the spatiotemporal co-evolution features of pollutants, yielding the identification results, including: The spatiotemporal sequence data output by the optimized pollutant diffusion model are processed to obtain the spatiotemporal distribution characteristics of pollutants; the changes in pollutant concentration gradient in the spatiotemporal distribution characteristics are analyzed to obtain pollutant propagation direction data. The pollutant propagation direction data is used with a ray algorithm to obtain the directional field data of the pollutant diffusion path; the directional field data and the spatiotemporal sequence data are input into a spatiotemporal convolutional neural network to obtain the spatiotemporal co-evolution characteristics of pollutants. By integrating the spatiotemporal co-evolution characteristics of pollutants with the directional field data of pollutant diffusion paths, the pollutant identification results are obtained.

6. The water pollution spatiotemporal distribution dynamic monitoring system based on multi-source remote sensing data fusion and deep learning according to claim 5, characterized in that, The pollutant propagation direction data is processed using a ray-mapping algorithm to obtain the directional field data of the pollutant diffusion path. This directional field data, along with spatiotemporal sequence data, is then input into a spatiotemporal convolutional neural network to obtain the spatiotemporal co-evolutionary characteristics of the pollutants, including: Process the pollutant propagation direction data to obtain the main direction set of pollutant diffusion; execute the ray tracing algorithm based on the main direction set to obtain the ray cluster distribution of pollutant diffusion; Spatial matching of the ray cluster distribution with real-time water quality parameter data yields the corrected ray distribution results; the corrected ray distribution results are then reconstructed into a directional field to obtain the directional field data of the pollutant diffusion path. The directional field data is fused with the spatiotemporal sequence data to obtain fused spatiotemporal feature data; the fused spatiotemporal feature data is then input into a spatiotemporal convolutional neural network for deep feature extraction to obtain the spatiotemporal co-evolutionary features of pollutants.

7. The water pollution spatiotemporal distribution dynamic monitoring system based on multi-source remote sensing data fusion and deep learning according to claim 6, characterized in that, The identification results and the real-time status of the optimized pollutant diffusion model are input into a spatiotemporal convolutional neural network to predict the pollutant diffusion time series. Risk level distribution data is obtained through risk level quantification and spatial interpolation, including: The identification results are fused with the real-time status of the optimized pollutant diffusion model to obtain the dynamic evolution characteristics of pollutants; the dynamic evolution characteristics of pollutants are then input into a spatiotemporal convolutional neural network for time series analysis to obtain the predicted data of pollutant diffusion. By comparing and analyzing the predicted data of pollutant diffusion with historical pollutant concentration data, the trend of pollutant concentration change is obtained; based on the trend of pollutant concentration change, the risk level is classified to obtain the quantitative result of pollutant risk level; the quantitative result of pollutant risk level is spatially interpolated to obtain the risk level distribution data.

8. The water pollution spatiotemporal distribution dynamic monitoring system based on multi-source remote sensing data fusion and deep learning according to claim 7, characterized in that, The identification results are fused with the real-time status of the optimized pollutant diffusion model to obtain the dynamic evolution characteristics of pollutants. By inputting the dynamic evolution characteristics of pollutants into a spatiotemporal convolutional neural network for time-series analysis, predictive data on pollutant diffusion are obtained, including: The identification results are fused with the real-time status of the optimized pollutant diffusion model to obtain the dynamic evolution features of pollutants; the dynamic evolution features of pollutants are then input into a spatiotemporal convolutional neural network for spatiotemporal feature extraction to obtain deep feature data. By correlating deep feature data with the directional field data of pollutant diffusion paths, enhanced spatiotemporal evolution features are obtained, and time-series prediction analysis is performed to obtain preliminary prediction results. The preliminary prediction results are then verified and calibrated with real-time water quality parameter data to obtain pollutant diffusion prediction data.

9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the system as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the system as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Rapid prediction and early warning method and system for short-term water quality pollution, medium and equipment

    CN118673407A

  • River water quality parameter supervision method and system based on deep learning

    CN120319366A

  • Water quality change trend rapid prediction method based on multi-source data fusion and physical constraint

    CN120598102A

  • River water quality monitoring method based on multi-source remote sensing data

    CN120690318A

Cited By

  • Steam condensate recovery system with automatic cleaning function

    CN121800380A