Ecological environment pollution source tracing method based on data analysis

Through the method of identifying pollution sources through data format standardization, abnormal detection and filling, pollution diffusion modeling and machine learning, combined with blockchain technology, data fusion and accuracy problems in pollution traceability methods are solved, accurate identification and diffusion simulation of pollution sources are achieved, and scientific pollution control support is provided.

CN120296329AInactive Publication Date: 2025-07-11甘肃省张掖生态环境监测中心
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510440247.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing pollution traceability methods have technical bottlenecks in data processing and pollution diffusion simulation, resulting in insufficient accuracy and traceability of pollution source identification. The data format, time accuracy and coordinate system differences of different equipment make it difficult to integrate data, and abnormal data and missing data exist, affecting the accuracy of pollution traceability analysis.

Method used

The time stamp synchronization algorithm and coordinate conversion method are used to unify the data format and geographical coordinates, and the local anomaly factor algorithm is used to detect abnormal data and combine wavelet transformation for denoising and completing. Gaussian smoke plume, fluid dynamics and Fick diffusion models are built for pollution diffusion simulation. Combined with machine learning models, identify pollution sources and calculate the contribution rate through Bayesian regression, and use blockchain technology to ensure the security and traceability of data.

Benefits of technology

It improves the accuracy and reliability of pollution traceability, realizes the optimization of accurate locking and diffusion simulation of pollution sources, provides a scientific basis for pollution control, and displays pollution paths and hot spots through the GIS visualization platform to ensure the safety and traceability of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296329A_ABST
    Figure CN120296329A_ABST
Patent Text Reader

Abstract

The invention discloses an ecological environment pollution source tracing method based on data analysis, and relates to the field of ecological environment, and the method comprises the steps: carrying out the type selection of equipment, obtaining original data, carrying out the arrangement of the equipment after the type selection of the equipment, carrying out the multi-source data access after the arrangement of the equipment, and carrying out the data format standardization based on the obtained original data. And performing data anomaly detection based on the original data after data format standardization, and performing vacancy value filling based on the original data after data anomaly detection. According to the method, through multi-source data fusion, data cleaning, pollution diffusion modeling, pollution source identification and contribution rate calculation and visual early warning, the precision and reliability of pollution tracing are improved, firstly, format standardization, time synchronization and anomaly detection are carried out on pollution monitoring data of different sources, and data consistency and integrity are ensured; and then combining a Gaussian plume model, a fluid dynamic model and a Fick diffusion model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of ecological environment, and in particular to a method for tracing the sources of ecological environment pollution based on data analysis. Background Art

[0002] With the rapid economic development and the acceleration of the industrialization process, the problem of ecological environment pollution has become increasingly serious. The diffusion process of pollutants in media such as the atmosphere, water bodies and soil is complex, and the causes of pollution are diverse, posing great challenges to the accurate tracing of pollution sources. Traditional pollution monitoring and tracing methods mainly rely on single-point sampling, laboratory analysis or simple mathematical model calculations, resulting in limited accuracy in inferring pollution diffusion paths and identifying pollution sources. In recent years, with the development of Internet of Things, data analysis and artificial intelligence technologies, pollution tracing methods based on multi-source data analysis have gradually become a research hotspot. By integrating multiple types of environmental data, constructing pollution diffusion models and intelligently identifying pollution sources, new technical means have been provided for environmental supervision and pollution control. However, there are still certain technical bottlenecks in the existing pollution tracing methods in terms of data processing and pollution diffusion simulation, which affect the accuracy of pollution source identification and the scientific nature of tracing.

[0003] However, in the prior art, due to the wide range of sources of pollution monitoring data, including air quality sensors, water quality analyzers, soil monitoring equipment and satellite remote sensing data, there are differences in data formats, time accuracies and coordinate systems of different devices, resulting in difficulty in directly integrating the data. At the same time, the monitoring data may be affected by factors such as sensor failures and environmental interferences, and there are abnormal data and missing data. Traditional data processing methods lack unified standardization, anomaly detection and data completion mechanisms, resulting in low data quality and affecting the accuracy of pollution tracing analysis. Summary of the Invention

[0004] The object of the present invention is to provide a method for tracing the sources of ecological environmental pollution based on data analysis in view of the deficiencies of the prior art, aiming to improve the accuracy of pollution source tracing through technical means such as data format standardization, abnormal data detection and filling, and pollution diffusion modeling. First, the present invention adopts a timestamp synchronization algorithm and a coordinate conversion method to unify the time base and geographical coordinate format of different data sources, ensuring data compatibility and consistency. At the same time, the Local Outlier Factor (LOF) algorithm is used to detect abnormal data, and methods such as wavelet transform and K-nearest neighbor interpolation are combined for data denoising and completion to improve data integrity. Secondly, the present invention combines the Gaussian plume model, the hydrodynamic model, and the Fick diffusion model to construct a refined diffusion path simulation according to the propagation medium of pollutants, comprehensively considering environmental factors such as wind speed, wind direction, water flow velocity, and soil permeability to improve the accuracy of pollution diffusion prediction. In addition, the present invention further uses a machine learning model to identify the categories of pollution sources, and calculates the pollution contribution rate through Bayesian regression to accurately quantify the impact of each pollution source on environmental pollution, ensuring the scientific nature and traceability of the pollution source tracing results, thereby providing reliable data support for ecological environment governance and pollution liability determination.

[0005] To this end, the present application provides a method for tracing the sources of ecological environmental pollution based on data analysis, including the following steps: Select equipment types and obtain raw data. Based on the selected equipment types, arrange the equipment, and based on the arranged equipment, access multi-source data.

[0006] Based on the obtained raw data, perform data format standardization. Based on the raw data after data format standardization, perform data anomaly detection. Based on the raw data after data anomaly detection, fill in the missing values.

[0007] Based on the raw data after filling in the missing values, perform atmospheric pollution diffusion simulation. Based on the raw data after atmospheric pollution diffusion simulation, perform water pollution diffusion modeling. Based on the raw data after water pollution diffusion modeling, perform soil pollution diffusion analysis.

[0008] Based on the raw data after soil pollution diffusion analysis, classify the pollution sources. Based on the classified pollution sources, calculate the pollution contribution rate. Based on the calculated pollution contribution rate, perform blockchain pollution tracing verification.

[0009] Generate a GIS pollution diffusion heat map based on the raw data. Based on the raw data after generating the GIS pollution diffusion heat map, perform intelligent pollution warning and set a warning threshold.

[0010] In some specific embodiments, selecting equipment types and obtaining raw data. Based on the selected equipment types, arrange the equipment, and based on the arranged equipment, access multi-source data specifically includes: Select equipment types and obtain raw data.

[0011] PM2.5, PM10, SO2, NO2, CO, O3 and volatile organic compounds are obtained through light scattering dust sensors, electrochemical gas sensors and lidar sensors. The sampling period of light scattering dust sensors, electrochemical gas sensors and lidar sensors is set to once every 10 seconds to 5 minutes.

[0012] The chemical oxygen demand, biochemical oxygen demand, ammonia nitrogen, total phosphorus, total nitrogen and heavy metals are obtained through the ultraviolet spectrum water quality analyzer, electrode sensor and ion selective electrode sensor. The collection cycle of the ultraviolet spectrum water quality analyzer, electrode sensor and ion selective electrode sensor is set to once every 30 seconds to 10 minutes.

[0013] Heavy metals, organic pollutants and soil pH are obtained through X-ray fluorescence spectrometer, ion chromatograph and infrared spectrometer. The sampling period of X-ray fluorescence spectrometer, ion chromatograph and infrared spectrometer is set to once every 1 to 12 hours.

[0014] Equipment layout is carried out based on equipment selection.

[0015] Equipment deployment includes: air pollution sensor deployment, water pollution sensor deployment and soil pollution sensor deployment.

[0016] Air pollution sensor deployment: install high-density sensor grids in high-pollution areas at intervals of 500m-1km, install more accurate monitoring equipment in environmentally sensitive areas at intervals of 1km-2km, use background monitoring points to obtain baseline data, and deploy 2-5 long-term monitoring points.

[0017] Water pollution sensors are deployed to monitor the background values ​​of inflowing pollutants in upstream waters, focus on monitoring industrial, agricultural and domestic sewage discharge near sewage outlets, monitor the mixing of pollution in different water bodies at river confluences, and monitor the degree of eutrophication of water bodies at central points of reservoirs and lakes.

[0018] Soil pollution sensors are deployed to focus on monitoring agricultural pollution in farmland areas, monitoring soil pollution caused by industrial wastewater and waste residue around industrial areas, and monitoring the accumulation of harmful substances in the soil in landfills.

[0019] Multi-source data access is carried out after equipment deployment.

[0020] The deployed sensors will be connected to the IoT cloud platform. Sensor data collection will collect pollutant concentration data according to the set period, and edge computing preprocessing will perform preliminary data cleaning in the sensor's local computing unit to eliminate obviously abnormal data points. The data will then be uploaded to the central server through 5G wireless communication technology.

[0021] Access satellite remote sensing data, use Landsat-8, MODIS, and Sentinel-2 satellite image data to monitor the distribution of large-scale pollution. Combine spectral analysis to extract water quality pollution and air pollution data, including: water body suspended solids, vegetation cover changes, and industrial emission concentrations.

[0022] Collect wind speed, wind direction, and precipitation through the meteorological center API, obtain the atmospheric boundary layer height through lidar, and collect factory pollutant emission data, including chimney waste gas and sewage discharge. Compare with the emission standards to identify enterprises that violate the standards and exceed the limits.

[0023] In some specific embodiments, perform data format standardization based on the obtained original data, perform data anomaly detection based on the original data after data format standardization, and perform missing value filling based on the original data after data anomaly detection. Specifically include: Perform data format standardization based on the obtained original data.

[0024] Adopt the timestamp synchronization algorithm to unify the time reference of all data. By setting a master clock server, all data acquisition devices request time synchronization from the master clock server, calculate the time deviation of each sensor device, and adjust the time of all devices.

[0025] The geographic coordinates of different data sources may use different coordinate systems and need to be uniformly converted to the WGS84 coordinate system. Use the ellipsoidal coordinate conversion formula for coordinate conversion.

[0026] Perform data anomaly detection based on the original data after data format standardization.

[0027] Adopt the local outlier factor algorithm to detect abnormal data points. The local outlier factor algorithm is used to detect outliers in the dataset, calculate the local outlier factor of each data point, calculate the reachable distance, local density, and LOF value.

[0028] Adopt wavelet transform for data denoising, use discrete wavelet transform for noise removal, select the Daubechies wavelet as the basis function for soft threshold denoising, and restore the denoised data through inverse wavelet transform.

[0029] Perform missing value filling based on the original data after data anomaly detection.

[0030] Adopt linear interpolation for small-scale data missing and K-nearest neighbor completion for large-scale data missing.

[0031] In some specific embodiments, atmospheric pollution diffusion simulation is performed based on the original data after missing value filling, water pollution diffusion modeling is performed based on the original data after atmospheric pollution diffusion simulation, and soil pollution diffusion analysis is performed based on the original data after water pollution diffusion modeling, specifically including: Perform atmospheric pollution diffusion simulation based on the original data after missing value filling.

[0032] Atmospheric pollution diffusion simulation includes: Gaussian plume model calculation and diffusion coefficient calculation.

[0033] The Gaussian plume model calculation is specifically: , In the formula: is the pollutant concentration at the spatial position , with the unit of , is the emission of the pollution source, with the unit of , is the wind speed, with the unit of , is the lateral diffusion coefficient, with the unit of m, is the vertical diffusion coefficient, with the unit of m, is the lateral offset distance of the pollution source, with the unit of m, is the vertical offset distance of the pollution source, with the unit of m, is the exponential function, that is, the calculation in the form of .

[0034] The diffusion coefficient calculated by the diffusion coefficient calculation depends on the atmospheric stability and the pollutant propagation distance.

[0035] Combine the original data to calculate the pollutant diffusion path, collect the wind speed, wind direction, and air temperature through the meteorological API to obtain the diffusion direction and speed of the pollutant in the air, monitor the atmospheric boundary layer height through lidar to obtain the vertical diffusion range of the pollutant, and combine the original data to correct the diffusion behavior of the pollutant in complex terrain.

[0036] Perform water pollution diffusion modeling based on the original data after atmospheric pollution diffusion simulation.

[0037] Water pollution diffusion modeling calculates the diffusion path of pollutants in water through a hydrodynamic model and pollutant degradation modeling.

[0038] The hydrodynamic model mainly calculates the hydraulic transport process of pollutants, including the diffusion, convection, and deposition of pollutants in water.

[0039] The QUAL2K water quality model is used to analyze the physical and chemical processes such as pollutant degradation, deposition, evaporation, and oxidation.

[0040] Collect the water flow velocity, direction, and water level through a hydrological monitoring station, calculate the convective transport process of pollutants, monitor the eutrophication of lake water bodies through remote sensing, evaluate the cumulative risk of pollutants, and calculate the scouring and runoff paths of pollutants in combination with the original data.

[0041] Conduct soil pollution diffusion analysis based on the original data after water pollution diffusion modeling.

[0042] The soil pollution diffusion analysis uses the Fick diffusion model to calculate the diffusion of pollutants in the soil.

[0043] Monitor the soil moisture content through a soil moisture sensor, calculate the infiltration rate of pollutants, calculate the flow path of pollutants with groundwater through groundwater monitoring, evaluate the groundwater pollution risk, and evaluate the chemical reaction characteristics of pollutants in combination with the original data.

[0044] In some specific embodiments, conduct pollution source classification based on the original data after soil pollution diffusion analysis, calculate the pollution contribution rate based on the pollution source classification, and conduct blockchain pollution tracing based on the calculation of the pollution contribution rate, specifically including: Conduct pollution source classification based on the original data after soil pollution diffusion analysis.

[0045] The purpose of pollution source classification is to classify pollution data into different pollution source categories, and use the random forest and XGBoost models for pollution source classification.

[0046] The random forest model conducts pollution source classification through the integration of multiple decision trees, and the XGBoost model uses gradient boosting trees and is improved based on the random forest.

[0047] Extract features from the pollution data to construct input variables for machine learning classification, including PM2.5, PM10, SO2, NO2, CO, O3, VOC, COD, BOD, ammonia nitrogen, total phosphorus, and heavy metals, the geographical location, terrain, and adjacent water body information of the pollution source, the daily and seasonal variations of pollutants, wind speed, wind direction, temperature, humidity, and precipitation.

[0048] Train the random forest classification model, specifically: , Where: is the final pollution source classification result, is the number of decision trees, is the prediction result of the

[0049] The random forest training process normalizes the pollutant data, divides the original data according to the ratio of 80% training set + 20% test set, sets the number of decision trees to 300, uses information gain or Gini coefficient as the feature selection criterion, uses the test set to evaluate the model accuracy, calculates the confusion matrix, recall rate, and F1-score, and based on the probability distribution of the pollution category output by the model.

[0050] Based on the pollution source classification, calculate the pollution contribution rate.

[0051] Adopt the Bayesian regression model for regression analysis, specifically: , In the formula: is the total pollutant concentration, with the unit of or , is the emission of the or , is the pollution contribution factor of the is the error term, representing unobservable noise factors.

[0052] Bayesian regression calculates the pollution contribution factor through maximum a posteriori estimation.

[0053] Based on the calculation of the pollution contribution rate, conduct blockchain pollution traceability verification.

[0054] Store the pollution tracking data through the blockchain. The block header includes the timestamp, the hash value of the previous block, and the Merkle root. The transaction data includes the pollutant concentration, pollution source classification, pollution contribution rate, and pollution diffusion path. The hash pointer is used to store the hash value of the pollution data.

[0055] Calculate the pollutant concentration and pollution contribution rate, generate the hash value of the pollution source data, generate a block and broadcast it to the blockchain network to confirm the transaction through the consensus mechanism, and store the pollution tracking data.

[0056] In some specific embodiments, generate a GIS pollution diffusion heat map based on the original data, conduct intelligent pollution warning based on the original data after generating the GIS pollution diffusion heat map, and set the warning threshold, specifically including: Generate a GIS pollution diffusion heat map based on the original data.

[0057] Build a pollution map by integrating ArcGIS and Leaflet.js. Perform data projection transformation using the WGS84 coordinate system. Calculate the distribution of pollutant concentrations in the geographical space through the Kriging interpolation method. Generate a pollution heat map using Leaflet.js, and use a gradient color, including red - yellow - green, to represent the pollution concentration.

[0058] Use a time - series animation to show the pollution diffusion trend. Combine wind speed, wind direction, and terrain data to dynamically adjust the pollution diffusion direction. Calculate the diffusion path of pollutants from the pollution source to the affected area using the Dijkstra shortest - path algorithm. Combine the Lagrangian particle - tracking model to simulate the movement trajectory of pollutants in the air / water.

[0059] Conduct intelligent pollution early warning based on the original data after generating the GIS pollution diffusion heat map.

[0060] Use a deep - learning model to predict the pollution trend and combine the original data to track pollution hotspots.

[0061] Use LSTM to predict the pollution trend. The input of the prediction model is the original data, wind speed, wind direction, temperature, precipitation, and seasonal factors.

[0062] Use the Adam optimizer to update parameters. Set the learning rate to 0.001. Use the mean - squared - error loss function. The training - set ratio is 80% and the test - set ratio is 20%. Conduct 50 rounds of training to predict the pollutant concentration. If the predicted pollutant concentration exceeds the set threshold, trigger an early warning.

[0063] Use a drone to detect pollution hotspots in real - time. Calculate the optimal cruise path through the A* search algorithm. The drone is equipped with a spectral imager to monitor industrial waste gas, oil spills, and river pollution in real - time. Use YOLOv5 to identify the pollution source. When the pollution source is identified, record the GPS coordinates of the pollution source.

[0064] In summary, the method for tracing the sources of ecological environmental pollution based on data analysis provided in this application improves the accuracy and reliability of pollution source tracing through multi-source data fusion, data cleaning, pollution diffusion modeling, pollution source identification and contribution rate calculation, and visualization warning. This method first standardizes the formats of pollution monitoring data from different sources, synchronizes the time, and detects anomalies to ensure data consistency and integrity. Then, in combination with the Gaussian plume model, the hydrodynamic model, and the Fick diffusion model, it establishes accurate diffusion path simulations for air, water, and soil pollution to improve the accuracy of pollution propagation prediction. Furthermore, machine learning techniques are used to classify pollution sources, and Bayesian regression is used to calculate the pollution contribution rate to achieve quantitative tracing of pollution liability. In addition, this invention combines blockchain technology to ensure the security, immutability, and traceability of pollution source tracing data, and constructs a GIS visualization platform to intuitively display the pollution diffusion path and pollution hotspots and provide intelligent warnings. Compared with traditional methods, this invention can accurately lock in pollution sources, optimize pollution diffusion simulations, and improve pollution control efficiency, providing a scientific basis for ecological environment monitoring and pollution liability determination. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 is the overall flowchart of the method for tracing the sources of ecological environmental pollution based on data analysis provided in an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] Please refer to Figure 1 , which shows the process of an embodiment of the method for tracing the sources of ecological environmental pollution based on data analysis according to the present disclosure As Figure 1 shown, the method for tracing the sources of ecological environmental pollution based on data analysis includes the following steps: Select equipment models and obtain original data. Based on the selected equipment models, arrange the equipment, and based on the arranged equipment, access multi-source data.

[0067] Standardize the data formats based on the obtained original data, detect data anomalies based on the original data after data format standardization, and fill in missing values based on the original data after data anomaly detection.

[0068] Conduct air pollution diffusion simulations based on the original data after filling in missing values, establish water pollution diffusion models based on the original data after air pollution diffusion simulations, and analyze soil pollution diffusion based on the original data after water pollution diffusion modeling.

[0069] Classify pollution sources based on the original data after soil pollution diffusion analysis, calculate the pollution contribution rate based on the classified pollution sources, and conduct blockchain pollution tracing certification based on the calculated pollution contribution rate.

[0070] Generate a GIS pollution diffusion heat map based on the original data, perform intelligent pollution early warning based on the original data after generating the GIS pollution diffusion heat map, and set an early warning threshold.

[0071] In some specific embodiments, select equipment and obtain the original data. Based on the equipment selection, conduct equipment layout, and based on the equipment layout, perform multi-source data access, specifically including: Select equipment and obtain the original data.

[0072] Through a light scattering dust sensor, an electrochemical gas sensor, and a lidar sensor, obtain PM2.5, PM10, SO2, NO2, CO, O3, and volatile organic compounds. Set the sampling period of the light scattering dust sensor, the electrochemical gas sensor, and the lidar sensor to collect once every 10 seconds - 5 minutes.

[0073] Through an ultraviolet spectrum water quality analyzer, an electrode sensor, and an ion-selective electrode sensor, obtain chemical oxygen demand, biochemical oxygen demand, ammonia nitrogen, total phosphorus, total nitrogen, and heavy metals. Set the collection period of the ultraviolet spectrum water quality analyzer, the electrode sensor, and the ion-selective electrode sensor to collect once every 30 seconds - 10 minutes.

[0074] Through an X-ray fluorescence spectrometer, an ion chromatograph, and an infrared spectrometer, obtain heavy metals, organic pollutants, and soil acidity. Set the sampling period of the X-ray fluorescence spectrometer, the ion chromatograph, and the infrared spectrometer to collect once every 1 hour - 12 hours.

[0075] Based on the equipment selection, conduct equipment layout.

[0076] Equipment layout includes: air pollution sensor layout, water pollution sensor layout, and soil pollution sensor layout.

[0077] Air pollution sensor layout: Install a high-density sensor grid in high-pollution areas at intervals of 500m - 1km, install more accurate monitoring equipment in environmentally sensitive areas at intervals of 1km - 2km, and use background monitoring points to obtain baseline data, with 2 - 5 long-term monitoring points arranged.

[0078] Water pollution sensor layout: Monitor the background value of pollutants flowing into the upstream water area, focus on monitoring the discharge of industrial, agricultural, and domestic sewage near the sewage outlet, monitor the mixing of different water body pollutions at the river confluence, and monitor the eutrophication degree of the water body at the central points of reservoirs and lakes.

[0079] Soil pollution sensor layout: Focus on monitoring agricultural pollution in farmland areas, monitor the pollution of industrial wastewater and waste residues to the soil around industrial areas, and monitor the accumulation of harmful substances in the soil at landfills.

[0080] Multi-source data access is carried out after equipment deployment.

[0081] The deployed sensors will be connected to the IoT cloud platform. Sensor data collection will collect pollutant concentration data according to the set period, and edge computing preprocessing will perform preliminary data cleaning in the sensor's local computing unit to eliminate obviously abnormal data points. The data will then be uploaded to the central server through 5G wireless communication technology.

[0082] Access satellite remote sensing data, use Landsat-8, MODIS, and Sentinel-2 satellite image data to monitor large-scale pollution distribution, and combine spectral analysis to extract water pollution and air pollution data, including: suspended matter in water bodies, changes in vegetation cover, and industrial emission concentrations.

[0083] Wind speed, wind direction, and precipitation are collected through the meteorological center API, the atmospheric boundary layer height is obtained through lidar, and factory pollutant emission data, including chimney exhaust gas and sewage emissions, are collected and compared with emission standards to identify companies that violate regulations and exceed standards.

[0084] In some specific implementations, data format standardization is performed based on the obtained raw data, data anomaly detection is performed based on the raw data after data format standardization, and missing value filling is performed based on the raw data after data anomaly detection, specifically including: The data format was standardized based on the obtained raw data.

[0085] A timestamp synchronization algorithm is used to unify the time base of all data. By setting a master clock server, all data acquisition devices request time synchronization from the master clock server, calculate the time deviation of each sensor device, and adjust the time of all devices.

[0086] The geographic coordinates of different data sources may use different coordinate systems and need to be uniformly converted to the WGS84 coordinate system using the ellipsoid coordinate conversion formula.

[0087] Data anomaly detection is performed based on the raw data after data format standardization.

[0088] The local anomaly factor algorithm is used to detect abnormal data points. The local anomaly factor algorithm is used to detect abnormal points in the data set, calculate the local anomaly factor of each data point, and calculate the reachable distance, local density and LOF value.

[0089] Wavelet transform is used for data denoising, discrete wavelet transform is used for noise removal, Daubechies wavelet is selected as the basis function, soft threshold denoising is performed, and the denoised data is restored by inverse wavelet transform.

[0090] Fill in the missing values based on the original data after data anomaly detection.

[0091] For small-scale data loss, linear interpolation is used; for large-scale data loss, K-nearest neighbor imputation is used.

[0092] In some specific embodiments, based on the original data after missing value filling, atmospheric pollution diffusion simulation is carried out; based on the original data after atmospheric pollution diffusion simulation, water pollution diffusion modeling is carried out; based on the original data after water pollution diffusion modeling, soil pollution diffusion analysis is carried out, specifically including: Carry out atmospheric pollution diffusion simulation based on the original data after missing value filling.

[0093] Atmospheric pollution diffusion simulation includes: Gaussian plume model calculation and diffusion coefficient calculation.

[0094] The Gaussian plume model calculation is specifically: , In the formula: is the pollutant concentration at the spatial position , with the unit of , is the emission of the pollution source, with the unit of , is the wind speed, with the unit of , is the horizontal diffusion coefficient, with the unit of m, is the vertical diffusion coefficient, with the unit of m, is the horizontal offset distance of the pollution source, with the unit of m, is the vertical offset distance of the pollution source, with the unit of m, is the exponential function, that is, the calculation in the form of .

[0095] The diffusion coefficient calculated by the diffusion coefficient calculation depends on the atmospheric stability and the pollutant propagation distance, specifically: , In the formula: is the horizontal diffusion coefficient, with the unit of m, is the vertical diffusion coefficient, with the unit of m, is the pollutant propagation distance, with the unit of m, is the empirical coefficient, which depends on the atmospheric stability.

[0096] Calculate the pollutant diffusion path by combining the original data, collect the wind speed, wind direction, and air temperature through the meteorological API to obtain the diffusion direction and speed of the pollutant in the air, monitor the atmospheric boundary layer height through lidar to obtain the vertical diffusion range of the pollutant, and combine the original data to correct the diffusion behavior of the pollutant in complex terrains.

[0097] Based on the original data after atmospheric pollution diffusion simulation, water pollution diffusion modeling is carried out.

[0098] Water pollution diffusion modeling calculates the diffusion path of pollutants in water through a hydrodynamic model and pollutant degradation modeling.

[0099] The hydrodynamic model mainly calculates the hydrodynamic transport process of pollutants, including the diffusion, convection and sedimentation of pollutants in water. The control equation for the change of pollutant concentration with time is specifically: , In the formula: is the pollutant concentration, with the unit of , is the time, with the unit of s, are the velocities of the water flow in the directions respectively, with the unit of , is the spatial coordinate of the location of the pollutant, with the unit of m, are the diffusion coefficients of the water body in the directions respectively, with the unit of , is the degradation rate of the pollutant, with the unit of .

[0100] The QUAL2K water quality model is used to analyze the physical and chemical processes such as the degradation, sedimentation, evaporation, and oxidation of pollutants. The change equation of pollutant concentration with time is specifically: , In the formula: is the pollutant concentration, with the unit of , is the time, with the unit of s, is the degradation rate of the pollutant, with the unit of , is the external pollutant input.

[0101] The pollutant degradation rate , and the calculation formula is: , In the formula: is the degradation rate constant at standard temperature, with the unit of , is the activation energy, with the unit of , is the gas constant, is the water temperature, with the unit of K.

[0102] Collect the water flow velocity, direction, and water level through a hydrological monitoring station, calculate the convective transport process of pollutants, monitor the eutrophication of lake water bodies through remote sensing, evaluate the cumulative risk of pollutants, and calculate the scour and runoff paths of pollutants in combination with the original data.

[0103] Conduct soil pollution diffusion analysis based on the original data after water pollution diffusion modeling.

[0104] For soil pollution diffusion analysis, the Fick diffusion model is used to calculate the diffusion of pollutants in the soil. Specifically: , In the formula: is the pollutant concentration, with the unit of , is the time, with the unit of s, is the soil depth, with the unit of m, is the diffusion coefficient, with the unit of .

[0105] Monitor the soil moisture content through a soil moisture sensor, calculate the infiltration rate of pollutants, calculate the flow path of pollutants with groundwater through groundwater monitoring, evaluate the risk of groundwater pollution, and evaluate the chemical reaction characteristics of pollutants in combination with the original data.

[0106] In some specific embodiments, based on the original data after soil pollution diffusion analysis, source classification of pollution sources is carried out. Based on the source classification of pollution sources, the pollution contribution rate is calculated. Based on the calculation of the pollution contribution rate, blockchain pollution tracing is carried out. Specifically, it includes: Carry out source classification of pollution sources based on the original data after soil pollution diffusion analysis.

[0107] The purpose of source classification of pollution sources is to classify pollution data into different pollution source categories, and the random forest and XGBoost models are used for source classification of pollution sources.

[0108] The random forest model conducts source classification of pollution sources by integrating multiple decision trees. The XGBoost model uses gradient boosting trees and is improved based on the random forest.

[0109] Extract features from the pollution data, construct input variables for machine learning classification, including PM2.5, PM10, SO2, NO2, CO, O3, VOC, COD, BOD, ammonia nitrogen, total phosphorus, and heavy metals, the geographical location, terrain, and adjacent water body information of the pollution source, the daily and seasonal variations of pollutants, wind speed, wind direction, temperature, humidity, and precipitation.

[0110] Train the random forest classification model. Specifically: , In the formula: is the final pollution source classification result, is the number of decision trees, is the prediction result of the

[0111] The random forest training process normalizes the pollutant data, divides the original data according to the ratio of 80% training set + 20% test set, sets the number of decision trees to 300, uses information gain or Gini coefficient as the feature selection criterion, evaluates the model accuracy using the test set, calculates the confusion matrix, recall rate, and F1-score, and calculates according to the probability distribution of the pollution category output by the model.

[0112] Calculate the pollution contribution rate based on the pollution source classification.

[0113] Use the Bayesian regression model for regression analysis, specifically: , where: is the total pollutant concentration, with the unit of or , is the emission of the or type of pollution source, with the unit of is the pollution contribution factor of the type of pollution source,

[0114] Bayesian regression calculates the pollution contribution factor through maximum a posteriori estimation, specifically: , where: is the posterior probability of the parameter after the given data , is the likelihood function, indicating the probability of the pollutant concentration given the value, is the prior distribution, usually assumed to follow a normal distribution , is the marginal probability, used to normalize the probability distribution.

[0115] Calculate the pollution source contribution rate, specifically: , where: is the contribution rate of the type of pollution source to the pollutant, with the unit of , is the pollution contribution factor, is the emission of the type of pollution source, and

[0116] Based on the calculated pollution contribution rate, blockchain pollution traceability is carried out.

[0117] Pollution tracking data is stored through the blockchain. The block header contains a timestamp, the hash value of the previous block, and the Merkle root. The transaction data includes pollutant concentration, pollution source classification, pollution contribution rate, and pollution diffusion path. Hash pointers are used to store the hash values of pollution data.

[0118] Calculate the pollutant concentration and pollution contribution rate, generate the hash value of the pollution source data, generate a block and broadcast it to the blockchain network to confirm the transaction through the consensus mechanism, and store the pollution tracking data.

[0119] In some specific embodiments, a GIS pollution diffusion heat map is generated based on the original data. Intelligent pollution warning is carried out based on the original data after the generation of the GIS pollution diffusion heat map, and a warning threshold is set, specifically including: Generate a GIS pollution diffusion heat map based on the original data.

[0120] Combine ArcGIS and Leaflet.js to construct a pollution map. Perform data projection transformation using the WGS84 coordinate system. Calculate the distribution of pollutant concentration in the geographical space through Kriging interpolation. Generate a pollution heat map using Leaflet.js and use gradient colors, including red - yellow - green, to represent the pollution concentration.

[0121] Use time - series animation to display the pollution diffusion trend. Combine wind speed, wind direction, and terrain data to dynamically adjust the pollution diffusion direction. Use the Dijkstra shortest path algorithm to calculate the diffusion path of pollutants from the pollution source to the affected area. Combine the Lagrangian particle tracking model to simulate the movement trajectory of pollutants in the air / water body.

[0122] Perform intelligent pollution warning based on the original data after the generation of the GIS pollution diffusion heat map.

[0123] Use a deep - learning model for pollution trend prediction and combine the original data for pollution hot - spot tracking.

[0124] Use LSTM to predict the pollution trend. The input of the prediction model is the original data, wind speed, wind direction, temperature, precipitation, and seasonal factors, specifically: , , , , , , In the formula: is the activation value of the forget gate, determining how much of the memory information from the previous time step is retained at the current moment. is the Sigmoid activation function, with an output range between indicating how much information is retained. is the weight matrix of the forget gate, used to calculate the weight parameters of . is the hidden state at the previous moment, containing past information. is the input feature vector at the current moment, containing pollution monitoring data or relevant environmental variables. is the bias term of the forget gate, used to optimize the calculation. is the activation value of the input gate, determining to what extent the current input information is memorized into the cell state. is the Sigmoid activation function, ensuring an output range between . is the weight matrix of the input gate, used to calculate the weight parameters of . is the hidden state at the previous moment, containing past pollution trend information. is the input data at the current moment, such as pollutant concentration, wind speed, wind direction, etc. is the bias term of the input gate, used to optimize the calculation. is the candidate cell state, representing new candidate information that will be used to update the cell state. is the hyperbolic tangent activation function, mapping the input to between to make the information have non-linear characteristics. is the weight matrix of the candidate cell state, used to calculate the weight parameters of . is the hidden state at the previous moment. is the input data at the current moment. is the bias term of the candidate cell state, used to optimize the calculation. is the activation value of the output gate, determining how the cell state affects the hidden state at the current moment. , is the Sigmoid activation function, with an output range between . is the weight matrix of the output gate, used to calculate the weight parameters of . is the hidden state at the previous moment, is the input data at the current moment, is the bias term of the output gate for optimizing the calculation, is the hidden state at the current moment, that is, the short-term memory of the LSTM, and is used as the final output, is the activation value of the output gate, which determines the weight of the output information, is the non-linear transformation of the cell state to ensure that the information is mapped to between, is the cell state at the current moment, which contains long-term memory information.

[0125] The Adam optimizer is used for parameter update, the learning rate is set to 0.001, the mean squared error loss function is adopted, the training set ratio is 80%, the test set is 20%, and 50 rounds of training are carried out to predict the pollutant concentration. If the predicted pollutant concentration exceeds the set threshold, an early warning will be triggered.

[0126] Drones are used to detect pollution hotspots in real time. The optimal cruise path is calculated through the A* search algorithm. The drones are equipped with spectral imagers to monitor industrial waste gas, oil spills and river pollution in real time. The YOLOv5 is used to identify pollution sources. When a pollution source is identified, the GPS coordinates of the pollution source are recorded.

[0127] Set the warning thresholds for air pollutants. The warning threshold for PM2.5 is 75 µg / m³. Exceeding this value triggers a light pollution warning. If it exceeds 115 µg / m³, a moderate pollution warning is triggered. If it exceeds 150 µg / m³, a severe pollution warning is triggered. The warning threshold for PM10 is 150 µg / m³. Exceeding this value triggers a light pollution warning. If it exceeds 250 µg / m³, a moderate pollution warning is triggered. If it exceeds 350 µg / m³, a severe pollution warning is triggered. The warning threshold for SO2 is 150 µg / m³. Exceeding this value triggers a light pollution warning. If it exceeds 500 µg / m³, a moderate pollution warning is triggered. If it exceeds 800 µg / m³, a severe pollution warning is triggered. The warning threshold for NO2 is 100 µg / m³. Exceeding this value triggers a light pollution warning. If it exceeds 200 µg / m³, a moderate pollution warning is triggered. If it exceeds 300 µg / m³, a severe pollution warning is triggered. The warning threshold for CO is 5000 µg / m³ (5 mg / m³). Exceeding this value triggers a light pollution warning. If it exceeds 10000 µg / m³ (10 mg / m³), a moderate pollution warning is triggered. If it exceeds 20000 µg / m³ (20 mg / m³), a severe pollution warning is triggered. The warning threshold for O3 is 160 µg / m³. Exceeding this value triggers a light pollution warning. If it exceeds 200 µg / m³, a moderate pollution warning is triggered. If it exceeds 400 µg / m³, a severe pollution warning is triggered. The warning threshold for VOC is 200 µg / m³. Exceeding this value triggers a light pollution warning. If it exceeds 300 µg / m³, a moderate pollution warning is triggered. If it exceeds 500 µg / m³, a severe pollution warning is triggered. Set the warning thresholds for water pollutants. The warning threshold for chemical oxygen demand is 15 mg / L. Exceeding this value triggers a light pollution warning. If it exceeds 20 mg / L, a moderate pollution warning is triggered. If it exceeds 30 mg / L, a severe pollution warning is triggered. The warning threshold for biochemical oxygen demand is 3 mg / L. Exceeding this value triggers a light pollution warning. If it exceeds 6 mg / L, a moderate pollution warning is triggered. If it exceeds 10 mg / L, a severe pollution warning is triggered. The warning threshold for ammonia nitrogen is 0.5 mg / L. Exceeding this value triggers a light pollution warning. If it exceeds 1.0 mg / L, a moderate pollution warning is triggered. If it exceeds 2.0 mg / L, a severe pollution warning is triggered. The warning threshold for total phosphorus is 0.02 mg / L. Exceeding this value triggers a light pollution warning. If it exceeds 0.05 mg / L, a moderate pollution warning is triggered. If it exceeds 0.1 mg / L, a severe pollution warning is triggered. The warning threshold for total nitrogen is 0.2 mg / L. Exceeding this value triggers a light pollution warning. If it exceeds 0.5 mg / L, a moderate pollution warning is triggered. If it exceeds 1.0 mg / L, a severe pollution warning is triggered. The warning threshold for heavy metals is 0.005 mg / L. Exceeding this value triggers a light pollution warning. If it exceeds 0.01 mg / L, a moderate pollution warning is triggered. If it exceeds 0.02 mg / L, a severe pollution warning is triggered. Set the warning thresholds for soil pollutants. The warning threshold for lead is 80 mg / kg. Exceeding this value triggers a mild pollution warning. If it exceeds 200 mg / kg, a moderate pollution warning is triggered. If it exceeds 400 mg / kg, a severe pollution warning is triggered. The warning threshold for cadmium is 0.3 mg / kg. Exceeding this value triggers a mild pollution warning. If it exceeds 1.0 mg / kg, a moderate pollution warning is triggered. If it exceeds 3.0 mg / kg, a severe pollution warning is triggered. The warning threshold for mercury is 0.3 mg / kg. Exceeding this value triggers a mild pollution warning. If it exceeds 1.0 mg / kg, a moderate pollution warning is triggered. If it exceeds 3.0 mg / kg, a severe pollution warning is triggered. The warning threshold for arsenic is 30 mg / kg. Exceeding this value triggers a mild pollution warning. If it exceeds 50 mg / kg, a moderate pollution warning is triggered. If it exceeds 100 mg / kg, a severe pollution warning is triggered. The warning threshold for hexavalent chromium is 150 mg / kg. Exceeding this value triggers a mild pollution warning. If it exceeds 300 mg / kg, a moderate pollution warning is triggered. If it exceeds 500 mg / kg, a severe pollution warning is triggered. The warning threshold for petroleum hydrocarbons is 100 mg / kg. Exceeding this value triggers a mild pollution warning. If it exceeds 500 mg / kg, a moderate pollution warning is triggered. If it exceeds 1000 mg / kg, a severe pollution warning is triggered. For the mild pollution warning, send text messages, APP push notifications, and emails to relevant management personnel, display the polluted area on the GIS visualization platform, and mark it with a yellow warning. For the moderate pollution warning, send text messages, APP push notifications, and emails to enterprises, environmental protection departments, and government supervision units, require them to take pollution control measures, mark the orange warning area on the GIS platform, and provide pollution trend prediction. The light alarm is orange. For the severe pollution warning, send text messages + phone notifications, show a red warning on the GIS platform, and the light alarm is red and there is a buzzer alarm.

[0128] In the above content, in actual application, first, based on the multi-source data fusion technology, the invention collects and preprocesses pollution data. Selects optical scattering dust sensors, electrochemical gas sensors, and lidar sensors to collect air pollution data, including PM2.5, PM10, SO2, NO2, CO, O3, and volatile organic compounds. The data collection period is set from 10 seconds to 5 minutes. Obtains water pollution data, including chemical oxygen demand, biochemical oxygen demand, ammonia nitrogen, total phosphorus, total nitrogen, and heavy metals, through ultraviolet spectrum water quality analyzers, electrode sensors, and ion-selective electrode sensors. The data collection period is set from 30 seconds to 10 minutes. Uses X-ray fluorescence spectrometers, ion chromatographs, and infrared spectrometers to monitor soil pollution data, including heavy metal content, organic pollutants, and soil pH. The data collection period is set from 1 hour to 12 hours. All collection devices are connected to the Internet of Things cloud platform, and real-time data transmission is achieved through 5G wireless communication technology to ensure the continuity and availability of data.

[0129] Next, perform data format standardization, abnormal data detection, and missing value filling on the collected pollution data. Use the timestamp synchronization algorithm to unify the time base of all data, and use the coordinate transformation method to ensure the accuracy of geographical location information. Detect abnormal data points using the Local Outlier Factor algorithm, eliminate invalid data caused by sensor failures, data interference, etc., and perform data denoising through wavelet transform and mean filtering. For the missing values in the pollution data, use linear interpolation to repair small-scale data missing, and use the K-nearest neighbor algorithm to complete large-scale missing data, improving data integrity and accuracy, thereby ensuring the reliability and consistency of the pollution data.

[0130] Then, construct a pollution diffusion model to simulate the propagation paths of pollutants in air, water, and soil. In the simulation of air pollution diffusion, use the Gaussian plume model to calculate the pollutant concentration distribution, and combine wind speed, wind direction, and atmospheric stability data to correct the pollution diffusion path. In the modeling of water pollution diffusion, combine the hydrodynamic model and the water quality model to calculate the diffusion process of pollutants in the water flow, and analyze the degradation and deposition of pollutants. In the simulation of soil pollution diffusion, use the Fick diffusion model to calculate the diffusion of pollutants in the soil, and combine soil moisture sensor and groundwater monitoring data to evaluate the penetration rate of pollutants and their impact on groundwater. By accurately simulating the diffusion of pollutants in different media, improve the accuracy of pollution source tracing.

[0131] At the same time, based on the results of the pollution diffusion modeling, conduct pollution source identification and pollution contribution rate calculation. Use random forest and XGBoost to train a pollution source classification model, classify pollution sources into industrial pollution sources, agricultural pollution sources, transportation pollution sources, and natural pollution sources, and use Bayesian regression method to calculate the contribution rate of different pollution sources to pollutants, realizing the quantitative attribution of pollution responsibility. Through pollution source identification and contribution rate calculation, the main pollution sources can be accurately locked, and targeted supervision and treatment measures can be implemented for key pollution sources based on the results to improve the scientificity and accuracy of pollution prevention and control.

[0132] Finally, construct a pollution visualization and intelligent early warning system. Based on GIS map technology, combine ArcGIS and Leaflet.js to generate a pollution diffusion heat map, visually display the spatial distribution of pollutants and the location of pollution sources, use the Long Short-Term Memory network to predict the future pollution change trend, and combine drone monitoring to detect pollution hotspots in real time. When the pollution exceeds the standard, the system will trigger a three-level early warning mechanism: mild pollution triggers APP notifications and GIS prompts, moderate pollution triggers text messages and orange light alarms, and severe pollution triggers red alarms, beeping sounds, and emergency control measures. At the same time, combine blockchain technology to store pollution tracking data to ensure the security, immutability, and traceability of pollution source tracing data, providing a scientific basis for ecological environment governance.

Claims

1. An ecological environmental pollution source tracing method based on data analysis, characterized in that, It includes the following steps: Select equipment types and obtain original data. Based on the selected equipment types, conduct equipment layout. Based on the equipment layout, conduct multi-source data access; Based on the obtained original data, conduct data format standardization. Based on the originally standardized data, conduct data anomaly detection. Based on the originally detected data with anomalies, fill in missing values; Based on the originally filled data with missing values, conduct air pollution diffusion simulation. Based on the originally simulated data of air pollution diffusion, conduct water pollution diffusion modeling. Based on the originally modeled data of water pollution diffusion, conduct soil pollution diffusion analysis; Based on the originally analyzed data of soil pollution diffusion, conduct pollution source classification. Based on the classified pollution sources, calculate pollution contribution rates. Based on the calculated pollution contribution rates, conduct blockchain pollution traceability verification; Based on the original data, generate a GIS pollution diffusion heat map. Based on the originally generated GIS pollution diffusion heat map, conduct intelligent pollution warning and set warning thresholds.

2. The method for tracing the source of ecological environmental pollution sources based on data analysis according to claim 1, characterized in that, Select equipment types and obtain original data. Based on the selected equipment types, conduct equipment layout. Based on the equipment layout, conduct multi-source data access, specifically including: Select equipment types and obtain original data; Through light scattering dust sensors, electrochemical gas sensors, and lidar sensors, obtain PM2.5, PM10, SO2, NO2, CO, O3, and volatile organic compounds. Set the sampling periods of the light scattering dust sensors, electrochemical gas sensors, and lidar sensors to collect once every 10 seconds to 5 minutes; Through ultraviolet spectrum water quality analyzers, electrode sensors, and ion-selective electrode sensors, obtain chemical oxygen demand, biochemical oxygen demand, ammonia nitrogen, total phosphorus, total nitrogen, and heavy metals. Set the collection periods of the ultraviolet spectrum water quality analyzers, electrode sensors, and ion-selective electrode sensors to collect once every 30 seconds to 10 minutes; Through X-ray fluorescence spectrometers, ion chromatographs, and infrared spectrometers, obtain heavy metals, organic pollutants, and soil acidity. Set the sampling periods of the X-ray fluorescence spectrometers, ion chromatographs, and infrared spectrometers to collect once every 1 hour to 12 hours; Based on the selected equipment types, conduct equipment layout; Equipment layout includes: air pollution sensor layout, water pollution sensor layout, and soil pollution sensor layout; For air pollution sensor layout, install a high-density sensor grid in high-pollution areas at intervals of 500m - 1km, install more accurate monitoring equipment in environmentally sensitive areas at intervals of 1km - 2km, and use background monitoring points to obtain baseline data, with 2 - 5 long-term monitoring points arranged; For water pollution sensor layout, monitor the background value of inflowing pollutants in the upper reaches of the water area, focus on monitoring industrial, agricultural, and domestic sewage discharges near sewage outlets, monitor the mixing of different water pollutions at river confluences, and monitor the eutrophication degree of water bodies at the central points of reservoirs and lakes; For soil pollution sensor layout, focus on monitoring agricultural pollution in farmland areas, monitor the pollution of industrial wastewater and waste residues to the soil around industrial areas, and monitor the accumulation of harmful substances in the soil at landfills; Access multi-source data based on equipment deployment; The deployed sensors will be connected to the IoT cloud platform. The sensor data collection will collect pollutant concentration data according to the set period, and the edge computing preprocessing will perform preliminary data cleaning in the sensor local computing unit to remove obvious abnormal data points, and upload the data to the central server through 5G wireless communication technology; Access satellite remote sensing data, use Landsat-8, MODIS, and Sentinel-2 satellite image data to monitor large-scale pollution distribution, and combine spectral analysis to extract water pollution and air pollution data, including: suspended matter in water bodies, changes in vegetation cover, and industrial emission concentrations; Wind speed, wind direction, and precipitation are collected through the meteorological center API, the atmospheric boundary layer height is obtained through lidar, and factory pollutant emission data, including chimney exhaust gas and sewage emissions, are collected and compared with emission standards to identify companies that violate regulations and exceed standards.

3. The method for tracing the sources of ecological environmental pollution based on data analysis according to claim 1, characterized in that, The data format is standardized based on the obtained original data, data anomaly detection is performed based on the original data after data format standardization, and missing value filling is performed based on the original data after data anomaly detection, specifically including: Standardize the data format based on the obtained raw data; The time stamp synchronization algorithm is used to unify the time base of all data. By setting a master clock server, all data acquisition devices request time synchronization from the master clock server, calculate the time deviation of each sensor device, and adjust the time of all devices; The geographic coordinates of different data sources may use different coordinate systems and need to be uniformly converted to the WGS84 coordinate system using the ellipsoid coordinate conversion formula; Perform data anomaly detection based on the raw data after data format standardization; The local anomaly factor algorithm is used to detect abnormal data points. The local anomaly factor algorithm is used to detect abnormal points in the data set, calculate the local anomaly factor of each data point, and calculate the reachable distance, local density and LOF value; Wavelet transform is used for data denoising, discrete wavelet transform is used for noise removal, Daubechies wavelet is selected as the basis function, soft threshold denoising is performed, and the denoised data is restored by inverse wavelet transform; Fill in missing values ​​based on the original data after data anomaly detection; Linear interpolation is used for small-scale data missing, and K-nearest neighbor completion is used for large-scale data missing.

4. The method for tracing the source of ecological environmental pollution sources based on data analysis according to claim 1, wherein Based on the original data after filling in the missing values, the atmospheric pollution diffusion simulation is carried out, the water pollution diffusion modeling is carried out based on the original data after the atmospheric pollution diffusion simulation, and the soil pollution diffusion analysis is carried out based on the original data after the water pollution diffusion modeling. Specifically, it includes: Simulate the diffusion of atmospheric pollution based on the original data after filling in the missing values; Atmospheric pollution diffusion simulation includes: Gaussian plume model calculation and diffusion coefficient calculation; Gaussian plume model calculation, specifically: , In the formula: is the pollutant concentration at the spatial position in the unit of , is the emission of the pollution source in the unit of , is the wind speed in the unit of , is the lateral diffusion coefficient in the unit of m, is the vertical diffusion coefficient in the unit of m, is the lateral offset distance of the pollution source in the unit of m, is the vertical offset distance of the pollution source in the unit of m, is the exponential function, that is in the form of calculation; Diffusion coefficientsThe diffusion coefficients calculated depend on the atmospheric stability and the distance the pollutants travel; Calculate the pollutant diffusion path by combining the original data. Collect the wind speed, wind direction, and air temperature through the meteorological API to obtain the diffusion direction and speed of pollutants in the air. Monitor the atmospheric boundary layer height through lidar to obtain the vertical diffusion range of pollutants. Combine the original data to correct the diffusion behavior of pollutants in complex terrains; Conduct water pollution diffusion modeling based on the original data after atmospheric pollution diffusion simulation; The water pollution diffusion modeling calculates the diffusion path of pollutants in water through a hydrodynamic model and pollutant degradation modeling; The hydrodynamic model mainly calculates the hydrodynamic transport process of pollutants, including the diffusion, convection, and deposition of pollutants in water; The QUAL2K water quality model is used to analyze the physical and chemical processes such as the degradation, deposition, evaporation, and oxidation of pollutants; Collect the water flow velocity, flow direction, and water level through a hydrological monitoring station to calculate the convective transport process of pollutants. Monitor the eutrophication of lake water through remote sensing to evaluate the pollutant accumulation risk. Combine the original data to calculate the scouring and runoff paths of pollutants; Conduct soil pollution diffusion analysis based on the original data after water pollution diffusion modeling; The soil pollution diffusion analysis uses the Fick diffusion model to calculate the diffusion of pollutants in soil; Monitor the soil moisture content through a soil moisture sensor to calculate the infiltration rate of pollutants. Calculate the flow path of pollutants with groundwater through groundwater monitoring to evaluate the groundwater pollution risk. Combine the original data to evaluate the chemical reaction characteristics of pollutants; 5. The method for tracing the source of ecological environmental pollution sources based on data analysis according to claim 1, characterized in that, Conduct pollution source classification based on the original data after soil pollution diffusion analysis. Calculate the pollution contribution rate based on the pollution source classification. Conduct blockchain pollution tracing based on the calculation of the pollution contribution rate, specifically including: Conduct pollution source classification based on the original data after soil pollution diffusion analysis; The purpose of pollution source classification is to classify pollution data into different pollution source categories. Use the random forest and XGBoost models for pollution source classification; The random forest model conducts pollution source classification by integrating multiple decision trees. The XGBoost model uses gradient boosting trees and is improved based on the random forest; Extract features from the pollution data to construct input variables for machine learning classification, including PM2.5, PM10, SO2, NO2, CO, O3, VOC, COD, BOD, ammonia nitrogen, total phosphorus, and heavy metals, the geographical location, terrain, and adjacent water body information of the pollution source, the daily and seasonal variations of pollutants, wind speed, wind direction, temperature, humidity, and precipitation; Train the random forest classification model, specifically: , In the formula: is the final pollution source classification result, is the number of decision trees, is the prediction result of the The random forest training process normalizes the pollutant data, divides the original data according to the ratio of 80% training set + 20% test set, sets the number of decision trees to 300, uses information gain or Gini coefficient as the feature selection criterion, uses the test set to evaluate the model accuracy, calculates the confusion matrix, recall rate, and F1-score, and according to the probability distribution of the pollution category output by the model; Calculate the pollution contribution rate based on the pollution source classification; Conduct regression analysis using the Bayesian regression model, specifically: , where: is the total pollutant concentration, with the unit of or , is the emission of the th type of pollution source, with the unit of or , is the pollution contribution factor of the th type of pollution source, is the error term, representing unobservable noise factors; The Bayesian regression calculates the pollution contribution factor through maximum a posteriori estimation; Perform blockchain pollution traceability verification based on the calculated pollution contribution rate; Store pollution tracking data through the blockchain. The block header contains a timestamp, the hash value of the previous block, and the Merkle root. The transaction data includes pollutant concentration, pollution source classification, pollution contribution rate, and pollution diffusion path. The hash pointer is used to store the hash value of the pollution data; Calculate the pollutant concentration and pollution contribution rate, generate the hash value of the pollution source data, generate a block and broadcast it to the blockchain network to confirm the transaction through the consensus mechanism, and store the pollution tracking data.

6. The method for tracing the source of ecological environmental pollution sources based on data analysis according to claim 1, wherein, Generate a GIS pollution diffusion heat map based on the original data, perform intelligent pollution warning based on the original data after the generation of the GIS pollution diffusion heat map, and set the warning threshold, specifically including: Generate a GIS pollution diffusion heat map based on the original data; Construct a pollution map by combining ArcGIS and Leaflet.js, perform data projection transformation using the WGS84 coordinate system, calculate the distribution of pollutant concentration in the geospatial area through Kriging interpolation method, generate a pollution heat map using Leaflet.js, and use gradient colors, including red - yellow - green, to represent the pollution concentration; Use time - series animation to display the pollution diffusion trend, combine wind speed, wind direction, and terrain data to dynamically adjust the pollution diffusion direction, use the Dijkstra shortest path algorithm to calculate the diffusion path of pollutants from the pollution source to the affected area, and combine the Lagrangian particle tracking model to simulate the movement trajectory of pollutants in the air / water body; Perform intelligent pollution warning based on the original data after the generation of the GIS pollution diffusion heat map; Use a deep - learning model to predict the pollution trend and combine the original data to track pollution hotspots; Use LSTM to predict the pollution trend. The input of the prediction model is the original data, wind speed, wind direction, temperature, precipitation, and seasonal factors; Use the Adam optimizer to update the parameters, set the learning rate to 0.001, use the mean squared error loss function, with an 80% training set ratio and 20% test set, perform 50 rounds of training, predict the pollutant concentration, and trigger a warning if the predicted pollutant concentration exceeds the set threshold; Use a drone to detect pollution hotspot areas in real - time, calculate the optimal cruise path through the A* search algorithm. The drone is equipped with a spectral imager to monitor industrial waste gas, oil spills, and river pollution in real - time. Use YOLOv5 to identify pollution sources, and record the GPS coordinates of the pollution sources when they are identified.

Citation Information

Cited By

  • Dynamic intelligent monitoring method and system for groundwater pollution

    CN120850792A

  • New pollutant multi-medium migration path intelligent tracing method and system

    CN120869892A

  • Water quality antibiotic tracing method and system based on data analysis

    CN120892686A

  • Cooperative supervision method and system for ecological environment of water area

    CN120952270A

  • Data integration method for intelligent ecological environment platform

    CN121255911A