Environment pollution detection data processing system and method based on Internet of Things

By constructing an environmental pollution data chain through data collection, preprocessing, and a 3D fusion framework, pollution sources and transmission paths are identified, solving the problems of heterogeneity and single-dimensional analysis of pollution detection data in the Internet of Things environment, and realizing efficient pollution source tracing and control.

CN120930935AInactive Publication Date: 2025-11-11GUANGZHOU DELONG ENVIRONMENTAL TESTING TECH CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202511081153.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-11-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The heterogeneity, unstable quality, and lack of standardization of pollution detection data in the Internet of Things environment, as well as the reliance of existing pollution source tracing methods on single-dimensional analysis, lead to low efficiency in pollution source tracing and governance.

Method used

By collecting data from IoT devices, preprocessing and standardizing it, and combining it with a three-dimensional fusion framework of time, space and logic, an environmental pollution data chain is constructed. Dynamic time warping algorithm, inverse distance weight interpolation and association rule algorithm are used to identify pollution sources and propagation paths.

Benefits of technology

It has achieved comprehensiveness and accuracy in pollution characteristic analysis, improved the precision of pollution source tracing, and ensured the reliability and traceability of source tracing conclusions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930935A_ABST
    Figure CN120930935A_ABST
Patent Text Reader

Abstract

The invention discloses an environment pollution detection data processing system and method based on the Internet of Things, relates to the technical field of data analysis and evidence tracing, and aims to solve the problems of data isomerism, time-space correlation analysis splitting and insufficient propagation path recognition precision in the prior art. According to the system and the method, multi-source information such as Internet of Things monitoring data and environment management data is collected, the time trend, the incidence relation and the spatial distribution characteristics of pollutants are extracted after preprocessing, an environment pollution data chain with time, space and logic three-dimensional fusion is constructed, then pollution sources are positioned, a propagation path and an influence range are analyzed, and a real-time environment pollution analysis result is obtained. And a result is visually presented. According to the invention, the precision and reliability of pollution traceability are improved, and a systematic solution is provided for precise treatment of environmental pollution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data analysis and evidence tracing technology, specifically to an environmental pollution detection data processing system and method based on the Internet of Things. Background Technology

[0002] With the deep application of IoT technology in environmental monitoring, various environmental monitoring equipment, sensor networks, and supporting systems have formed a massive multi-source data acquisition capability, providing a data foundation for real-time perception and precise control of environmental pollution. However, in the actual processing of environmental pollution detection data, problems such as inconsistent data quality, fragmented spatiotemporal correlation analysis, and insufficient accuracy in propagation path identification still restrict the efficiency of pollution source tracing and control, specifically manifested as follows: Pollution detection data in the Internet of Things (IoT) environment exhibits significant heterogeneity: The data comes from diverse sources, including real-time sensor monitoring values, meteorological data, geographic information, pollution source emission records, and equipment operation logs, with significant differences in format and units. Unstable quality: Sensor failures can easily lead to data jumps, network latency may cause timestamp errors, some monitoring points have missing data due to lack of maintenance, and traditional cleaning methods are difficult to adapt to complex environmental interference. Lack of standardization: The monitoring data formats of different equipment manufacturers are not uniform, and the units of numerical indicators are inconsistent, making it difficult to directly integrate and analyze multi-source data; Existing pollution source tracing methods mostly rely on single-dimensional analysis, making it difficult to capture the dynamic characteristics of pollution events. Isolated time dimension: Concentration changes are analyzed only through simple time series models, without linking the temporal coupling relationship between equipment operating status and pollution events; Spatial dimension fragmentation: Although pollution hotspot maps can be generated through interpolation, the location of pollution sources, topographic barriers, and meteorological conditions are not taken into account, resulting in distortion of spatial correlation strength calculation. Lack of logical connection: Ignoring the co-occurrence patterns among pollutant indicators and the causal relationship of influencing factors makes it difficult to construct a complete link from pollution source to transmission path to monitoring point.

[0003] To address the aforementioned problems, this invention provides an environmental pollution detection data processing system and method based on the Internet of Things (IoT) to solve these issues. Summary of the Invention

[0004] The purpose of this invention is to provide an environmental pollution detection data processing system and method based on the Internet of Things to solve the problems raised in the prior art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: A method for processing environmental pollution detection data based on the Internet of Things includes the following steps: S1. Collect raw data, obtain environmental monitoring data, meteorological data, geographic information data, and pollution source-related data from IoT environmental monitoring equipment, sensor networks, meteorological monitoring instruments, geographic information system databases, and pollution source information management systems, and at the same time collect data from environmental management agencies and pollution risk factors in the operation of IoT equipment; S2. Preprocess the collected raw data, including data cleaning and standardization, to unify the data format and provide usable data for subsequent analysis; S3. Based on standardized data, analyze the trends in environmental pollutant concentrations, the relationships between pollutants, and their spatial distribution characteristics to form pollution information; S4. Integrate pollution information with IoT data in terms of time and space to build an environmental pollution data chain; S5. Based on the constructed environmental pollution data chain, determine the location of the pollution source, the direction of spread, and the scope of impact, and output the processing results; S6. Transform the processing results into a visual format and present it to the user.

[0006] S1 further includes the following steps: S1.1: Collect full-volume structured data from IoT environmental monitoring devices and sensor networks, including structured fields such as device identifier, monitoring time, pollutant type, concentration value, and monitoring point location; specifically: adopt a real-time incremental acquisition method, and use a data interface program written in Python to locate sensors by device ID and obtain real-time monitoring data based on the API interface of the IoT platform. Then, use SQL statements to extract data from the sensor status table, monitoring data table, and device operation and maintenance record table through IoT database views. S1.2: Upload environmental management agency data in the form of a structured form through a data acquisition plugin deployed on the agency's internal IoT platform. Collect risk factors of environmental management agency data and IoT device operation from the environmental management agency data. Then, collect device operation logs in real time by embedding a monitoring program in the IoT gateway. The log fields contain environmental management data and device operation risk factors. Finally, store the above-mentioned risk factors and device operation logs in JSON format on the IoT log server.

[0007] S2 further includes the following steps: S2.1: Perform outlier removal, duplicate record deduplication, and missing value repair on the collected IoT environmental monitoring data, environmental management data, and equipment operation risk factors. For IoT monitoring data, outlier data is detected by setting standard value ranges for pollutant concentrations and combining them with the 3σ principle. Duplicate records are removed using the unique identifier of device ID + timestamp. For environmental management data, a regression model is built based on the monitoring point location using a multiple interpolation method to generate replacement values ​​for missing values. Cleaning rules are executed in stages according to data type. For equipment operation risk factors, duplicate and erroneous records are removed using timestamp continuity verification and sensor parameter threshold detection. Missing values ​​are filled by pattern matching from historical operation logs. S2.2: Transform heterogeneous IoT data into a unified format system. The monitoring time field is parsed into a preset standard format using regular expressions, and the device ID is given a classification prefix to form a standardized code. Numerical indicators such as pollutant concentration are mapped to a unified dimension using the Z-score standardization algorithm. The text description of the monitoring point location is geocoded to generate latitude and longitude feature vectors. Numerical fields in environmental management data are normalized. Equipment operation data is stored according to the standard format specifications in the IoT field. A field format conversion rule library for various types of data is established, and field mapping is performed based on the output data structure of the data acquisition.

[0008] S3 further includes the following steps: S3.1: Based on the standardized data obtained in S2.2, the concentration data is sorted in chronological order to construct a time series. A moving average model is used to filter short-term fluctuations, and a linear regression algorithm is combined to fit the long-term trend of changes, identifying concentration increase, decrease and periodic fluctuation patterns. Outlier correction based on the 3σ principle is performed on the time series data. Finally, the periodic components are decomposed through frequency domain analysis, and finally a pollutant concentration-time trend curve with characteristic parameters is generated. S3.2: Based on standardized data, an association rule algorithm is used to identify strong association rules by generating frequent itemsets and filtering rules that meet the threshold, in order to analyze the co-occurrence patterns among pollutant indicators. At the same time, a pollutant-influencing factor network is constructed with pollutant concentration as nodes and pollution source emission records in equipment operating parameters and environmental management data as edge weights. Then, the betweenness centrality algorithm is used to calculate the node centrality. Specifically, the concentration anomalies in the monitoring data are associated with equipment operating parameter records through key field mapping. Association rule mining and network topology analysis are performed to generate a pollutant association matrix and an importance ranking of influencing factors. The pollutant association matrix records the rule support and confidence, and the importance ranking is arranged in descending order of centrality value. S3.3: Based on standardized data, analyze the spatial distribution characteristics of pollutants, combine the latitude and longitude coordinates of monitoring points, and use spatial interpolation algorithms to generate a spatial distribution heat map of pollutants to identify high concentration areas.

[0009] S4 further includes the following steps: S4.1: In the time dimension, the pollutant concentration trend curve is aligned with the time series of equipment operation logs. A dynamic time warping algorithm is used to calculate the time series similarity, identifying the time coupling points between concentration anomalies and equipment anomalies. The time coupling degree C... t The calculation formula is as follows: ; Where T1 is the pollution event time series, T2 is the data anomaly time series, DTW(·) is the dynamic time warp distance, and C t The value ranges from [0,1]. The larger the value, the stronger the time correlation. len represents the length of the time series, which is the number of data points in the time series. It is used to normalize the Dynamic Time Warped Distance (DTW). S4.2: In the spatial dimension, the latitude and longitude coordinates of the monitoring points are overlaid with the spatial distribution hotspots of pollutants. An inverse distance weighted interpolation method is used to generate a pollution propagation heat map, determining the spatial correlation strength between the monitoring points and pollution hotspots, and the spatial influence weight W. s The calculation formula for W is s =1 / d, where d is the Euclidean distance between the monitoring point and the pollution hotspot, in km; S4.3: In terms of logic, the strongly associated rules generated by the association rule algorithm are used as chains to connect pollution sources, propagation paths, and influencing factors into a directed acyclic graph; S4.4: Establish a time-space-logic ternary association rule base, define time window threshold, spatial distance threshold and logical confidence threshold, and use graph database technology to associate pollution information with IoT raw data nodes. Node attributes include timestamp, geographic coordinates and association strength, with values ​​[0,1]. Finally, a multi-level environmental pollution data chain graph is generated, and each node can be traced back to the original data table fields and feature extraction results.

[0010] S5 further includes the following steps: S5.1: Based on a multi-level environmental pollution data chain diagram, determine the location of pollution sources and extract the temporal coupling degree Ct and spatial influence weight W of all pollution event nodes in the data chain. s and logical association strength W j W j The confidence level is used for confirmation, and then the comprehensive traceability score S is calculated using the following formula: ; Where α, β, and γ are weighting coefficients; Then, for each node, the time difference between its timestamp and the earliest pollution event is checked. Then, the weighted values ​​of spatial distance and logical association are combined to finally determine the node with the largest S value as the pollution source. S5.2: Analyze the direction and scope of pollution propagation. By analyzing the topological structure of the directed acyclic graph of the data chain, and combining the concentration time trend and spatial weights, the propagation weight W is calculated based on the betweenness centrality algorithm, starting from the source node. p The calculation formula is as follows: ; Where Δt is the time interval between adjacent nodes; Then, the data chain graph is topologically sorted, nodes are traversed in chronological order, and path priorities are dynamically adjusted through propagation weights. Next, the shortest path algorithm is used to identify the main propagation path, and a heatmap of the pollution impact range is generated through spatial interpolation, where the boundary of the impact range is determined by W. p The set of nodes with a value of ≥0.2 is determined, and the final output includes the propagation direction arrow, main path identifier, and boundary of the impact range. This output forms a closed loop verification with the spatiotemporal correlation data of the data chain and the concentration trend analysis of pollutant feature extraction. S6 further includes the following steps: S6.1: Based on the processing results, combined with timestamp and spatial coordinate information, the latitude and longitude coordinates of the source node and the timestamp of the pollution event are obtained from the processing results. Then, the pollution impact range is rendered into a heat map using GIS spatial interpolation technology, with the color gradient corresponding to the spatial impact weight W. s The value; the pollution event sequence is displayed in time axis form, with each time point associated with the pollutant concentration trend curve, and the temporal coupling degree C of key events is marked. t The spatial coordinates are converted into a map layer in the WGS84 coordinate system, and a time axis control is superimposed to realize dynamic time series display, generating an interactive spatiotemporal detection map. Users can zoom in and out to view the spatial distribution of pollution at different time points, and the visualization results can be traced back to the original monitoring data and processing results. S6.2: Transform the directed acyclic graph and propagation path analysis results of the environmental pollution data chain into a logically related visualization chart. Draw the environmental pollution data chain graph in the form of nodes to edges. The node size corresponds to the comprehensive score S, and the color depth of the edges reflects the propagation weight W. p The main propagation path is marked with a bold arrow, and the time interval Δt and the logical correlation strength W are labeled on the side. jSimultaneously, it generates a histogram ranking the importance of influencing factors, arranged in descending order of betweenness centrality, and associates them with equipment operating parameters. Based on the topology of the graph database, it uses a force-directed layout algorithm to optimize the graph visualization effect. When a node is hovered, its complete attributes are displayed, allowing users to trace the evidence source of any node to the original data table fields through interactive operations, thus forming a closed loop verification between the visualization results and the spatiotemporal logic analysis of the data chain.

[0011] An IoT-based environmental pollution detection data processing system includes a data acquisition module, a data preprocessing module, a pollutant feature extraction module, a data correlation analysis module, a detection result analysis module, and a visualization module. The data acquisition module is responsible for collecting raw data from IoT environmental monitoring equipment, sensor networks, and related systems, including but not limited to environmental monitoring, meteorological, and geographic information, as well as data from environmental management agencies and equipment operation risk factors, providing the system with a full range of raw data input. The data preprocessing module receives the raw data, cleans it by removing outliers, deduplicating, and repairing missing values, and standardizes heterogeneous data into a unified format to provide standardized data for subsequent analysis. The pollutant feature extraction module, based on the standardized data, analyzes the temporal trends of pollutant concentrations, the correlations between indicators, and spatial distribution characteristics to form pollution information, providing a foundation for data correlation analysis. The data correlation analysis module integrates the temporal, spatial, and logical correlations between pollution information and raw data to construct a multi-level environmental pollution data chain, making the data chain nodes traceable and supporting the analysis of detection results. The detection result analysis module, based on the environmental pollution data chain, determines the location of pollution sources and analyzes the direction of pollution propagation and the scope of impact. The output processing results form a closed-loop verification with relevant data; the visualization module transforms the processing results into spatiotemporal dimension visualization results and logical relationship charts, presenting them to users in an intuitive form and supporting interactive traceability.

[0012] Compared with the prior art, the beneficial effects of the present invention are: 1. A three-dimensional fusion framework of time, space, and logic is adopted: The time dimension calculates the coupling degree through a dynamic time warping algorithm to accurately identify the temporal synchronization between pollution events and equipment anomalies; the spatial dimension combines inverse distance weighted interpolation with geographic coordinate overlay to quantify the spatial correlation between monitoring points and pollution hotspots; the logical dimension constructs a directed acyclic graph with association rules as chains to reveal the causal relationship between pollutants and influencing factors. This modeling method breaks through the limitations of traditional single-dimensional analysis, improves the comprehensiveness and accuracy of pollution feature analysis, and can effectively identify implicit correlations in the detection and monitoring process.

[0013] 2. Dynamic path analysis improves the accuracy of pollution source tracing: The propagation weight is dynamically calculated by the betweenness centrality algorithm, which integrates time interval, spatial distance and logical strength. Combined with the shortest path algorithm, the main propagation path is identified, and the scope of influence is defined by the set of nodes with Wp≥0.2. This solves the problems of fixed weight and ambiguous scope in traditional path analysis. At the same time, the analysis results are verified by the spatiotemporal correlation and concentration trend of the original data to ensure the reliability of the source tracing conclusions. Attached Figure Description

[0014] Figure 1 This is a flowchart of an environmental pollution detection data processing method based on the Internet of Things according to the present invention; Figure 2 This is a flowchart illustrating the evidence chain tracing process of an environmental pollution detection data processing method based on the Internet of Things (IoT) according to the present invention. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] Example: Figures 1-2 As shown, the present invention provides a technical solution. A method for processing environmental pollution detection data based on the Internet of Things includes the following steps: S1. Collect raw data, obtain environmental monitoring data, meteorological data, geographic information data, and pollution source-related data from IoT environmental monitoring equipment, sensor networks, meteorological monitoring instruments, geographic information system databases, and pollution source information management systems, and at the same time collect data from environmental management agencies and pollution risk factors in the operation of IoT equipment; S2. Preprocess the collected raw data, including data cleaning and standardization, to unify the data format and provide usable data for subsequent analysis; S3. Based on standardized data, analyze the trends in environmental pollutant concentrations, the relationships between pollutants, and their spatial distribution characteristics to form pollution information; S4. Integrate pollution information with IoT data in terms of time and space to build an environmental pollution data chain; S5. Based on the constructed environmental pollution data chain, determine the location of the pollution source, the direction of spread, and the scope of impact, and output the processing results; S6. Transform the processing results into a visual format and present it to the user.

[0017] S1 further includes the following steps: S1.1: Collect full-volume structured data from IoT environmental monitoring devices and sensor networks, including structured fields such as device identifier, monitoring time, pollutant type, concentration value, and monitoring point location; specifically: adopt a real-time incremental acquisition method, and use a data interface program written in Python to locate sensors by device ID and obtain real-time monitoring data based on the API interface of the IoT platform. Then, use SQL statements to extract data from the sensor status table, monitoring data table, and device operation and maintenance record table through IoT database views. S1.2: Upload environmental management agency data in the form of a structured form through a data acquisition plugin deployed on the agency's internal IoT platform. Collect risk factors of environmental management agency data and IoT device operation from the environmental management agency data. Then, collect device operation logs in real time by embedding a monitoring program in the IoT gateway. The log fields contain environmental management data and device operation risk factors. Finally, store the above-mentioned risk factors and device operation logs in JSON format on the IoT log server.

[0018] S2 further includes the following steps: S2.1: Perform outlier removal, duplicate record deduplication, and missing value repair on the collected IoT environmental monitoring data, environmental management data, and equipment operation risk factors. For IoT monitoring data, outlier data is detected by setting standard value ranges for pollutant concentrations and combining them with the 3σ principle. Duplicate records are removed using the unique identifier of device ID + timestamp. For environmental management data, a regression model is built based on the monitoring point location using a multiple interpolation method to generate replacement values ​​for missing values. Cleaning rules are executed in stages according to data type. For equipment operation risk factors, duplicate and erroneous records are removed using timestamp continuity verification and sensor parameter threshold detection. Missing values ​​are filled by pattern matching from historical operation logs. S2.2: Transform heterogeneous IoT data into a unified format system. The monitoring time field is parsed into a preset standard format using regular expressions, and the device ID is given a classification prefix to form a standardized code. Numerical indicators such as pollutant concentration are mapped to a unified dimension using the Z-score standardization algorithm. The text description of the monitoring point location is geocoded to generate latitude and longitude feature vectors. Numerical fields in environmental management data are normalized. Equipment operation data is stored according to the standard format specifications in the IoT field. A field format conversion rule library for various types of data is established, and field mapping is performed based on the output data structure of the data acquisition.

[0019] S3 further includes the following steps: S3.1: Based on the standardized data obtained in S2.2, the concentration data is sorted in chronological order to construct a time series. A moving average model is used to filter short-term fluctuations, and a linear regression algorithm is combined to fit the long-term trend of changes, identifying concentration increase, decrease and periodic fluctuation patterns. Outlier correction based on the 3σ principle is performed on the time series data. Finally, the periodic components are decomposed through frequency domain analysis, and finally a pollutant concentration-time trend curve with characteristic parameters is generated. S3.2: Based on standardized data, an association rule algorithm is used to identify strong association rules by generating frequent itemsets and filtering rules that meet the threshold, in order to analyze the co-occurrence patterns among pollutant indicators. At the same time, a pollutant-influencing factor network is constructed with pollutant concentration as nodes and pollution source emission records in equipment operating parameters and environmental management data as edge weights. Then, the betweenness centrality algorithm is used to calculate the node centrality. Specifically, the concentration anomalies in the monitoring data are associated with equipment operating parameter records through key field mapping. Association rule mining and network topology analysis are performed to generate a pollutant association matrix and an importance ranking of influencing factors. The pollutant association matrix records the rule support and confidence, and the importance ranking is arranged in descending order of centrality value. S3.3: Based on standardized data, analyze the spatial distribution characteristics of pollutants, combine the latitude and longitude coordinates of monitoring points, and use spatial interpolation algorithms to generate a spatial distribution heat map of pollutants to identify high concentration areas.

[0020] S4 further includes the following steps: S4.1: In the time dimension, the pollutant concentration trend curve is aligned with the time series of equipment operation logs. A dynamic time warping algorithm is used to calculate the time series similarity, identifying the time coupling points between concentration anomalies and equipment anomalies. The time coupling degree C... t The calculation formula is as follows: ; Where T1 is the pollution event time series, T2 is the data anomaly time series, DTW(·) is the dynamic time warp distance, and C t The value ranges from [0,1]. The larger the value, the stronger the time correlation. len represents the length of the time series, which is the number of data points in the time series. It is used to normalize the Dynamic Time Warped Distance (DTW). S4.2: In the spatial dimension, the latitude and longitude coordinates of the monitoring points are overlaid with the spatial distribution hotspots of pollutants. An inverse distance weighted interpolation method is used to generate a pollution propagation heat map, determining the spatial correlation strength between the monitoring points and pollution hotspots, and the spatial influence weight W. s The calculation formula for W is s =1 / d, where d is the Euclidean distance between the monitoring point and the pollution hotspot, in km; S4.3: In terms of logic, the strongly associated rules generated by the association rule algorithm are used as chains to connect pollution sources, propagation paths, and influencing factors into a directed acyclic graph; S4.4: Establish a time-space-logic ternary association rule base, define time window threshold, spatial distance threshold and logical confidence threshold, and use graph database technology to associate pollution information with IoT raw data nodes. Node attributes include timestamp, geographic coordinates and association strength, with values ​​[0,1]. Finally, a multi-level environmental pollution data chain graph is generated, and each node can be traced back to the original data table fields and feature extraction results.

[0021] S5 further includes the following steps: S5.1: Based on a multi-level environmental pollution data chain diagram, determine the location of pollution sources and extract the temporal coupling degree Ct and spatial influence weight W of all pollution event nodes in the data chain. s and logical association strength W j W j The confidence level is used for confirmation, and then the comprehensive traceability score S is calculated using the following formula: ; Where α, β, and γ are weighting coefficients; Then, for each node, the time difference between its timestamp and the earliest pollution event is checked. Then, the weighted values ​​of spatial distance and logical association are combined to finally determine the node with the largest S value as the pollution source. S5.2: Analyze the direction and scope of pollution propagation. By analyzing the topological structure of the directed acyclic graph of the data chain, and combining the concentration time trend and spatial weights, the propagation weight W is calculated based on the betweenness centrality algorithm, starting from the source node. p The calculation formula is as follows: ; Where Δt is the time interval between adjacent nodes; Then, the data chain graph is topologically sorted, nodes are traversed in chronological order, and path priorities are dynamically adjusted through propagation weights. Next, the shortest path algorithm is used to identify the main propagation path, and a heatmap of the pollution impact range is generated through spatial interpolation, where the boundary of the impact range is determined by W. p The set of nodes with a value of ≥0.2 is determined, and the final output includes the propagation direction arrow, main path identifier, and boundary of the impact range. This output forms a closed loop verification with the spatiotemporal correlation data of the data chain and the concentration trend analysis of pollutant feature extraction. S6 further includes the following steps: S6.1: Based on the processing results, combined with timestamp and spatial coordinate information, the latitude and longitude coordinates of the source node and the timestamp of the pollution event are obtained from the processing results. Then, the pollution impact range is rendered into a heat map using GIS spatial interpolation technology, with the color gradient corresponding to the spatial impact weight W. s The value; the pollution event sequence is displayed in time axis form, with each time point associated with the pollutant concentration trend curve, and the temporal coupling degree C of key events is marked. t The spatial coordinates are converted into a map layer in the WGS84 coordinate system, and a time axis control is superimposed to realize dynamic time series display, generating an interactive spatiotemporal detection map. Users can zoom in and out to view the spatial distribution of pollution at different time points, and the visualization results can be traced back to the original monitoring data and processing results. S6.2: Transform the directed acyclic graph and propagation path analysis results of the environmental pollution data chain into a logically related visualization chart. Draw the environmental pollution data chain graph in the form of nodes to edges. The node size corresponds to the comprehensive score S, and the color depth of the edges reflects the propagation weight W. p The main propagation path is marked with a bold arrow, and the time interval Δt and the logical correlation strength W are labeled on the side. j Simultaneously, it generates a histogram ranking the importance of influencing factors, arranged in descending order of betweenness centrality, and associates them with equipment operating parameters. Based on the topology of the graph database, it uses a force-directed layout algorithm to optimize the graph visualization effect. When a node is hovered, its complete attributes are displayed, allowing users to trace the evidence source of any node to the original data table fields through interactive operations, thus forming a closed loop verification between the visualization results and the spatiotemporal logic analysis of the data chain.

[0022] An IoT-based environmental pollution detection data processing system includes a data acquisition module, a data preprocessing module, a pollutant feature extraction module, a data correlation analysis module, a detection result analysis module, and a visualization module. The data acquisition module is responsible for collecting raw data from IoT environmental monitoring equipment, sensor networks, and related systems, including but not limited to environmental monitoring, meteorological, and geographic information, as well as data from environmental management agencies and equipment operation risk factors, providing the system with a full range of raw data input. The data preprocessing module receives the raw data, cleans it by removing outliers, deduplicating, and repairing missing values, and standardizes heterogeneous data into a unified format to provide standardized data for subsequent analysis. The pollutant feature extraction module, based on the standardized data, analyzes the temporal trends of pollutant concentrations, the correlations between indicators, and spatial distribution characteristics to form pollution information, providing a foundation for data correlation analysis. The data correlation analysis module integrates the temporal, spatial, and logical correlations between pollution information and raw data to construct a multi-level environmental pollution data chain, making the data chain nodes traceable and supporting the analysis of detection results. The detection result analysis module, based on the environmental pollution data chain, determines the location of pollution sources and analyzes the direction of pollution propagation and the scope of impact. The output processing results form a closed-loop verification with relevant data; the visualization module transforms the processing results into spatiotemporal dimension visualization results and logical relationship charts, presenting them to users in an intuitive form and supporting interactive traceability.

[0023] Suppose that on August 10, 2024, the IoT monitoring network around an industrial park shows that the PM2.5 concentration has exceeded the national standard (75 μg / m³) for three consecutive days, reaching a maximum of 120 μg / m³. The park has 10 IoT monitoring points (numbered S1-S10), including atmospheric sensors, weather stations, and pollution source emission monitoring equipment. This system is needed to trace the pollution source and its propagation path.

[0024] First, raw data is collected, including two parts: IoT monitoring data and environmental management and equipment operation data. The IoT monitoring data is collected in real-time incrementally by device ID (e.g., S1-S10) through a data interface program written in Python, based on the park's IoT platform API. This includes sensor data such as PM2.5 concentration, SO2 concentration, monitoring time, and latitude and longitude of the monitoring point; meteorological data such as wind speed, wind direction, and humidity; and sensor status tables and equipment maintenance records extracted from the IoT database view. For example, S3 experienced a calibration anomaly at 08:00 on August 9th, and S5 had its filter replaced at 12:00 on August 8th. Regarding environmental management and equipment operation data, the park's environmental protection bureau uploads structured forms through its internal IoT platform, which include the exhaust emission records of three companies, A, B, and C. Company A's emissions on August 8 were 1200 kg, exceeding the limit by 200 kg. At the same time, a monitoring program is embedded in the IoT gateway to capture equipment operation logs, such as the S3 calibration parameter error log: 2024-08-09 08:00, calibration value deviated from the standard by 15%, and stored in JSON format on the log server.

[0025] The collected raw data is then preprocessed, consisting of two stages: data cleaning and standardization. During data cleaning, for IoT monitoring data, a standard PM2.5 range of 0-500 μg / m³ is set. Following the 3σ principle, a mean of 80 and a standard deviation of 30 are used. An outlier of 250 μg / m³ (actually a calibration error) at 08:00 on August 9th is removed from S3. Duplicate data is removed using "device ID + timestamp," and duplicate data from 14:00 on August 8th uploaded by S7 is deleted. For environmental management data, multiple interpolation is used, and a regression model is built based on historical emission data from Company A to correct the missing emission record at 16:00 on August 8th (predicted value 1100 kg). For equipment operation logs, duplicate calibration records at 12:05 on August 8th for S5 are deleted through timestamp continuity verification, and the missing operating status of S2 at 00:00 on August 9th is removed. The data is filled in as "normal" by matching the historical log pattern; in the data standardization process, the monitoring time is uniformly parsed as "YYYY-MM-DDHH:MM:SS" (e.g., "2024 / 8 / 88:00" → "2024-08-0808:00:00"); the device ID is prefixed with "Sensor-", e.g., S1 → Sensor-S1; in terms of numerical indicator standardization, PM2.5 concentration is mapped using the Z-score algorithm, e.g., 120μg / m³ → (120-80) / 30≈1.33, and enterprise emissions are normalized to the [0,1] interval, e.g., 1200kg for enterprise A → 1200 / 1500=0.8; the monitoring point address, such as "northeast of the park", is generated as a latitude and longitude feature vector through geocoding.

[0026] Subsequently, pollutant characteristics were extracted based on standardized data, focusing on three dimensions: temporal trend, correlation, and spatial distribution. In the temporal trend analysis, a PM2.5 concentration time series was constructed, using a 5-point moving average to filter short-term fluctuations. A linear regression model was used to fit the trend equation y=5x+70 (x being the number of days), identifying August 8th at 06:00 as the inflection point for concentration increases. Frequency domain analysis showed that the daily peak concentration occurred between 08:00 and 18:00 (a 24-hour cycle), coinciding with the company's production periods. The correlation analysis employed an association rule algorithm, generating a frequent itemset {PM2.5 exceeding the standard, Company A's emissions >1000kg}, with a support of 0.6 (co-occurring on 1.8 days out of 3) and a confidence level of 0.9 (90% probability of PM2.5 exceeding the standard when Company A exceeds its emission limits). Simultaneously, a "pollutant-influencing factor network" was constructed, with nodes: PM2.5, SO2, Company A's emissions, and S3 calibration; edge weights: association rule confidence. Betweenness centrality calculations showed that the "Company A's emissions" node had a centrality of 0.75 (the highest), making it a key influencing factor. Spatial distribution analysis combined with the latitude and longitude of monitoring points and used inverse distance weighted interpolation to generate a PM2.5 spatial hotspot map, showing that the hotspot area is concentrated around Company A (coordinates 30.13°N, 120.06°E), which overlaps with monitoring points S1-S3.

[0027] Based on this, an environmental pollution data chain is constructed, which is achieved through time coupling, spatial correlation, logical link construction, and the generation of multi-level data chains. In the time coupling analysis, the time series of excessive emissions from Company A, such as T1: August 8, 06:00-18:00, is aligned with the PM2.5 exceeding the standard time series, such as T2: August 8, 08:00-20:00. The dynamic time warping algorithm calculates the DTW distance as 5, and the time coupling degree C. t =1-5 / max(12,12)=0.58 (strong correlation). In spatial correlation analysis, the coordinates of company A (30.13°N, 120.06°E) are d=0.5km away from the pollution hotspot, and the spatial influence weight W s =1 / 0.5=2 (normalized to 1.0); the distance between monitoring point S1 and the hotspot is d=1km, W s =0.5. The logical link is constructed using the strong association rule "Company A emissions → PM2.5 exceedance" as the chain, connecting "S3 calibration anomaly → data distortion" and "northeast wind → pollution diffusion" to form a directed acyclic graph. Finally, a ternary association rule base is established (time window ±2h, spatial distance <2km, confidence ≥0.8), and nodes are associated through the graph database with attributes: timestamp, coordinates, and association strength 0.8, generating a data chain graph. Nodes can be traced back to the original emission record table and sensor data table.

[0028] Next, pollution source tracing analysis will be conducted, including locating the pollution source and analyzing its propagation path and scope. When locating the pollution source, the C-value of the data link node will be extracted. t=0.58, W s =1.0, Logical Association Strength W j =0.9 (Confidence level of A Company's emissions and PM2.5 exceedance), weighting coefficients α=0.3, β=0.4, γ=0.3, calculated comprehensive score S=0.3×0.58+0.4×1.0+0.3×0.9=0.824; after verifying the node timestamp, it was found that A Company's excessive emissions (August 8th 06:00) were earlier than the earliest PM2.5 exceedance (August 8th 08:00), a time difference of 2 hours. Combining spatial distance and logical correlation, A Company was determined to be the pollution source. When analyzing the transmission path and scope, A Company was taken as the starting point, and the transmission weight W was calculated. p = (Ws × Wj) / (Δt² + 1), for example, Δt = 2h from company A to S1, W p = (1.0 × 0.9) / (4 + 1) = 0.18; After topological sorting, nodes are traversed by time, and the main path "Company A → S1 → S2" (W) is identified using the shortest path algorithm. p (Mean 0.15); Spatial interpolation generates a heatmap of the influence range, W p The boundary of the node set with a value ≥0.2 is defined as a 1.5km range in the northeast of the park, which is consistent with the northeast wind direction and concentration trend (closed-loop verification).

[0029] Finally, visualization is presented, divided into spatiotemporal visualization and logical visualization. In the spatiotemporal visualization, the GIS heat map shows that the area surrounding Company A is a high-concentration red zone (W). s =1.0), the timeline marks the A company's excessive emission event at 06:00 on August 8th, and the associated concentration trend curve shows that PM2.5 continued to rise thereafter, labeled C. t =0.58; Users can zoom to view the concentration value of S1 at 10:00 on August 8th, which is 100 μg / m³. In the logical visualization, the node of Company A (S=0.824) in the node-edge graph has the largest diameter, the arrow of the main path "Company A → S1" is bolded, and Δt=2h and W are marked. j =0.9; The bar chart shows that "Company A's emissions" has a centrality of 0.75, making it the primary influencing factor. Hovering over the node allows you to view the original emission records, such as the emission of 1200kg on August 8th.

[0030] This system traces the pollution source to excessive emissions from Company A on August 8th, which, driven by northeasterly winds, spread towards monitoring points S1-S2, affecting an area of ​​1.5 km in the northeast of the industrial park. A calibration anomaly at sensor S3 caused localized data distortion, but did not affect the source location. The system's output evidence chain can be traced back to the original emission data (Company Record ID: EP2024080801) and sensor logs (Log ID: LOG2024080908), verifying the accuracy and traceability of the method.

[0031] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

Claims

1. A method for processing environmental pollution detection data based on the Internet of Things, characterized in that: Includes the following steps: S1. Collect raw data, obtain environmental monitoring data, meteorological data, geographic information data, and pollution source-related data from IoT environmental monitoring equipment, sensor networks, meteorological monitoring instruments, geographic information system databases, and pollution source information management systems, and at the same time collect data from environmental management agencies and pollution risk factors in the operation of IoT equipment; S2. Preprocess the collected raw data, including data cleaning and standardization, to unify the data format and provide usable data for subsequent analysis; S3. Based on standardized data, analyze the trends in environmental pollutant concentrations, the relationships between pollutants, and their spatial distribution characteristics to form pollution information; S4. Integrate pollution information with IoT data in terms of time and space to build an environmental pollution data chain; S5. Based on the constructed environmental pollution data chain, determine the location of the pollution source, the direction of spread, and the scope of impact, and output the processing results; S6. Transform the processing results into a visual format and present it to the user.

2. The method for processing environmental pollution detection data based on the Internet of Things according to claim 1, characterized in that: S1 further includes the following steps: S1.1: Collect full-volume structured data from IoT environmental monitoring devices and sensor networks, including structured fields such as device identifier, monitoring time, pollutant type, concentration value, and monitoring point location; specifically: adopt a real-time incremental acquisition method, and use a data interface program written in Python to locate sensors by device ID and obtain real-time monitoring data based on the API interface of the IoT platform. Then, use SQL statements to extract data from the sensor status table, monitoring data table, and device operation and maintenance record table through IoT database views. S1.2: Upload environmental management agency data in the form of a structured form through a data acquisition plugin deployed on the agency's internal IoT platform. Collect risk factors of environmental management agency data and IoT device operation from the environmental management agency data. Then, collect device operation logs in real time by embedding a monitoring program in the IoT gateway. The log fields contain environmental management data and device operation risk factors. Finally, store the above-mentioned risk factors and device operation logs in JSON format on the IoT log server.

3. The method for processing environmental pollution detection data based on the Internet of Things according to claim 1, characterized in that: S2 further includes the following steps: S2.1: Perform outlier removal, duplicate record deduplication, and missing value repair on the collected IoT environmental monitoring data, environmental management data, and equipment operation risk factors. For IoT monitoring data, outlier data is detected by setting standard value ranges for pollutant concentrations and combining them with the 3σ principle. Duplicate records are removed using the unique identifier of device ID + timestamp. For environmental management data, a regression model is built based on the monitoring point location using a multiple interpolation method to generate replacement values ​​for missing values. Cleaning rules are executed in stages according to data type. For equipment operation risk factors, duplicate and erroneous records are removed using timestamp continuity verification and sensor parameter threshold detection. Missing values ​​are filled by pattern matching from historical operation logs. S2.2: Transform heterogeneous IoT data into a unified format system. The monitoring time field is parsed into a preset standard format using regular expressions, and the device ID is given a classification prefix to form a standardized code. Numerical indicators such as pollutant concentration are mapped to a unified dimension using the Z-score standardization algorithm. The text description of the monitoring point location is geocoded to generate latitude and longitude feature vectors. Numerical fields in environmental management data are normalized. Equipment operation data is stored according to the standard format specifications in the IoT field. A field format conversion rule library for various types of data is established, and field mapping is performed based on the output data structure of the data acquisition.

4. The method for processing environmental pollution detection data based on the Internet of Things according to claim 1, characterized in that: S3 further includes the following steps: S3.1: Based on the standardized data obtained in S2.2, the concentration data is sorted in chronological order to construct a time series. A moving average model is used to filter short-term fluctuations, and a linear regression algorithm is combined to fit the long-term trend of changes, identifying concentration increase, decrease and periodic fluctuation patterns. Outlier correction based on the 3σ principle is performed on the time series data. Finally, the periodic components are decomposed through frequency domain analysis, and finally a pollutant concentration-time trend curve with characteristic parameters is generated. S3.2: Based on standardized data, an association rule algorithm is used to identify strong association rules by generating frequent itemsets and filtering rules that meet the threshold, in order to analyze the co-occurrence patterns among pollutant indicators. At the same time, a pollutant-influencing factor network is constructed with pollutant concentration as nodes and pollution source emission records in equipment operating parameters and environmental management data as edge weights. Then, the betweenness centrality algorithm is used to calculate the node centrality. Specifically, the concentration anomalies in the monitoring data are associated with equipment operating parameter records through key field mapping. Association rule mining and network topology analysis are performed to generate a pollutant association matrix and an importance ranking of influencing factors. The pollutant association matrix records the rule support and confidence, and the importance ranking is arranged in descending order of centrality value. S3.3: Based on standardized data, analyze the spatial distribution characteristics of pollutants, combine the latitude and longitude coordinates of monitoring points, and use spatial interpolation algorithms to generate a spatial distribution heat map of pollutants to identify high concentration areas.

5. The method for processing environmental pollution detection data based on the Internet of Things according to claim 1, characterized in that: S4 further includes the following steps: S4.1: In the time dimension, the pollutant concentration trend curve is aligned with the time series of equipment operation logs. A dynamic time warping algorithm is used to calculate the time series similarity, identifying the time coupling points between concentration anomalies and equipment anomalies. The time coupling degree C... t The calculation formula is as follows: ; Where T1 is the pollution event time series, T2 is the data anomaly time series, DTW(·) is the dynamic time warp distance, and C t The value ranges from [0,1]. The larger the value, the stronger the time correlation. len represents the length of the time series, which is the number of data points in the time series. It is used to normalize the Dynamic Time Warped Distance (DTW). S4.2: In the spatial dimension, the latitude and longitude coordinates of the monitoring points are overlaid with the spatial distribution hotspots of pollutants. An inverse distance weighted interpolation method is used to generate a pollution propagation heat map, determining the spatial correlation strength between the monitoring points and pollution hotspots, and the spatial influence weight W. s The calculation formula for W is s =1 / d, where d is the Euclidean distance between the monitoring point and the pollution hotspot, in km; S4.3: In terms of logic, the strongly associated rules generated by the association rule algorithm are used as chains to connect pollution sources, propagation paths, and influencing factors into a directed acyclic graph; S4.4: Establish a time-space-logic ternary association rule base, define time window threshold, spatial distance threshold and logical confidence threshold, and use graph database technology to associate pollution information with IoT raw data nodes. Node attributes include timestamp, geographic coordinates and association strength, with values ​​[0,1]. Finally, a multi-level environmental pollution data chain graph is generated, and each node can be traced back to the original data table fields and feature extraction results.

6. The method for processing environmental pollution detection data based on the Internet of Things according to claim 1, characterized in that: S5 further includes the following steps: S5.1: Based on a multi-level environmental pollution data chain diagram, determine the location of pollution sources and extract the temporal coupling degree Ct and spatial influence weight W of all pollution event nodes in the data chain. s and logical association strength W j W j The confidence level is used for confirmation, and then the comprehensive traceability score S is calculated using the following formula: ; Where α, β, and γ are weighting coefficients; Then, for each node, the time difference between its timestamp and the earliest pollution event is checked. Then, the weighted values ​​of spatial distance and logical association are combined to finally determine the node with the largest S value as the pollution source. S5.2: Analyze the direction and scope of pollution propagation. By analyzing the topological structure of the directed acyclic graph of the data chain, and combining the concentration time trend and spatial weights, the propagation weight W is calculated based on the betweenness centrality algorithm, starting from the source node. p The calculation formula is as follows: ; Where Δt is the time interval between adjacent nodes; Then, the data chain graph is topologically sorted, nodes are traversed in chronological order, and path priorities are dynamically adjusted through propagation weights. Next, the shortest path algorithm is used to identify the main propagation path, and a heatmap of the pollution impact range is generated through spatial interpolation, where the boundary of the impact range is determined by W. p The set of nodes with a value of ≥0.2 is determined, and the final output includes the propagation direction arrow, the main path identifier, and the boundary of the impact range. This output forms a closed loop verification with the spatiotemporal correlation data of the data chain and the concentration trend analysis of pollutant feature extraction.

7. The method for processing environmental pollution detection data based on the Internet of Things according to claim 1, characterized in that: S6 further includes the following steps: S6.1: Based on the processing results, combined with timestamp and spatial coordinate information, the latitude and longitude coordinates of the source node and the timestamp of the pollution event are obtained from the processing results. Then, the pollution impact range is rendered into a heat map using GIS spatial interpolation technology, with the color gradient corresponding to the spatial impact weight W. s The value; the pollution event sequence is displayed in time axis form, with each time point associated with the pollutant concentration trend curve, and the temporal coupling degree C of key events is marked. t The spatial coordinates are converted into a map layer in the WGS84 coordinate system, and a time axis control is superimposed to realize dynamic time series display, generating an interactive spatiotemporal detection map. Users can zoom in and out to view the spatial distribution of pollution at different time points, and the visualization results can be traced back to the original monitoring data and processing results. S6.2: Transform the directed acyclic graph and propagation path analysis results of the environmental pollution data chain into a logically related visualization chart. Draw the environmental pollution data chain graph in the form of nodes to edges. The node size corresponds to the comprehensive score S, and the color depth of the edges reflects the propagation weight W. p The main propagation path is marked with a bold arrow, and the time interval Δt and logical correlation strength W are labeled on the side. j Simultaneously, it generates a histogram ranking the importance of influencing factors, arranged in descending order of betweenness centrality, and associates them with equipment operating parameters. Based on the topology of the graph database, it uses a force-directed layout algorithm to optimize the graph visualization effect. When a node is hovered, its complete attributes are displayed, allowing users to trace the evidence source of any node to the original data table fields through interactive operations, thus forming a closed loop verification between the visualization results and the spatiotemporal logic analysis of the data chain.

8. An Internet of Things (IoT)-based environmental pollution detection data processing system, applied to any one of the IoT-based environmental pollution detection data processing methods described in claims 1-7, characterized in that: The system comprises a data acquisition module, a data preprocessing module, a pollutant feature extraction module, a data correlation analysis module, a detection result analysis module, and a visualization module. The data acquisition module is responsible for collecting raw data from IoT environmental monitoring equipment, sensor networks, and related systems, including but not limited to environmental monitoring, meteorological, and geographic information, as well as data from environmental management agencies and equipment operation risk factors, providing the system with a full range of raw data input. The data preprocessing module receives the raw data, cleans it by removing outliers, deduplicating, and repairing missing values, and standardizes heterogeneous data into a unified format, providing standardized data for subsequent analysis. The pollutant feature extraction module, based on standardized data, analyzes the temporal trends of pollutant concentrations, the correlations between indicators, and spatial distribution characteristics to form pollution information, providing a foundation for data correlation analysis. The data correlation analysis module integrates the temporal, spatial, and logical correlations between pollution information and raw data, constructing a multi-level environmental pollution data chain, making data chain nodes traceable and supporting detection result analysis. The detection result analysis module, based on the environmental pollution data chain, determines the location of pollution sources and analyzes the direction of pollution propagation and the scope of impact. Output the processing results and form a closed-loop verification with relevant data; The visualization module transforms the processing results into spatiotemporal visualizations and logical relationship charts, presenting them to users in an intuitive way and supporting interactive traceability.

Citation Information

Cited By

  • Intelligent flight decision-making auxiliary system and method based on multi-modal flight data

    CN121300093A

  • Multi-source heterogeneous data association method and system for environmental assessment detection

    CN121834251A

  • Multi-source heterogeneous data association method and system for environmental impact assessment detection

    CN121834251B

  • Intelligent environment monitoring method and system for meat product processing

    CN122108276A

  • An artificial intelligence-based pollution source monitoring data anomaly detection method and system

    CN122490383A