A system and method for detecting air pollution based on big data
Through big data technology and sensor network, real-time monitoring and analysis of air environmental pollution is achieved, and the problems of low data coverage density and poor real-time performance of traditional air detection technology are solved, the accuracy and timeliness of air quality monitoring are improved, and abnormal types are automatically identified and scientific emergency response is provided.
Patent Information
- Application Number
- CN202510224565.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Traditional air detection technology has low data coverage density, poor real-time performance, insufficient multi-factor coupling analysis capabilities, making it difficult to accurately distinguish sudden pollution and long-term accumulated pollution, and the weight allocation method is fixed, which cannot reflect the actual impact of different pollutants under different time and space conditions.
The air environment pollution detection method based on big data is adopted, and multi-source environmental data is collected through a distributed sensor network, sliding window filtering and ARIMA missing value filling are performed, the pollution comprehensive index is calculated, a dual-modal early warning mechanism is established, and an abnormal type is automatically judged using the correlation intensity coefficient and matched the emergency response plan.
Real-time monitoring and analysis of air environmental pollution is realized, the calculation accuracy of the comprehensive pollution index and the accuracy and timeliness of air quality monitoring are improved, abnormal types are automatically identified, false alarms or missed reports are avoided, and the continuity and stability of air quality monitoring are ensured.
Smart Images

Figure CN120102792B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of air pollution detection, and specifically relates to an air environment pollution detection system and method based on big data. Background Art
[0002] With the rapid development of industrialization and urbanization, the problem of air pollution has become increasingly serious. In order to ensure environmental safety, real-time monitoring and early warning of air quality have become particularly important. This will enable us to grasp the air quality status in a timely and accurate manner, prevent the occurrence of environmental pollution incidents, and enable relevant departments to take timely response measures to reduce pollutant emissions, improve air quality, and provide corresponding protection for public health.
[0003] Traditional air detection technology mainly relies on fixed monitoring sites and single pollutant analysis, and has defects such as low data coverage density, poor real-time performance, and insufficient multi-factor coupling analysis capabilities. The existing methods are mostly based on sensor network monitoring systems. Although basic data collection can be achieved, the problem of fuzzy classification of abnormal events is not solved. In particular, there is a lack of dynamic decision-making mechanism in the judgment of sudden pollution and long-term cumulative pollution, making it impossible to issue accurate alarm information. In addition, the weight distribution method mostly uses fixed coefficients, which makes it difficult to accurately reflect the actual impact of different pollutants under different time and space conditions. Based on this, the present invention proposes an air environment pollution detection method based on big data to solve the above problems. Summary of the Invention
[0004] The purpose of the present invention is to provide an air pollution detection system and method based on big data, which can realize real-time monitoring and analysis of multi-source environmental data, accurately determine the air pollution level, and accurately classify specific abnormal types to facilitate matching corresponding emergency response plans.
[0005] The technical solutions adopted by the present invention are as follows:
[0006] A method for detecting air pollution based on big data, comprising:
[0007] Collect multi-source environmental data of the target area through a distributed sensor network, including concentrations of PM2.5, PM10, SO2, NOx, O3, as well as temperature, humidity, wind speed, and wind direction parameters;
[0008] Preprocess multi-source environmental data, including sliding window filtering and ARIMA missing value filling;
[0009] Calculating a comprehensive pollution index within the monitoring area based on the preprocessed multi-source environmental data, and determining an air quality state of the environment within the monitoring area based on the comprehensive pollution index, wherein the air quality state includes a normal state and an abnormal state;
[0010] A dual-modal early warning mechanism including transient anomalies and normalized anomalies is established, and the correlation between anomaly types is automatically determined through the correlation strength coefficient, and corresponding emergency response plans are matched according to different anomaly types.
[0011] In a preferred embodiment, the preprocessing of multi-source environmental data, including the steps of sliding window filtering and ARIMA missing value filling, comprises:
[0012] The dynamic threshold method is used to clean the collected multi-source environmental data to remove noise data and outliers;
[0013] Use sliding window filtering technology to reduce random fluctuations in multi-source environmental data;
[0014] The ARIMA model is applied to fill missing values, and based on the time series characteristics of multi-source environmental data, missing data points are predicted and filled.
[0015] In a preferred embodiment, the step of calculating the comprehensive pollution index in the monitoring area based on the pre-processed multi-source environmental data includes:
[0016] Multiple monitoring points are set up in the monitoring area, and the pollutant concentration of the pollution index at each monitoring point is collected and simultaneously recorded as the first characteristic parameter;
[0017] Obtaining the proportion of the first characteristic parameter at each monitoring point and recording it as a second characteristic parameter, and then determining the initial weight of each pollution indicator based on the second characteristic parameter;
[0018] The initial weights of each pollution index are corrected and integrated to obtain the dynamic weights of each pollution index;
[0019] The dynamic weights of pollution indicators and pollution concentrations at each monitoring point are aggregated and calculated to obtain a comprehensive pollution index within the monitoring area.
[0020] In a preferred embodiment, the pollution index includes a positive index and a negative index, the positive index is positively correlated with the concentration of the pollution index, and the negative index is negatively correlated with the concentration of the pollutant;
[0021] The positive indicator and the negative indicator are normalized by a preset normalization function, and the normalized result is output as the first characteristic parameter.
[0022] In a preferred embodiment, the step of determining the air quality state of the environment in the monitoring area based on the comprehensive pollution index includes:
[0023] Obtaining a preset comprehensive assessment threshold, and comparing the comprehensive assessment threshold with a comprehensive pollution index;
[0024] When the comprehensive pollution index is higher than or equal to the assessment threshold, it indicates that the air quality in the monitored area is abnormal and an alarm signal is issued simultaneously;
[0025] When the comprehensive pollution index is lower than the assessment threshold, it indicates that the air quality of the environment in the monitoring area is normal, and routine monitoring of the monitoring area will continue.
[0026] In a preferred embodiment, the step of establishing a dual-mode early warning mechanism including transient anomalies and normalized anomalies includes:
[0027] Under abnormal conditions, the time window threshold and volatility threshold are preset, and based on the time window threshold and volatility threshold;
[0028] The pollutant concentration change rate of the pollution indicators in the monitoring area is collected in real time and recorded as the third characteristic parameter. When the third characteristic parameter exceeds the fluctuation rate threshold, the duration of the third characteristic parameter exceeding the fluctuation rate threshold is counted and recorded as the fourth characteristic parameter.
[0029] The fourth characteristic parameter is compared with the time window threshold, and when the fourth characteristic parameter is greater than the time window threshold, the abnormal state corresponding to the fourth characteristic parameter is recorded as a transient abnormality, otherwise it is recorded as a normalized abnormality.
[0030] In a preferred embodiment, the step of automatically determining the correlation between abnormality types by using the correlation strength coefficient includes:
[0031] Construct a spatiotemporal correlation matrix and calculate the linear correlation between each pollutant through the spatiotemporal correlation matrix. Then calculate the nonlinear spatiotemporal weight by the overlap times and time windows between normalized anomalies and transient anomalies.
[0032] By fusing the nonlinear spatiotemporal weight with the linear correlation, the correlation strength coefficient between anomaly types is output;
[0033] According to the preset classification interval, the correlation strength coefficient is compared with the classification interval;
[0034] When the correlation strength coefficient exceeds the upper limit of the classification interval, it indicates that the correlation between the transient anomaly and the normalized anomaly is strong, and a high-risk warning signal is issued;
[0035] When the correlation strength coefficient falls within the classification interval, it indicates that the correlation between the transient anomaly and the normalized anomaly is moderate, and a normalized risk warning signal is issued;
[0036] When the correlation strength coefficient is lower than the lower limit of the classification interval, it indicates that the correlation between the transient anomaly and the normalized anomaly is weak, and they are independent events, and a low-risk warning signal is issued.
[0037] In a preferred embodiment, the step of matching corresponding emergency response plans according to different abnormality types includes:
[0038] Obtaining the type of abnormal state, and based on the type of abnormal state, retrieving an emergency plan that matches the type of abnormal state from a preset emergency response plan library, wherein the emergency plan includes a high-risk emergency response plan and a normalized emergency response plan;
[0039] When a transient anomaly is detected, a correlation strength coefficient between the transient anomaly and the normalized anomaly is continuously determined, and when the correlation strength coefficient corresponds to a high-risk warning signal, a high-risk emergency response plan is initiated;
[0040] When a moderate correlation is detected between transient anomalies and normalized anomalies, the normalized emergency response plan is directly called, and the air quality in the monitoring area is continuously monitored until the air quality returns to normal.
[0041] The present invention also provides a big data-based air pollution detection system, which uses the above-mentioned big data-based air pollution detection method, including:
[0042] A data acquisition module is used to collect multi-source environmental data of the target area through a distributed sensor network, including concentration values of PM2.5, PM10, SO2, NOx, O3, as well as temperature, humidity, wind speed, and wind direction parameters;
[0043] A preprocessing module is used to preprocess multi-source environmental data, including sliding window filtering and ARIMA missing value filling;
[0044] A quality assessment module, the quality assessment module is used to calculate the comprehensive pollution index within the monitoring area based on the preprocessed multi-source environmental data, and determine the air quality status of the environment within the monitoring area based on the comprehensive pollution index, wherein the air quality status includes normal state and abnormal state;
[0045] The abnormality identification module is used to establish a dual-modal early warning mechanism including transient abnormalities and normalized abnormalities, automatically identify the correlation between abnormality types through the correlation strength coefficient, and match corresponding emergency response plans according to different abnormality types.
[0046] And, an electronic device, comprising:
[0047] at least one processor;
[0048] and a memory communicatively coupled to the at least one processor;
[0049] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above-mentioned air pollution detection method based on big data.
[0050] The technical effects achieved by the present invention are:
[0051] The present invention realizes the detection and early warning of air pollution by mining and analyzing multi-source environmental data. When determining the comprehensive pollution index in the monitoring area, a dynamic weight allocation strategy is adopted, thereby effectively improving the calculation accuracy of the comprehensive pollution index and improving the accuracy and timeliness of air quality monitoring. In addition, by constructing a spatiotemporal correlation matrix and calculating the correlation strength coefficient, the present invention can automatically identify the correlation between anomaly types, providing a scientific basis for emergency response. At the same time, the dual-modal early warning mechanism proposed in the present invention can distinguish between transient anomalies and normalized anomalies, avoiding the waste of resources and untimely emergency response caused by false alarms or missed alarms, and ensuring the continuity and stability of air quality monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a schematic flow chart of the method of the present invention;
[0053] Figure 2 It is a schematic diagram of the system modules of the present invention;
[0054] Figure 3 It is a schematic structural diagram of an electronic device of the present invention. DETAILED DESCRIPTION
[0055] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0056] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0057] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in a preferred embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it constitute a separate or selective embodiment that is mutually exclusive of other embodiments.
[0058] See also Figure 1 As shown, the present invention provides an air pollution detection method based on big data, comprising:
[0059] S1. Collect multi-source environmental data of the target area through a distributed sensor network, including concentration values of PM2.5, PM10, SO2, NOx, O3, as well as temperature, humidity, wind speed, and wind direction parameters;
[0060] In step S1, when collecting the air pollution level in the area to be monitored, it is necessary to use a distributed sensor network to monitor the target area and collect multi-source environmental data in the area. The multi-source environmental data covers the concentration values of various pollutants, such as fine particulate matter PM2.5, coarse particulate matter PM10, sulfur dioxide SO2, nitrogen oxides NOx and ozone O3. In addition, it is also necessary to monitor meteorological parameters such as temperature, humidity, wind speed, and wind direction to comprehensively reflect the air quality status of the target area and possible influencing factors.
[0061] S2. Preprocessing of multi-source environmental data, including sliding window filtering and ARIMA missing value filling;
[0062] In step S2, after the multi-source environmental data in the area to be monitored is collected, it is necessary to preprocess the collected multi-source environmental data. The preprocessing steps include using sliding window filtering technology to remove noise, applying the ARIMA model to fill missing values in the data, and performing correlation analysis between pollutants to ensure the accuracy and reliability of the data. The preprocessing of the multi-source environmental data, including the sliding window filtering and ARIMA missing value filling steps, includes:
[0063] The dynamic threshold method is used to clean the collected multi-source environmental data to remove noise data and outliers;
[0064] Use sliding window filtering technology to reduce random fluctuations in multi-source environmental data;
[0065] Apply the ARIMA model to fill missing values, predict and fill missing data points based on the time series characteristics of multi-source environmental data;
[0066] Specifically, when processing multi-source environmental data, the dynamic threshold method is first used to perform preliminary cleaning on the multi-source environmental data that has been collected. This process is mainly to remove noise data and outliers to ensure the accuracy of subsequent analysis. Then, the sliding window filtering technology is used to reduce the random fluctuations in the multi-source environmental data. In this way, the data can be effectively smoothed, unnecessary interference can be reduced, and a more stable data basis can be provided for subsequent analysis. The ARIMA model is then applied to fill the missing values in the data. Since multi-source environmental data has obvious time series characteristics, the ARIMA model is used to predict missing data points and fill these gaps accordingly to ensure the integrity and continuity of the data.
[0067] S3. Calculating a comprehensive pollution index within the monitoring area based on the preprocessed multi-source environmental data, and determining an air quality state within the monitoring area based on the comprehensive pollution index, wherein the air quality state includes a normal state and an abnormal state;
[0068] In step S3, after the multi-source environmental data in the monitored area is preprocessed, a comprehensive pollution index in the monitored area is calculated based on the preprocessed multi-source environmental data, and then the ambient air quality state in the monitored area is evaluated using the comprehensive pollution index. In this embodiment, the air quality state is divided into two types: a normal state and an abnormal state. The step of calculating the comprehensive pollution index in the monitored area based on the preprocessed multi-source environmental data includes:
[0069] Multiple monitoring points are set up in the monitoring area, and the pollutant concentration of the pollution index at each monitoring point is collected and simultaneously recorded as the first characteristic parameter;
[0070] Obtain the proportion of the first characteristic parameter at each monitoring point and record it as the second characteristic parameter, and then determine the initial weight of each pollution indicator based on the second characteristic parameter;
[0071] The initial weights of each pollution index are corrected and integrated to obtain the dynamic weights of each pollution index;
[0072] The dynamic weights of pollution indicators and pollution concentrations at each monitoring point are aggregated and calculated to obtain a comprehensive pollution index within the monitoring area.
[0073] Specifically, in order to accurately calculate the comprehensive pollution index in the monitoring area, first, multiple monitoring points are set up in the monitoring area, and then pollution indicators at each monitoring point are collected, including but not limited to the concentration of pollutants in the air. The collected data will be synchronously recorded as a first characteristic parameter for storage, and then the proportion of the first characteristic parameter at each monitoring point will be obtained and recorded as a second characteristic parameter. Based on the second characteristic parameter, the initial weight of each pollution indicator can be further determined, laying the foundation for subsequent weight correction and fusion processing. Then, the initial weight of each pollution indicator is corrected and fused accordingly, so that a more accurate and dynamic weight can be obtained, thereby reflecting the actual impact of the pollution indicator under different conditions;
[0074] Here, pollution indicators include positive indicators and negative indicators. Positive indicators are positively correlated with the concentration of pollution indicators, and negative indicators are negatively correlated with the concentration of pollutants.
[0075] For example, the higher the concentration of pollutants such as sulfur dioxide and nitrogen oxides, the worse the air quality. These are positive indicators, while negative indicators, such as ozone content, generally indicate better air quality as the concentration increases within a certain range. Subsequently, the positive and negative indicators need to be normalized using a preset normalization function, and the normalized result is output as the first characteristic parameter.
[0076] The expression of the normalization function is: ;
[0077] Where, represents the first characteristic parameter, Indicates the The pollutants in The original concentration value at each monitoring point, Indicates the The minimum concentration of a pollutant at all monitoring points, Indicates the The maximum concentration of a pollutant at all monitoring points;
[0078] After normalizing the positive and negative indicators, the proportion of the first characteristic parameter at each monitoring point can be calculated to obtain the second characteristic parameter, which is calculated as follows:
[0079] ;
[0080] Where, Indicates the The pollutants in The proportion of monitoring points, Indicates the total number of monitoring points;
[0081] After the second characteristic parameter is output, the preset initial weight calculation function is introduced to calculate the initial weight of each pollution index, wherein the expression of the initial weight calculation function is:
[0082] ;
[0083] Where, represents the initial weight of the pollution index, Indicates the number of pollutant types;
[0084] After the initial weights of the pollution indicators are output, a reference sequence and a comparison sequence are constructed. The reference sequence contains the standard values of various pollutant concentrations (for details, please refer to the international root values of pollutants), while the comparison sequence contains the measured values of pollutant concentrations. Based on this, a preset correlation coefficient calculation function is introduced to calculate the correlation coefficient of each monitoring point, which is used as the basis for weight correction. The expression of the correlation coefficient calculation function is:
[0085] ;
[0086] Where, Indicates the The correlation coefficient of each monitoring point, represents the resolution coefficient (a constant value, usually 0.5);
[0087] Finally, the correction weights of each pollution index are calculated based on the above correlation coefficients. The calculation formula for the correction weights is:
[0088] ;
[0089] Where, represents the correction weight of pollution index;
[0090] After the correction weights of the pollution indicators are output, they will be dynamically fused to output the final dynamic weights. The fusion formula is:
[0091] ;
[0092] Where, represents the dynamic weight after fusion, is the dominant coefficient (constant value);
[0093] Finally, the dynamic weights of pollution indicators and pollution concentration data at each monitoring point are summarized and calculated to obtain the comprehensive pollution index in the monitoring area. The calculation formula for the comprehensive pollution index is:
[0094] ;
[0095] Where, represents the comprehensive pollution index, represents the measured value of the pollutant concentration, Indicates the standard value of pollutant concentration;
[0096] The comprehensive pollution index can comprehensively reflect the overall pollution status within the monitoring area. The steps of determining the air quality status of the environment within the monitoring area based on the comprehensive pollution index include:
[0097] Obtaining a preset comprehensive assessment threshold, and comparing the comprehensive assessment threshold with the comprehensive pollution index;
[0098] When the comprehensive pollution index is higher than or equal to the assessment threshold, it indicates that the air quality in the monitored area is abnormal and an alarm signal is issued simultaneously;
[0099] When the comprehensive pollution index is lower than the assessment threshold, it indicates that the air quality in the monitoring area is normal, and routine monitoring of the monitoring area will continue;
[0100] Specifically, after the comprehensive pollution index is output, a pre-set comprehensive assessment threshold will be introduced. The comprehensive assessment threshold is obtained based on a comprehensive analysis of relevant environmental protection standards and historical data. The obtained comprehensive assessment threshold will then be carefully compared and analyzed with the currently monitored comprehensive pollution index. In this process, it is necessary to ensure the accuracy and timeliness of the data in order to make correct judgments. When the comprehensive pollution index is higher than or equal to the comprehensive assessment threshold, it indicates that the ambient air quality in the monitoring area is already in an abnormal state, and there may be pollution exceeding the standard or other environmental problems. At this time, an alarm signal will be immediately issued to notify relevant departments and personnel to take corresponding emergency measures to prevent further spread and deterioration of pollution. On the contrary, when the comprehensive pollution index is lower than the assessment threshold, this indicates that the ambient air quality in the monitoring area is normal and meets the requirements of environmental protection standards. At this time, routine monitoring of the monitoring area will continue to be carried out to ensure the stability and controllability of environmental quality and to promptly discover and deal with possible environmental problems.
[0101] S4. Establish a dual-modal early warning mechanism that includes transient anomalies and normalized anomalies, automatically identify the correlation between anomaly types through the correlation strength coefficient, and match corresponding emergency response plans according to different anomaly types.
[0102] In step S4, under abnormal conditions, this embodiment further proposes a dual-modal early warning mechanism that includes transient abnormalities and normalized abnormalities. The dual-modal early warning mechanism automatically determines the correlation between abnormality types through the correlation strength coefficient, and automatically matches and executes corresponding emergency response plans based on different abnormality types, thereby achieving rapid response and effective control of air pollution. The steps of establishing the dual-modal early warning mechanism that includes transient abnormalities and normalized abnormalities include:
[0103] Under abnormal conditions, the time window threshold and volatility threshold are preset, and based on the time window threshold and volatility threshold;
[0104] The pollutant concentration change rate of the pollution indicators in the monitoring area is collected in real time and recorded as the third characteristic parameter. When the third characteristic parameter exceeds the fluctuation rate threshold, the duration of the third characteristic parameter exceeding the fluctuation rate threshold is counted and recorded as the fourth characteristic parameter.
[0105] Compare the fourth characteristic parameter with the time window threshold, and when the fourth characteristic parameter is greater than the time window threshold, record the abnormal state corresponding to the fourth characteristic parameter as a transient abnormality; otherwise, record it as a normalized abnormality;
[0106] Specifically, when establishing a dual-modal early warning mechanism that includes transient anomalies and normalized anomalies, first, in an abnormal state, a time window threshold and a fluctuation threshold are pre-set. The time window threshold and the fluctuation threshold will serve as the prerequisite for subsequent judgment of the abnormal type. The time window threshold is preferably set to 15 to 60 minutes, and the fluctuation threshold is preferably set to 150% to 300%. Then, the pollutant concentration change rate of various pollution indicators in the monitoring area is collected in real time and recorded as the third characteristic parameter. During the monitoring process, once the third characteristic parameter is found to exceed the preset fluctuation threshold, the duration of the third characteristic parameter exceeding the fluctuation threshold is immediately counted and recorded as the fourth characteristic parameter. Subsequently, the fourth characteristic parameter is compared and analyzed with the preset time window threshold. If the fourth characteristic parameter is greater than the time window threshold, the corresponding abnormal state is determined to be a transient anomaly and recorded. Conversely, if the fourth characteristic parameter is less than or equal to the time window threshold, the abnormal state is determined to be a normalized anomaly and recorded accordingly. In this way, different types of abnormal states can be effectively distinguished and recorded, providing accurate data support for subsequent early warning and response measures.
[0107] Secondly, the steps of automatically determining the correlation between abnormal types through the correlation strength coefficient include:
[0108] Construct a spatiotemporal correlation matrix and calculate the linear correlation between each pollutant through the spatiotemporal correlation matrix. Then calculate the nonlinear spatiotemporal weight by the overlap times and time windows between normalized anomalies and transient anomalies.
[0109] By fusing the nonlinear spatiotemporal weight with the linear correlation, the correlation strength coefficient between anomaly types is output;
[0110] According to the preset classification interval, the correlation strength coefficient is compared with the classification interval;
[0111] When the correlation strength coefficient exceeds the upper limit of the classification interval, it indicates that the correlation between the transient anomaly and the normalized anomaly is strong, and a high-risk warning signal is issued;
[0112] When the correlation strength coefficient falls within the classification interval, it indicates that the correlation between transient anomalies and normalized anomalies is moderate, and a normalized risk warning signal is issued;
[0113] When the correlation strength coefficient is lower than the lower limit of the classification interval, it indicates that the correlation between the transient anomaly and the normalized anomaly is weak, and they are independent events, and a low-risk warning signal is issued;
[0114] Specifically, when discriminating the specific abnormal type under the abnormal state, a spatiotemporal correlation matrix is first constructed, and the linear correlation between each pollutant is calculated using the spatiotemporal correlation matrix, which can be specifically calculated using the Pearson correlation coefficient. On this basis, this embodiment further calculates the nonlinear spatiotemporal weight by analyzing the number of overlaps and time windows between normalized anomalies and transient anomalies, and fuses the calculated nonlinear spatiotemporal weight with the previously obtained linear correlation to output the correlation strength coefficient between the anomaly types. The specific calculation formula of the correlation strength coefficient is:
[0115] ;
[0116] Where, represents the correlation strength coefficient, In the spatiotemporal correlation matrix, the instantaneous abnormal intensity sequence The instantaneous abnormal pollutant concentration fluctuation rate, In the normalized abnormal intensity sequence under the spatiotemporal correlation matrix, The pollution index, which is abnormal at all times, exceeds the standard. represents the arithmetic mean of the instantaneous abnormal intensity series, represents the arithmetic mean of the normalized anomaly intensity series, represents the number of spatiotemporal overlaps (the number of times transient anomalies and normalized anomalies occur together within the time window), represents the analysis time window (the observation duration of the correlation calculation), Indicates the abnormal time difference (the difference between the end time of the instantaneous abnormality and the start time of the most recent normalized abnormality);
[0117] Then, according to the pre-set classification interval (the value range of the classification interval is preferably 0.3-0.7), the calculated correlation strength coefficient is compared with the classification interval. When the correlation strength coefficient exceeds the upper limit of the classification interval, it indicates that the correlation between the transient anomaly and the normalized anomaly is very strong. At this time, a high-risk warning signal will be issued, prompting relevant departments to take emergency measures. When the correlation strength coefficient is within the classification interval, it indicates that the correlation between the transient anomaly and the normalized anomaly is moderate. At this time, a normalized risk warning signal will be issued, reminding relevant departments to pay attention and take corresponding preventive measures. When the correlation strength coefficient is lower than the lower limit of the classification interval, it indicates that the correlation between the transient anomaly and the normalized anomaly is weak, and it is an independent event. At this time, a low-risk warning signal will be issued, informing relevant departments to maintain basic monitoring intensity.
[0118] Secondly, the steps of matching the corresponding emergency response plan according to different abnormality types include:
[0119] Obtain the type of abnormal state, and based on the type of abnormal state, retrieve the emergency plan that matches the abnormal state type from the preset emergency response plan library. The emergency plan includes a high-risk emergency response plan and a normalized emergency response plan;
[0120] When a transient anomaly is detected, the correlation strength coefficient between the transient anomaly and the normalized anomaly is further determined, and when the correlation strength coefficient corresponds to a high-risk warning signal, a high-risk emergency response plan is initiated;
[0121] When a moderate correlation is detected between transient anomalies and normalized anomalies, the normalized emergency response plan is directly called, and the air quality in the monitoring area is continuously monitored until the air quality returns to normal.
[0122] Specifically, after the abnormal type in the abnormal state is determined, it is first necessary to obtain the specific type of the current abnormal state, and based on the obtained abnormal state type, retrieve the emergency plan that matches the abnormal state type from the pre-set emergency response plan library. The emergency plan covers two categories, namely high-risk emergency response plans for high-risk situations, and normalized emergency response plans for normalized situations and low-risk situations. When a transient abnormality is detected, it is necessary not only to confirm the existence of the transient abnormality, but also to further analyze and determine the correlation strength coefficient between the transient abnormality and the normalized abnormality. In this process, if the correlation strength coefficient reaches or exceeds the threshold of the high-risk warning signal, the high-risk emergency response will be immediately activated. In addition, the corresponding risk assessment can be made according to the degree of excess of pollutant concentration under normalized anomalies. Specifically, the pollutant concentration under normalized anomalies is compared with the preset standard excess degree. When the excess degree of pollutants is higher than the standard excess degree, the high-risk emergency response plan will also be activated. On the other hand, when the correlation strength between the transient anomaly and the normalized anomaly is detected to be at a moderate level, the normalized emergency response plan will be directly called for processing. At the same time, the air quality in the monitoring area must be continuously and closely monitored to ensure that any new abnormal situation can be discovered and responded to in a timely manner until the air quality is fully restored to normal, thereby ensuring environmental safety.
[0123] See also Figure 2 A big data-based air pollution detection system, using the above-mentioned big data-based air pollution detection method, includes:
[0124] Data acquisition module, which is used to collect multi-source environmental data of the target area through a distributed sensor network, including the concentration values of PM2.5, PM10, SO2, NOx, O3, as well as temperature, humidity, wind speed, and wind direction parameters;
[0125] Preprocessing module, which is used to preprocess multi-source environmental data, including sliding window filtering and ARIMA missing value filling;
[0126] The quality assessment module is used to calculate the comprehensive pollution index in the monitoring area based on the pre-processed multi-source environmental data, and determine the air quality status of the environment in the monitoring area based on the comprehensive pollution index, wherein the air quality status includes normal state and abnormal state;
[0127] The abnormality identification module is used to establish a dual-modal early warning mechanism that includes transient abnormalities and normalized abnormalities. It realizes automatic identification of the correlation between abnormality types through the correlation strength coefficient, and matches the corresponding emergency response plan according to different abnormality types.
[0128] In the above, the main function of the data acquisition module is to comprehensively collect multi-source environmental data in the target monitoring area through a widely deployed distributed sensor network. The specific data types collected cover the concentration values of key pollutants such as fine particulate matter PM2.5, inhalable particulate matter PM10, sulfur dioxide SO2, nitrogen oxides NOx, ozone O3, as well as basic meteorological parameters such as ambient temperature, relative humidity, wind speed, and wind direction, providing a corresponding data basis for subsequent environmental pollution analysis. The role of the preprocessing module is to perform preliminary processing and optimization on the collected multi-source environmental data. The preprocessing process includes using sliding window filtering technology to smooth the data to eliminate the influence of random noise, and using the ARIMA (autoregressive integrated moving average) model to fill in the missing data to ensure the integrity and continuity of the data. The main task of the quality assessment module is based on The pre-processed multi-source environmental data is used to scientifically calculate the comprehensive air pollution index in the monitoring area. By comprehensively analyzing the concentrations of various pollutants and their impact on the environment, a comprehensive index reflecting the overall pollution status of the region is calculated. Based on the comprehensive pollution index, the system further determines the air quality status of the environment in the monitoring area, which is specifically divided into two categories: normal and abnormal states, providing an intuitive reference for environmental management and decision-making. The function of the anomaly discrimination module is to establish a dual-modal early warning mechanism that includes instantaneous anomalies and normalized anomalies. By introducing the correlation strength coefficient, the system can automatically judge the correlation between different anomaly types, thereby accurately identifying abnormal conditions in the environment. For different anomaly types, the system will match the corresponding emergency response plan to ensure that when an environmental pollution incident occurs, response measures can be taken quickly and effectively to minimize environmental pollution.
[0129] See also Figure 3 , an electronic device, the electronic device comprising:
[0130] at least one processor;
[0131] and a memory communicatively coupled to the at least one processor;
[0132] The memory stores a computer program that can be executed by at least one processor, and the computer program is executed by at least one processor so that the at least one processor can execute the above-mentioned air pollution detection method based on big data.
[0133] The processor of the above-mentioned electronic device can be a central processing unit (CPU), a graphics processing unit (GPU) or a digital signal processor (DSP), etc. The processor implements all or part of the steps of the above-mentioned air pollution detection method based on big data by reading and executing a computer program stored in the memory. The memory can be a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory, etc., which is used to store computer programs and data to ensure the normal operation of the electronic device. In addition, the electronic device may also have components such as an arithmetic unit, input and output devices, and a network interface. The arithmetic unit is used to perform various arithmetic and logical operations to ensure the accuracy and efficiency of data processing. Input and output devices, such as a keyboard, a mouse, a display, etc., provide users with an interface for interacting with the electronic device, allowing users to easily input instructions and view processing results. The network interface is used to realize communication between the electronic device and other devices or networks, facilitating data transmission and sharing.
[0134] The foregoing is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art may make various improvements and modifications without departing from the principles of the present invention, and such improvements and modifications are also within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described or explained herein shall, unless otherwise specified or limited, be implemented in accordance with conventional means in the art.
Claims
1. A method for detecting air pollution based on big data, characterized by: include: Collect multi-source environmental data of the target area through a distributed sensor network, including concentrations of PM2.5, PM10, SO2, NOx, O3, as well as temperature, humidity, wind speed, and wind direction parameters; Preprocess multi-source environmental data, including sliding window filtering and ARIMA missing value filling; Calculating a comprehensive pollution index within the monitoring area based on the preprocessed multi-source environmental data, and determining an air quality state of the environment within the monitoring area based on the comprehensive pollution index, wherein the air quality state includes a normal state and an abnormal state; Establish a dual-modal early warning mechanism that includes transient anomalies and normalized anomalies, automatically identify the correlation between anomaly types through the correlation strength coefficient, and match corresponding emergency response plans according to different anomaly types; The step of calculating the comprehensive pollution index in the monitoring area based on the pre-processed multi-source environmental data includes: Multiple monitoring points are set up in the monitoring area, and the pollutant concentration of the pollution index at each monitoring point is collected and simultaneously recorded as the first characteristic parameter; Obtaining the proportion of the first characteristic parameter at each monitoring point and recording it as a second characteristic parameter, and then determining the initial weight of each pollution indicator based on the second characteristic parameter; The initial weights of each pollution index are corrected and integrated to obtain the dynamic weights of each pollution index; The dynamic weights of pollution indicators and pollution concentrations at each monitoring point are aggregated and calculated to obtain a comprehensive pollution index within the monitoring area. The pollution index includes a positive index and a negative index, wherein the positive index is positively correlated with the concentration of the pollution index, and the negative index is negatively correlated with the concentration of the pollutant; Normalizing the positive and negative indicators using a preset normalization function, and outputting the normalized result as a first characteristic parameter; The expression of the normalization function is: ; Where, represents the first characteristic parameter, Indicates the The pollutants in The original concentration value at each monitoring point, Indicates the The minimum concentration of a pollutant at all monitoring points, Indicates the The maximum concentration of a pollutant at all monitoring points; After normalizing the positive and negative indicators, the proportion of the first characteristic parameter at each monitoring point can be calculated to obtain the second characteristic parameter, which is calculated as follows: ; Where, Indicates the The pollutants in The proportion of monitoring points, Indicates the total number of monitoring points; After the second characteristic parameter is output, the preset initial weight calculation function is introduced to calculate the initial weight of each pollution index, wherein the expression of the initial weight calculation function is: ; Where, represents the initial weight of the pollution index, Indicates the number of pollutant types; After the initial weights of the pollution indicators are output, a reference sequence and a comparison sequence are constructed. The reference sequence contains the standard values of various pollutant concentrations (for details, please refer to the international root values of pollutants), while the comparison sequence contains the measured values of pollutant concentrations. Based on this, a preset correlation coefficient calculation function is introduced to calculate the correlation coefficient of each monitoring point, which is used as the basis for weight correction. The expression of the correlation coefficient calculation function is: ; Where, Indicates the The correlation coefficient of each monitoring point, represents the resolution coefficient (a constant value, usually 0.5); Finally, the correction weights of each pollution index are calculated based on the above correlation coefficients. The calculation formula for the correction weights is: ; Where, represents the correction weight of pollution index; After the correction weights of the pollution indicators are output, they will be dynamically fused to output the final dynamic weights. The fusion formula is: ; Where, represents the dynamic weight after fusion, is the dominant coefficient (constant value); Finally, the dynamic weights of pollution indicators and pollution concentration data at each monitoring point are summarized and calculated to obtain the comprehensive pollution index in the monitoring area. The calculation formula for the comprehensive pollution index is: ; Where, represents the comprehensive pollution index, represents the measured value of the pollutant concentration, Indicates the standard value of pollutant concentration.
2. The method for detecting air pollution based on big data according to claim 1, characterized in that: The preprocessing of multi-source environmental data, including sliding window filtering and ARIMA missing value filling, includes: The dynamic threshold method is used to clean the collected multi-source environmental data to remove noise data and outliers; Use sliding window filtering technology to reduce random fluctuations in multi-source environmental data; The ARIMA model is applied to fill missing values, and based on the time series characteristics of multi-source environmental data, missing data points are predicted and filled.
3. The method for detecting air pollution based on big data according to claim 1, characterized in that: The step of determining the air quality status of the environment in the monitoring area based on the comprehensive pollution index includes: Obtaining a preset comprehensive assessment threshold, and comparing the comprehensive assessment threshold with a comprehensive pollution index; When the comprehensive pollution index is higher than or equal to the assessment threshold, it indicates that the air quality in the monitored area is abnormal and an alarm signal is issued simultaneously; When the comprehensive pollution index is lower than the assessment threshold, it indicates that the air quality of the environment in the monitoring area is normal, and routine monitoring of the monitoring area will continue.
4. The method for detecting air pollution based on big data according to claim 1, characterized in that: The steps of establishing a dual-mode early warning mechanism including transient anomalies and normalized anomalies include: Under abnormal conditions, the time window threshold and volatility threshold are preset, and based on the time window threshold and volatility threshold; The pollutant concentration change rate of the pollution indicators in the monitoring area is collected in real time and recorded as the third characteristic parameter. When the third characteristic parameter exceeds the fluctuation rate threshold, the duration of the third characteristic parameter exceeding the fluctuation rate threshold is counted and recorded as the fourth characteristic parameter. The fourth characteristic parameter is compared with the time window threshold, and when the fourth characteristic parameter is greater than the time window threshold, the abnormal state corresponding to the fourth characteristic parameter is recorded as a transient abnormality, otherwise it is recorded as a normalized abnormality.
5. The method for detecting air pollution based on big data according to claim 1, characterized in that: The step of automatically determining the correlation between abnormal types by using the correlation strength coefficient includes: Construct a spatiotemporal correlation matrix and calculate the linear correlation between each pollutant through the spatiotemporal correlation matrix. Then calculate the nonlinear spatiotemporal weight by the overlap times and time windows between normalized anomalies and transient anomalies. By fusing the nonlinear spatiotemporal weight with the linear correlation, the correlation strength coefficient between anomaly types is output; According to the preset classification interval, the correlation strength coefficient is compared with the classification interval; When the correlation strength coefficient exceeds the upper limit of the classification interval, it indicates that the correlation between the transient anomaly and the normalized anomaly is strong, and a high-risk warning signal is issued; When the correlation strength coefficient falls within the classification interval, it indicates that the correlation between the transient anomaly and the normalized anomaly is moderate, and a normalized risk warning signal is issued; When the correlation strength coefficient is lower than the lower limit of the classification interval, it indicates that the correlation between the transient anomaly and the normalized anomaly is weak, and they are independent events, and a low-risk warning signal is issued.
6. The method for detecting air pollution based on big data according to claim 1, characterized in that: The step of matching corresponding emergency response plans according to different abnormality types includes: Obtaining the type of abnormal state, and based on the type of abnormal state, retrieving an emergency plan that matches the type of abnormal state from a preset emergency response plan library, wherein the emergency plan includes a high-risk emergency response plan and a normalized emergency response plan; When a transient anomaly is detected, a correlation strength coefficient between the transient anomaly and the normalized anomaly is continuously determined, and when the correlation strength coefficient corresponds to a high-risk warning signal, a high-risk emergency response plan is initiated; When a moderate correlation is detected between transient anomalies and normalized anomalies, the normalized emergency response plan is directly called, and the air quality in the monitoring area is continuously monitored until the air quality returns to normal.
7. An air pollution detection system based on big data, characterized by: The method for detecting air pollution based on big data according to any one of claims 1 to 6 comprises: A data acquisition module is used to collect multi-source environmental data of the target area through a distributed sensor network, including concentration values of PM2.5, PM10, SO2, NOx, O3, as well as temperature, humidity, wind speed, and wind direction parameters; A preprocessing module is used to preprocess multi-source environmental data, including sliding window filtering and ARIMA missing value filling; A quality assessment module, the quality assessment module is used to calculate the comprehensive pollution index within the monitoring area based on the preprocessed multi-source environmental data, and determine the air quality status of the environment within the monitoring area based on the comprehensive pollution index, wherein the air quality status includes normal state and abnormal state; The abnormality identification module is used to establish a dual-modal early warning mechanism including transient abnormalities and normalized abnormalities, automatically identify the correlation between abnormality types through the correlation strength coefficient, and match the corresponding emergency response plan according to different abnormality types.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; Wherein, the memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the air pollution detection method based on big data as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Atmospheric environment control system based on Internet of Things
CN117851900A
Intelligent monitoring method and system for air environment monitoring station
CN119397458A