An intelligent environmental monitoring method and system based on big data

By adopting methods such as big data processing and information entropy analysis in environmental monitoring technology, the problem of insufficient data processing speed and real-time response in the existing technology is solved, and more efficient and accurate environmental monitoring is achieved, and more effective environmental management and policy decisions are supported.

CN119128768BActive Publication Date: 2025-05-30YANTAI PENGLAI ENVIRONMENTAL MONITORING CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411588224.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-08
Publication Date
2025-05-30
Estimated Expiration
2044-11-08

AI Technical Summary

Technical Problem

Existing environmental monitoring technologies have problems with insufficient processing speed and real-time response when processing large-scale and diversified data, especially lack of flexibility and automation in data verification and format standardization, which affects timeliness and timeliness of decision-making.

Method used

An intelligent environmental monitoring method based on big data is adopted. By collecting real-time environmental monitoring data, verifying and format standardization, information entropy calculation and abnormal identification are performed, non-parametric estimation and trend detection of probability distribution are carried out, and environmental monitoring logs are generated.

Benefits of technology

It improves the processing accuracy and real-time nature of environmental monitoring data, enhances the ability to identify environmental abnormalities and rare events, and supports more effective environmental management and policy decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119128768B_ABST
    Figure CN119128768B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of environmental monitoring, and specifically to an intelligent environmental monitoring method and system based on big data, which includes the following steps: collecting real-time environmental monitoring data, recording indicators of temperature, humidity, and pollutant concentration from different sensors, uploading and updating the data, and obtaining a processed environmental data set through data verification and format standardization. In the present invention, by calculating and analyzing the information entropy, the volatility of each environmental indicator is quantified, effectively identifying anomalies and rare events in the data. The annotation of abnormal indicators, locations, and times further improves the positioning accuracy of environmental anomalies, facilitating rapid response, and helping to accurately identify and predict the trends of environmental changes. Through in-depth processing and analysis of environmental monitoring data, the real-time performance, accuracy, and prediction ability of environmental monitoring are greatly improved, better supporting environmental management and policy decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of environmental monitoring, and particularly to an intelligent environmental monitoring method and system based on big data. Background Art

[0002] The technical field of environmental monitoring aims to track and evaluate changes in various parameters in the environment, covering the whole process from basic data collection to complex analysis and processing. The concerns include air quality, water quality, soil conditions, noise levels, and radiation, etc. The development of environmental monitoring technology can obtain data in real time and predict future environmental changes through advanced algorithms to support the decision-making process of environmental protection and sustainable development. Emerging intelligent environmental monitoring technologies, such as the use of Internet of Things (IoT) devices and big data analysis, are improving the frequency and accuracy of data collection, making environmental monitoring more efficient and comprehensive.

[0003] Among them, the intelligent environmental monitoring method based on big data refers to using big data technology to optimize and improve the efficiency and accuracy of environmental monitoring. By integrating the data collected by environmental monitoring devices and applying machine learning and data analysis technologies, it can not only monitor the environmental conditions in real time, but also predict and simulate the future change trends of the environment. Its main uses include improving the response speed to environmental pollution, more accurately conducting environmental risk assessment, and optimizing resource allocation and environmental protection policy formulation, effectively supporting environmental management and policy decision-making.

[0004] Although the existing environmental monitoring technologies cover the whole process from data collection to analysis and processing, they show deficiencies in processing speed and real-time response when dealing with large-scale and diverse data. Especially in data verification and format standardization, there is a lack of sufficient flexibility and automation, resulting in data processing delays and affecting timeliness and the timeliness of decision-making. The existing technologies rely on traditional statistical methods when predicting future environmental changes, which require assumptions about the data distribution, restricting the accuracy of prediction and the wide application. In terms of anomaly detection, there is a lack of effective tools to quantitatively analyze and label anomalies, making the response to environmental anomalies not timely or accurate enough. For example, pollution events that are not identified and reacted to in time lead to an increase in environmental and public health risks. Therefore, the existing technologies have obvious deficiencies in the real-time performance, accuracy, and prediction ability of data processing, directly affecting the efficiency of environmental monitoring and the formulation of environmental policies. Summary of the Invention

[0005] The purpose of the present invention is to solve the deficiencies existing in the prior art, and to propose an intelligent environmental monitoring method and system based on big data.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions. An intelligent environmental monitoring method based on big data includes the following steps:

[0007] S1: Collect real-time environmental monitoring data, record the indicators of temperature, humidity, and pollutant concentration from different sensors, upload and update the data, and obtain the processed environmental data set through data verification and format standardization;

[0008] S2: Perform information entropy calculation on the indicators in the processed environmental data set, including analyzing the frequency distribution of each environmental indicator and calculating the volatility to obtain the information entropy analysis result;

[0009] S3: Identify the indicators with abnormal entropy values through the information entropy analysis result, mark the indicators with the corresponding time and location, analyze the areas of environmental anomalies and rare events, and generate abnormal indicator records;

[0010] S4: Perform non-parametric estimation of the probability distribution on the data in the abnormal indicator records, count the frequency of each marked data point and estimate the distribution shape to obtain the probability distribution estimation result;

[0011] S5: Utilize the probability distribution estimation result, analyze the change trend of environmental indicators in the time series according to big data analysis, and determine the change of each indicator through trend testing to obtain the trend detection result;

[0012] S6: Based on the trend detection result, calculate the cumulative anomaly value of each environmental indicator, identify the deviation and abnormal changes in the long-term trend through the comparison of the accumulation and average of long-term data, and generate an environmental monitoring log.

[0013] As a further solution of the present invention, the processed environmental data set includes the standardized and verified data of temperature, humidity, and pollutant concentration indicators collected from multiple different sensors. The information entropy analysis result includes the volatility analysis result of the frequency distribution of temperature, humidity, and pollutant concentration, and the information entropy calculation result of each indicator. The abnormal indicator record includes the environmental indicators with abnormal information entropy values, the associated time and location, and the areas of environmental anomalies and rare events. The probability distribution estimation result includes the estimated frequency distribution and distribution shape of the abnormal indicator record data points. The trend detection result includes the test result of the change trend of each environmental indicator determined, the stability of the trend, and the change point. The environmental monitoring log includes the calculated cumulative anomaly value of each environmental indicator, and the deviation and abnormal change conditions in the long-term trend identified through long-term data comparison.

[0014] As a further solution of the present invention, the steps of collecting real-time environmental monitoring data, recording the indicators of temperature, humidity, and pollutant concentration from different sensors, uploading and updating the data, and obtaining the processed environmental data set through data verification and format standardization are specifically as follows:

[0015] S101: Deploy a variety of environmental sensors to monitor the temperature, humidity, and pollutant concentration in the target area. The sensors automatically collect data every minute, and the data is transmitted to the data center in real time via a wireless network to obtain real-time environmental data;

[0016] S102: Use the real-time environmental data to perform data verification, including checking the data integrity, range, and type, removing outliers, and performing data updates once an hour to obtain verified environmental data;

[0017] S103: Based on the verified environmental data, standardize the data format, including unifying the timestamp format, numerical units, and data structure, and checking whether the data meets the standard requirements for analysis and storage to obtain a processed environmental data set.

[0018] As a further solution of the present invention, perform information entropy calculation on the indicators in the processed environmental data set, including analyzing the frequency distribution of each environmental indicator and calculating the volatility. The steps for obtaining the information entropy analysis result are specifically as follows:

[0019] S201: Through the processed environmental data set, extract the data frequency of each environmental indicator, including temperature, humidity, and pollutant concentration, count the occurrence frequency of data points in the data set, and process each indicator separately to obtain an indicator frequency distribution diagram;

[0020] S202: Based on the indicator frequency distribution diagram, calculate the volatility of each environmental indicator. By analyzing the standard deviation and coefficient of variation of data points, determine the fluctuation range of indicator data, perform volatility calculation once a week to monitor the stability of environmental conditions, and obtain the indicator volatility analysis result;

[0021] S203: Use the indicator volatility analysis result to perform information entropy calculation, use the information entropy formula to calculate the volatility of each environmental indicator, identify the randomness and information content of the data, and obtain the information entropy analysis result.

[0022] As a further solution of the present invention, through the information entropy analysis result, identify the indicators with abnormal entropy values, mark the indicators with the corresponding time and location, analyze the areas of environmental anomalies and rare events, and the steps for generating abnormal indicator records are specifically as follows:

[0023] S301: Use the information entropy analysis result, adopt the decision tree algorithm, screen the environmental indicators whose entropy values deviate from the normal range, including temperature, humidity, and pollutant concentration, calculate the degree of deviation of the entropy value of each indicator, and obtain a list of abnormal entropy value indicators;

[0024] S302: Based on the list of abnormal entropy value metrics, associate each abnormal metric with the corresponding timestamp and geographical location information, verify that the occurrence time and location of each abnormal event are recorded, and update the environmental monitoring data in real time according to the abnormal metrics to obtain the associated data of abnormal metrics with time and location.

[0025] S303: According to the associated data of abnormal metrics with time and location, analyze the environmental anomalies and rare events in the target area, including climate change and sudden environmental events, and evaluate the environmental stability of each area to obtain the abnormal metric records.

[0026] As a further solution of the present invention, the formula of the decision tree algorithm is as follows:

[0027] ;

[0028] Wherein, is the weighted average information entropy, represents the probability of the th subset under the splitting attribute , is the information entropy of the th subset, represents the number of values of the attribute , is the adjustment coefficient, represents the average value of the subset size, is the th subset size, is the total number of subsets.

[0029] As a further solution of the present invention, the steps for non-parametric estimation of the probability distribution of the data in the abnormal metric records, statistically counting the frequency of each marked data point and estimating the distribution shape to obtain the probability distribution estimation result are specifically as follows:

[0030] S401: Based on the abnormal metric records, extract the data points marked as abnormal, statistically count the frequency of each data point, record the number of occurrences in the entire data set, and obtain the abnormal data frequency statistical table;

[0031] S402: Use the abnormal data frequency statistical table to estimate the probability distribution of each abnormal data point, without relying on the prior setting of the data distribution, analyze the distribution shape of the data, and obtain the approximate shape of the probability distribution;

[0032] S403: Use the approximate shape of the probability distribution to integrate the distribution estimation results of the abnormal data points, identify and analyze the distribution characteristics of the abnormal environmental metrics, including peaks, skewness, and tail characteristics, to obtain the probability distribution estimation result.

[0033] As a further solution of the present invention, using the probability distribution estimation result, according to the change trend of environmental indicators in the time series of big data analysis, the step of determining the change of each indicator through trend testing to obtain the trend detection result is specifically as follows:

[0034] S501: Based on the probability distribution estimation result, extract the distribution data of environmental indicators, form a time series using the data, and arrange the time series data points of each indicator in the chronological order of real-time occurrence, verify the continuity and integrity of the time series to obtain the environmental indicator time series;

[0035] S502: According to the environmental indicator time series, use trend testing to analyze each time series. The test does not require the data to follow a target distribution, detect the trend change situation in the data set to obtain the initial trend analysis result;

[0036] S503: Use the initial trend analysis result to determine the change trend of each environmental indicator, distinguish the indicators with stable trends, and evaluate the change degree of environmental quality to obtain the trend detection result.

[0037] As a further solution of the present invention, based on the trend detection result, calculate the cumulative anomaly value of each environmental indicator, and identify the deviation and abnormal changes in the long-term trend by comparing the accumulation and average of long-term data. The step of generating the environmental monitoring log is specifically as follows:

[0038] S601: Based on the trend detection result, extract the time series data of each environmental indicator, calculate the difference between each data point and the long-term average value, and continuously track the deviation degree of the environmental indicator to obtain the deviation data table;

[0039] S602: Use the deviation data table to accumulate the deviation values of environmental indicators in the time series, integrate the accumulated values and arrange them in chronological order to generate a cumulative anomaly graph;

[0040] S603: Through the cumulative anomaly graph, compare the cumulative anomaly values at multiple time points with the long-term average cumulative anomaly value, screen the time periods deviating from the normal, and calibrate the abnormal changes according to the deviation to obtain the environmental monitoring log.

[0041] An intelligent environmental monitoring system based on big data, the intelligent environmental monitoring system based on big data is used to execute the above-mentioned intelligent environmental monitoring method based on big data, and the system includes:

[0042] The data collection and processing module collects data of temperature, humidity and pollutant concentration from a variety of environmental sensors, executes data upload, and performs verification and format standardization processing on the original data to obtain a standardized environmental data set;

[0043] The information entropy calculation module conducts a frequency distribution analysis on the environmental indicators in the standardized environmental dataset, calculates the volatility of the indicators to determine the information entropy, and generates an information entropy analysis result;

[0044] The anomaly recognition module identifies the environmental indicators with abnormal information entropy values based on the information entropy analysis result, marks the abnormal indicators associated with the target time and location, and obtains an abnormal indicator record set through regional analysis of abnormal and rare events;

[0045] The data point frequency statistics module conducts a non-parametric estimation of the probability distribution of the data in the abnormal indicator record set, counts the frequency of data points, estimates the distribution shape, and generates a probability distribution estimation result;

[0046] The trend analysis module uses the probability distribution estimation result to analyze the change trend of environmental indicators in the big data time series, conducts a trend test to verify the change of indicators, calculates the cumulative anomaly value of environmental indicators, identifies the deviation in the long-term trend, and generates an environmental monitoring log.

[0047] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0048] In the present invention, by collecting real-time environmental monitoring data and uploading and updating it, validating the data and standardizing the format, it is ensured that the processed dataset has high accuracy and reliability. The calculation and analysis of information entropy quantify the volatility of each environmental indicator, and the quantitative analysis method can effectively identify abnormal and rare events in the data. The annotation of abnormal indicators, locations, and times further improves the positioning accuracy of environmental anomalies, facilitating a rapid response. The non-parametric probability distribution estimation enables accurate counting of the frequency and distribution shape of each data point without assuming the specific distribution form of the data, providing a deeper understanding of the fluctuation pattern of environmental indicators. The introduction of trend testing, combined with the comparison of long-term data cumulative values, helps to accurately identify and predict the trend of environmental changes, providing data support for long-term environmental protection and policy formulation. Through in-depth processing and analysis of environmental monitoring data, the real-time performance, accuracy, and prediction ability of environmental monitoring are greatly improved, better supporting environmental management and policy decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is a schematic diagram of the working process of the present invention;

[0050] Figure 2 is a detailed flowchart of S1 of the present invention;

[0051] Figure 3 is a detailed flowchart of S2 of the present invention;

[0052] Figure 4 is a detailed flowchart of S3 of the present invention;

[0053] Figure 5 It is the detailed flowchart of S4 of the present invention;

[0054] Figure 6 It is the detailed flowchart of S5 of the present invention;

[0055] Figure 7 It is the detailed flowchart of S6 of the present invention;

[0056] Figure 8 It is the system flowchart of the present invention. Specific embodiments

[0057] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0058] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, in the description of the present invention, the meaning of "a plurality" is two or more, unless otherwise specifically defined.

[0059] Please refer to Figure 1 , the present invention provides a technical solution, an intelligent environmental monitoring method based on big data, including the following steps:

[0060] S1: Collect real-time environmental monitoring data, record the indicators of temperature, humidity, and pollutant concentration from different sensors, upload and update the data, and obtain a processed environmental data set through data verification and format standardization;

[0061] S2: Perform information entropy calculation on the indicators in the processed environmental data set, including analyzing the frequency distribution of each environmental indicator and calculating the volatility to obtain the information entropy analysis result;

[0062] S3: Identify the indicators with abnormal entropy values through the information entropy analysis result, mark the indicators with the corresponding time and location, analyze the areas of environmental anomalies and rare events, and generate abnormal indicator records;

[0063] S4: Perform non-parametric estimation of the probability distribution on the data in the abnormal index record, count the frequency of each marked data point and estimate the distribution shape to obtain the probability distribution estimation result;

[0064] S5: Utilize the probability distribution estimation result to analyze the change trend of environmental indicators in the time series according to big data analysis, and determine the change situation of each indicator through trend testing to obtain the trend detection result;

[0065] S6: Based on the trend detection result, calculate the cumulative anomaly value of each environmental indicator, and identify the deviations and abnormal changes in the long-term trend by comparing the accumulation and average of long-term data to generate the environmental monitoring log.

[0066] The processed environmental data set includes the standardized and verified data of temperature, humidity, and pollutant concentration indicators collected from multiple differentiated sensors. The information entropy analysis result includes the volatility analysis result of the frequency distribution of temperature, humidity, and pollutant concentration, and the information entropy calculation result of each indicator. The abnormal index record includes the environmental indicators with abnormal information entropy values, the associated time and location, and the regions of environmental anomalies and rare events. The probability distribution estimation result includes the estimated frequency distribution and distribution shape of the abnormal index record data points. The trend detection result includes the test results of the change trend of each environmental indicator determined, the stability of the trend, and the change points. The environmental monitoring log includes the calculated cumulative anomaly value of each environmental indicator, and the deviations and abnormal change situations in the long-term trend identified by comparing long-term data.

[0067] Please refer to Figure 2 , for the specific steps of collecting real-time environmental monitoring data, recording the indicators of temperature, humidity, and pollutant concentration from differentiated sensors, uploading and updating the data, and obtaining the processed environmental data set through data verification and format standardization are as follows:

[0068] S101: Deploy a variety of environmental sensors to monitor the temperature, humidity, and pollutant concentration in the target area. The sensors automatically collect data every minute, and the data is transmitted to the data center in real time through the wireless network. The execution process of obtaining the real-time environmental data is as follows;

[0069] Deploy a variety of environmental sensors in the target area to monitor temperature, humidity, and pollutant concentration. The sensor settings need to precisely match the environmental characteristics to ensure the validity and accuracy of the data. The selection of sensors includes temperature and humidity sensors and air quality monitoring devices. The devices automatically collect data every minute. The data collection involves the timed activation and data reading process of the sensors to ensure that the data per minute can reflect the environmental conditions at that moment. The collected data is transmitted to the data center in real time through a wireless network. During the data transmission process, it is necessary to ensure the stability and transmission speed of the network to prevent data loss or errors during transmission. After the data is successfully transmitted to the data center, real-time environmental data is obtained.

[0070] S102: Use the real-time environmental data to perform data verification, including checking the data integrity, range, and type, removing outliers, and performing data updates once an hour. The execution process of obtaining the verified environmental data is as follows;

[0071] Use the real-time environmental data to perform data verification, which involves checking the data integrity, range, and type. During the verification process, systematic screening of the data is carried out to ensure that all data is within the predetermined reasonable range. The data verification process includes the automatic identification and removal of outliers. For data outside the range, the system will mark and exclude it to ensure that it is not included in subsequent data analysis. The process is performed once an hour to ensure the accuracy and reliability of each batch of data. The data update process includes data cleaning, updating error records, and refreshing the database operation to obtain the verified environmental data.

[0072] S103: Based on the verified environmental data, perform data format standardization, including unifying the timestamp format, numerical units, and data structure, and checking that the data meets the standard requirements for analysis and storage. The execution process of obtaining the processed environmental data set is as follows;

[0073] Based on the verified environmental data, perform data format standardization, according to the formula Calculate the average value of each environmental data. In the formula, represents the average value of the formatted data, represents a single data point, represents the total number of data points. Explanation of the formula and the derivation process of the formula calculation: Consider a simple example. If the temperature values of five data points are 20°C, 22°C, 21°C, 23°C, and 19°C respectively. Then the average value can be calculated by inserting into the formula as follows. The specific steps are as follows: . The calculation process shows how to convert the actual collected data values into average values, which are then used for data standardization processing. The process not only involves data averaging but also includes unifying the timestamp format and organizing the data structure to ensure that the data meets the standard requirements for analysis and storage.

[0074] Please refer to Figure 3 to perform information entropy calculation on the indicators in the processed environmental dataset, including analyzing the frequency distribution of each environmental indicator and calculating the volatility, and the steps to obtain the information entropy analysis result are specifically as follows:

[0075] S201: Through the processed environmental dataset, extract the data frequency of each environmental indicator, including temperature, humidity, and pollutant concentration, count the occurrence frequency of data points in the dataset, and process each indicator separately. The execution process to obtain the indicator frequency distribution diagram is as follows;

[0076] Through the processed environmental dataset, extract the data frequency of each environmental indicator, which involves data processing of indicators such as temperature, humidity, and pollutant concentration. The processing includes counting the occurrence frequency of each indicator data point in the entire dataset. The frequency is calculated by dividing the occurrence times of each data point by the total number of data points. The frequency statistical method ensures that the processing of each indicator is independent and accurately reflects the normal distribution of each environmental factor, and obtains the frequency distribution diagram.

[0077] S202: Based on the indicator frequency distribution diagram, calculate the volatility of each environmental indicator, determine the fluctuation range of the indicator data by analyzing the standard deviation and coefficient of variation, perform the volatility calculation once a week, and monitor the stability of environmental conditions. The execution process to obtain the indicator volatility analysis result is as follows;

[0078] Based on the indicator frequency distribution diagram, calculate the volatility of each environmental indicator, and calculate the standard deviation and coefficient of variation according to the formulas and . In the formula, represents the standard deviation, represents a single data point, represents the average value, represents the total number of data points, represents the coefficient of variation. Formula details and formula calculation derivation process: Set a data point set of an environmental indicator as [18, 20, 22, 20, 19], calculate the average value of the data: , insert the value of each data point into the formula of the standard deviation for calculation: , calculate the coefficient of variation: . The calculation process shows that the volatility of the indicator is quantitatively expressed through the standard deviation and the coefficient of variation. The calculation of volatility helps to determine the stability of environmental conditions, and the volatility calculation performed once a week provides data support for monitoring environmental conditions.

[0079] S203: Using the analysis result of index volatility, perform information entropy calculation. Use the information entropy formula to calculate the volatility of each environmental index, identify the randomness and information content of the data. The execution process for obtaining the information entropy analysis result is as follows;

[0080] Using the analysis result of index volatility, perform information entropy calculation. Information entropy is a method to measure the randomness and information content of data. Use the information entropy formula to calculate the volatility of each environmental index, which involves substituting the probability distribution of each index into the information entropy formula. During the calculation process, the calculation formula for information entropy is , where represents the probability of the index value occurrence. The information entropy calculation for each environmental index is performed independently and is not affected by the remaining indices. The calculation result of information entropy can identify and quantify the volatility and complexity of the data, providing important reference data for the environmental monitoring and early warning system to obtain the information entropy analysis result.

[0081] Please refer to Figure 4 , through the information entropy analysis result, identify the indices with abnormal entropy values, mark the indices with the corresponding time and location, and analyze the areas of environmental anomalies and rare events. The specific steps for generating the abnormal index record are as follows:

[0082] S301: Using the information entropy analysis result, adopt the decision tree algorithm to screen the environmental indices whose entropy values deviate from the normal range, including temperature, humidity, and pollutant concentration. Calculate the degree of deviation of the entropy value of each index to obtain the execution process of the abnormal entropy value index list as follows;

[0083] The formula for the decision tree algorithm is as follows:

[0084] ;

[0085] where is the weighted average information entropy, represents the probability of the th subset under the splitting attribute , is the information entropy of the th subset, represents the number of values of the attribute , is the adjustment coefficient, represents the average value of the subset size, is the th subset size, is the total number of subsets.

[0086] To demonstrate the specific application and calculation process of the formula, consider a simple environmental monitoring data set, including temperature, humidity, and pollutant concentration, which is divided into three subsets. The specific parameters and data are from on-site measurements.

[0087] Probability of subsets : Represents the proportion of each subset in the total dataset. For example, if the three subsets are evenly distributed, the probability of each subset .

[0088] Information entropy : The information entropy of each subset is calculated based on the volatility of the content. For example, assume that the information entropies of the subsets calculated are 0.8, 0.5, and 0.9 respectively.

[0089] Attribute Number of value : The number of values of the attribute. For example, temperature, humidity, and pollutant concentration each have 3 different states, so .

[0090] Adjustment coefficient : Determined through the study of the volatility of actual data, set to 0.5 to balance information gain and data volatility.

[0091] Subset size and average subset size : If each subset has an equal number of data points. For example, if the size of each subset is 100, then .

[0092] Substitute actual values for calculation: Set the information entropies of the three subsets to be 0.8, 0.5, and 0.9 respectively.

[0093] ;

[0094] The results show that under the current environmental monitoring settings, the weighted average information entropy of the attributes (i.e., temperature, humidity, pollutant concentration) is 0.7333. The value reflects the overall volatility of the data. The correlation lies in determining the overall entropy of the dataset under given environmental conditions, which can be used for the classification and evaluation of environmental quality.

[0095] S302: Based on the list of abnormal entropy value indicators, associate each abnormal indicator with the corresponding timestamp and geographical location information, verify that the occurrence time and location of each abnormal event are recorded, and update the environmental monitoring data in real time according to the abnormal indicators. The execution process of obtaining the time-location associated data of abnormal indicators is as follows;

[0096] Based on the list of abnormal entropy value indicators, verify the time and location of each abnormal event, and associate each abnormal indicator with the timestamp and geographical location according to the formula . In the formula, represents the comprehensive entropy value of the abnormal event, represents the timestamp of the event, The encoded value representing geographical location information, represents the number of events. Detailed formula explanation and formula calculation derivation process: Consider there are three abnormal events with timestamps (in minutes starting from the day), and the geographical location encoding is . Substitute the values into the formula for calculation: =(158· )+(162· )+(170· ) 158·1.099 + 162·1.609 + 170·1.386 173.662 + 260.658 + 235.62 = 669.94. This result indicates that by correlating the timestamp and geographical location information, a comprehensive entropy value can be calculated, which represents the complexity and volatility of the spatio-temporal distribution of abnormal events and provides a quantitative basis for further data analysis.

[0097] S303: According to the time-location associated data of abnormal indicators, analyze the environmental anomalies and rare events in the target area, including climate change and sudden environmental events, and evaluate the environmental stability of each area. The execution process of obtaining the abnormal indicator records is as follows;

[0098] According to the time-location associated data of abnormal indicators, analyze the environmental anomalies and rare events in the target area, including evaluating the anomalies caused by climate change and sudden environmental events. During the analysis process, it involves comparing the time-location associated data with the known environmental stability records, identifying the data points significantly different from the normal state, and ensuring that the data of each area reflects the latest environmental state by real-time updating the environmental monitoring data. The evaluation process includes a comprehensive analysis of the volatility and frequency distribution of environmental indicators to obtain the abnormal indicator records.

[0099] Please refer to Figure 5 to perform a non-parametric estimation of the probability distribution of the data in the abnormal indicator records, count the frequency of each marked data point and estimate the distribution shape. The specific steps to obtain the probability distribution estimation result are as follows:

[0100] S401: Based on the abnormal indicator records, extract the data points marked as abnormal, perform a frequency count on each data point, and record the number of occurrences in the entire dataset. The execution process of obtaining the abnormal data frequency statistical table is as follows;

[0101] Based on the abnormal index records, extract the data points marked as abnormal, and conduct frequency statistics on the data points, which involves calculating the number of occurrences of each data point in the entire dataset. The frequency statistics method checks the status of each data point in the dataset one by one, records the points marked as abnormal, and the creation of the statistical table not only depends on the identification and marking process of the data points but also requires accurately calculating the number of occurrences of each abnormal point. This process ensures the accuracy and reliability of the statistical results, providing the basic data for subsequent analysis. The data point frequency reveals the occurrence frequency of abnormal events, providing a basis for in-depth analysis of the causes and impacts of abnormalities, and obtaining the abnormal data frequency statistical table.

[0102] S402: Use the abnormal data frequency statistical table to estimate the probability distribution of each abnormal data point, without relying on the prior setting of the data distribution. Analyze the distribution shape of the data, and the execution process for obtaining the approximate shape of the probability distribution is as follows;

[0103] Use the abnormal data frequency statistical table to estimate the probability distribution of each abnormal data point, according to the formula Calculate the probability of the data point occurrence. In the formula, represents the occurrence probability of the data point , represents the number of occurrences of the data point , represents the total number of data points. Explanation of the formula and the derivation process of the formula calculation: Consider a simple example, set that the abnormal data point A appears 15 times in the dataset, and the total number of data points in the dataset is 500. Substitute the values into the probability formula: . The calculation shows that the occurrence probability of the data point is 3%. The probability estimation is based on the actual number of occurrences of the data point, rather than relying on any assumption of the prior data distribution. It provides a basis for further statistical analysis, enabling the analysis of the actual distribution shape of the data.

[0104] S403: Use the approximate shape of the probability distribution to integrate the distribution estimation results of the abnormal data points, identify and analyze the distribution characteristics of the abnormal environmental indicators, including peak value, skewness, and tail characteristics, and the execution process for obtaining the probability distribution estimation result is as follows;

[0105] Use the approximate shape of the probability distribution to integrate the distribution estimation results of the abnormal data points. The process includes identifying and analyzing the distribution characteristics of the abnormal environmental indicators. The analysis of the distribution characteristics focuses on the peak value, skewness, and tail characteristics of the abnormal data. These characteristics are obtained through statistical analysis, based on the actual data point distribution, not just the theoretical model. By observing the actual distribution of the data, we can more accurately understand the nature of the abnormal events and the environmental impacts, and obtain the probability distribution estimation result.

[0106] Please refer to Figure 6, using the probability distribution estimation result, according to the big data analysis of the change trend of environmental indicators in the time series, and determining the change situation of each indicator through trend testing, the steps to obtain the trend detection result are specifically as follows:

[0107] S501: Based on the probability distribution estimation result, extract the distribution data of environmental indicators, and use the data to form a time series. The time series data points of each indicator are arranged in the chronological order of real-time occurrence. Verify the continuity and integrity of the time series. The execution process of obtaining the environmental indicator time series is as follows;

[0108] Based on the probability distribution estimation result, extract the distribution data of environmental indicators, and use the data to form a time series. The creation of the time series involves arranging the data points of each indicator in the chronological order of real-time occurrence. The process ensures the continuity and integrity of the time series. Verifying the continuity of the time series is by checking whether there are missing or inconsistent intervals between timestamps. Verifying continuity is a key step to ensure the accuracy of time series analysis. The integrity of the time series is ensured by completely recording all data points from the start of observation to the current. Obtain the environmental indicator time series.

[0109] S502: According to the environmental indicator time series, use trend testing to analyze each time series. The test does not require the data to follow a target distribution, and detect the trend change situation in the dataset. The execution process of obtaining the initial trend analysis result is as follows;

[0110] According to the environmental indicator time series, use trend testing to analyze each time series, according to the formula Test the trend change in the dataset. In the formula, T represents the trend value, represents the data points in the time series, represents the corresponding time mark, represents the number of data points. Formula details and formula calculation derivation process: Set a time series data point set as [10, 12, 13, 14, 15], associated with the time marks [1, 2, 3, 4, 5]. Substitute the values into the trend testing formula: . The calculation shows the square root of the product sum of the data points and time marks in the time series divided by the sum of the squares of the time marks. This value provides a measure of the trend change in the dataset and can determine the growth or decline trend of the data.

[0111] S503: Using the initial trend analysis result, determine the change trend of each environmental indicator, distinguish the indicators with stable trends, and evaluate the degree of change in environmental quality. The execution process of obtaining the trend detection result is as follows;

[0112] Using the initial trend analysis results, determine the change trend of each environmental indicator. During the process, identify the indicators with stable trends. Analyze the results based on time series data and trend tests. The indicators with stable trends are evaluated by the magnitude and direction of the trend values. Evaluating the degree of change in environmental quality involves explaining the relationship between the trend test results and the actual changes in environmental indicators, revealing the stability and trend of each environmental indicator over time, providing important information for environmental management and policy making, helping to guide future environmental protection measures, and obtaining the trend detection results.

[0113] Please refer to Figure 7 , based on the trend detection results, calculate the cumulative anomaly value of each environmental indicator. By comparing the accumulation and average of long-term data, identify the deviations and abnormal changes in the long-term trend. The steps to generate the environmental monitoring log are as follows:

[0114] S601: Based on the trend detection results, extract the time series data of each environmental indicator, calculate the difference between each data point and the long-term average value, and continuously track the deviation degree of the environmental indicator. The execution process to obtain the deviation data table is as follows;

[0115] Based on the trend detection results, extract the time series data of each environmental indicator, calculate the difference between each data point and the long-term average value, according to the formula Continuously track the deviation degree of the environmental indicator.

[0116] In the formula, represents the data point and the long-term average value The difference. Formula details and formula calculation derivation process:

[0117] Set the time series data points as [15, 18, 20, 14, 19], and the long-term average value is 17. Insert the data to calculate the deviation value of each point: , , , , . This calculation process shows that the deviation degree of each data point is calculated and recorded. This continuous tracking helps to identify the abnormal changes in environmental indicators, and generating the deviation data table provides important basic data for further analysis.

[0118] S602: Using the deviation data table, accumulate the deviation values of the environmental indicators in time series, integrate the accumulated values and arrange them in chronological order. The execution process to generate the cumulative anomaly graph is as follows;

[0119] Using a deviation data table, the deviation values of environmental indicators are accumulated over time. During the process, the accumulated values are integrated and arranged in chronological order. The accumulation process involves summing up consecutive deviation values, and the deviation value at each new time point is added to the previous total accumulated value. Such accumulation shows the dynamic change of deviation over time, clearly demonstrating the cumulative deviation trend of data points, which is particularly important for environmental monitoring and trend analysis. It provides a visual tool to observe and compare the changes in environmental data over the long term, generating a cumulative anomaly map.

[0120] S603: Through the cumulative anomaly map, compare the cumulative anomaly values at multiple time points with the long-term average cumulative anomaly value, screen out the time periods that deviate from the norm, and calibrate the abnormal changes according to the deviation to obtain the execution process of the environmental monitoring log as follows;

[0121] Through the cumulative anomaly map, compare the cumulative anomaly values at multiple time points with the long-term average cumulative anomaly value. The comparison process involves screening out the time periods that deviate from the norm. For the time periods, calibrate the abnormal changes according to the degree of deviation, which not only helps to identify when the data points deviate from the normal state but also reveals the trend and severity of the deviation. By detailed recording of the environmental quality changes, it provides important information for understanding the patterns and causes of environmental changes, contributing to formulating response measures and long-term environmental protection strategies to obtain the environmental monitoring log.

[0122] Please refer to Figure 8 , an intelligent environmental monitoring system based on big data. The intelligent environmental monitoring system based on big data is used to execute the above-mentioned intelligent environmental monitoring method based on big data. The system includes:

[0123] The data collection and processing module collects data on temperature, humidity, and pollutant concentration from various environmental sensors, performs data upload, and verifies and standardizes the format of the raw data to obtain a standardized environmental data set;

[0124] The information entropy calculation module conducts frequency distribution analysis on the environmental indicators in the standardized environmental data set, calculates the volatility of the indicators to determine the information entropy, and generates an information entropy analysis result;

[0125] The anomaly recognition module identifies the environmental indicators with abnormal information entropy values according to the information entropy analysis result, marks the abnormal indicators associated with the target time and location, and obtains an abnormal indicator record set through regional analysis of abnormal and rare events;

[0126] The data point frequency statistics module conducts non-parametric estimation of the probability distribution of the data in the abnormal indicator record set, counts the frequency of the data points, estimates the distribution shape, and generates a probability distribution estimation result;

[0127] The trend analysis module uses the probability distribution estimation results to analyze the change trend of environmental indicators in the big data time series, conducts trend tests to verify the changes of the indicators, calculates the cumulative anomaly value of the environmental indicators, identifies the deviations in the long-term trend, and generates environmental monitoring logs.

[0128] The above are only the preferred embodiments of the present invention, and the present invention is not limited to other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. An intelligent environmental monitoring method based on big data, characterized in that: The following steps are involved: Collect real-time environmental monitoring data, record indicators of temperature, humidity, and pollutant concentration from differentiated sensors, upload and update the data, and obtain processed environmental data sets through data verification and format standardization; Performing information entropy calculation on the indicators in the processed environmental data set, including analyzing the frequency distribution of each environmental indicator and calculating volatility, and obtaining information entropy analysis results; Through the information entropy analysis results, identify indicators with abnormal entropy values, mark the indicators with corresponding time and place, analyze the areas of environmental anomalies and rare events, and generate abnormal indicator records; Performing non-parametric estimation of probability distribution on the data in the abnormal indicator record, counting the frequency of each marked data point and estimating the distribution shape, to obtain a probability distribution estimation result; Using the probability distribution estimation result, analyzing the change trend of environmental indicators in the time series according to big data, determining the change of each indicator through trend testing, and obtaining the trend detection result; Based on the trend detection results, the cumulative deviation value of each environmental indicator is calculated, and the deviation and abnormal changes in the long-term trend are identified by comparing the accumulation and average of long-term data, and an environmental monitoring log is generated; Based on the abnormal index record, extract the data points marked as abnormal, perform frequency statistics on each data point, record the number of occurrences in the entire data set, and obtain an abnormal data frequency statistics table; The abnormal data frequency statistics table is used to estimate the probability distribution of each abnormal data point, without relying on the prior setting of the data distribution, and the distribution shape of the data is analyzed to obtain the approximate shape of the probability distribution; Using the approximate shape of the probability distribution, integrating the distribution estimation results of the abnormal data points, identifying and analyzing the distribution characteristics of the abnormal environment indicators, including peak, skewness and tail characteristics, and obtaining the probability distribution estimation results; Based on the probability distribution estimation result, the distribution data of the environmental indicators are extracted, and the data are used to form a time series. The time series data points of each indicator are arranged in the order of real-time occurrence, and the continuity and integrity of the time series are verified to obtain the environmental indicator time series; According to the environmental indicator time series, each time series is analyzed using a trend test. The test does not require the data to follow the target distribution. The trend changes in the data set are detected to obtain an initial trend analysis result. Using the initialization trend analysis results, determine the change trend of each environmental indicator, distinguish indicators with stable trends, evaluate the degree of change in environmental quality, and obtain trend detection results; Extract the data frequency of each environmental indicator from the processed environmental data set, including temperature, humidity and pollutant concentration, calculate the frequency of occurrence of statistical points in the data set, process each indicator separately, and obtain an indicator frequency distribution diagram; Based on the indicator frequency distribution diagram, the volatility of each environmental indicator is calculated, and the fluctuation range of the indicator data is determined by analyzing the standard deviation and coefficient of variation of the data points. The volatility calculation is performed once a week to monitor the stability of the environmental conditions and obtain the indicator volatility analysis results; Calculate the standard deviation and coefficient of variation according to the formula: ; and ; In the formula, represents the standard deviation, represents a single data point, represents the average value, represents the total number of data points, represents the coefficient of variation; The information entropy is calculated using the indicator volatility analysis results, and the volatility of each environmental indicator is calculated using the information entropy formula to identify the randomness and information content of the data to obtain the information entropy analysis results.

2. The intelligent environment monitoring method based on big data according to claim 1 is characterized in that: The processed environmental data set includes standardized and verified data of air temperature, humidity and pollutant concentration indicators collected from multiple differentiated sensors, the information entropy analysis results include volatility analysis results of the frequency distribution of air temperature, humidity and pollutant concentration, and information entropy calculation results of each indicator, the abnormal indicator records include environmental indicators with abnormal information entropy values, associated time and place, and areas of environmental anomalies and rare events, the probability distribution estimation results include the estimated frequency distribution and distribution shape of abnormal indicator record data points, the trend detection results include the test results of the determined change trend of each environmental indicator, the stability of the trend and the change points, and the environmental monitoring log includes the calculated cumulative deviation value of each environmental indicator, deviations and abnormal changes in long-term trends identified by long-term data comparison.

3. The intelligent environment monitoring method based on big data according to claim 1 is characterized in that: The steps to collect real-time environmental monitoring data, record the indicators of temperature, humidity, and pollutant concentration from differentiated sensors, upload and update the data, and obtain the processed environmental data set through data verification and format standardization are as follows: Deploy a variety of environmental sensors to monitor the temperature, humidity and pollutant concentration in the target area. The sensors automatically collect data every minute, and the data is transmitted to the data center in real time via wireless networks to obtain real-time environmental data; Using the real-time environmental data, perform data verification, including checking the data integrity, scope and type, removing abnormal values, and performing data updates once an hour to obtain verified environmental data; Based on the verified environmental data, data format standardization is performed, including unified timestamp format, numerical units and data structure, and the data is checked to meet standard requirements for analysis and storage to obtain a processed environmental data set.

4. The intelligent environment monitoring method based on big data according to claim 1 is characterized in that: Through the information entropy analysis results, the indicators with abnormal entropy values ​​are identified, the indicators are marked with corresponding time and place, the areas of environmental anomalies and rare events are analyzed, and the steps of generating abnormal indicator records are specifically as follows: Using the information entropy analysis results and a decision tree algorithm, environmental indicators whose entropy values ​​deviate from the normal range are screened, including temperature, humidity, and pollutant concentration, and the degree of entropy deviation of each indicator is calculated to obtain a list of abnormal entropy value indicators; Based on the abnormal entropy value indicator list, each abnormal indicator is associated with the corresponding timestamp and geographic location information, and the time and location of each abnormal event are verified to be recorded. The environmental monitoring data is updated in real time according to the abnormal indicator to obtain the abnormal indicator time and location association data; According to the time and place correlation data of the abnormal indicators, the environmental anomalies and rare events in the target area, including climate change and sudden environmental events, are analyzed, the environmental stability of each area is evaluated, and the abnormal indicator records are obtained.

5. The intelligent environment monitoring method based on big data according to claim 4 is characterized in that: The formula of the decision tree algorithm is as follows: ; in, is the weighted average information entropy, Represents the split attribute Next The probability of a subset, It is The information entropy of the subsets is Representation attributes The number of values ​​of is the adjustment factor, represents the average size of the subsets, It is The size of the subset, is the total number of subsets.

6. The intelligent environment monitoring method based on big data according to claim 1 is characterized in that: Based on the trend detection results, the cumulative deviation value of each environmental indicator is calculated, and the deviation and abnormal changes in the long-term trend are identified by comparing the accumulation and average of long-term data. The steps of generating an environmental monitoring log are specifically as follows: Based on the trend detection results, extract the time series data of each environmental indicator, calculate the difference between each data point and the long-term average value, and continuously track the deviation degree of the environmental indicator to obtain a deviation data table; Using the deviation data table, the deviation values ​​of the environmental indicators are accumulated in time series, the accumulated values ​​are integrated and arranged in time sequence, and a cumulative deviation graph is generated; By comparing the cumulative anomaly values ​​at multiple time points with the long-term average cumulative anomaly value through the cumulative anomaly diagram, the time periods that deviate from the norm are screened, and the abnormal changes are calibrated according to the deviations to obtain the environmental monitoring log.

7. An intelligent environment monitoring system based on big data, characterized in that: According to any one of claims 1 to 6, the intelligent environment monitoring method based on big data comprises: The data collection and processing module collects data on temperature, humidity, and pollutant concentration from a variety of environmental sensors, performs data upload, verifies and standardizes the format of the raw data, and obtains a standardized environmental data set; The information entropy calculation module performs frequency distribution analysis on the environmental indicators in the standardized environmental data set, calculates the volatility of the indicators to determine the information entropy, and generates information entropy analysis results; The anomaly identification module identifies environmental indicators with abnormal information entropy values ​​based on the information entropy analysis results, marks abnormal indicators associated with target time and location, and obtains an abnormal indicator record set through regional analysis of abnormal and rare events; The data point frequency statistics module performs non-parametric estimation of the probability distribution of the data in the abnormal indicator record set, calculates the frequency of the data points, estimates the distribution shape, and generates a probability distribution estimation result; The trend analysis module uses the probability distribution estimation results to analyze the changing trends of environmental indicators in the big data time series, conducts trend tests to verify the changes in indicators, calculates the cumulative deviations of environmental indicators, identifies deviations in long-term trends, and generates environmental monitoring logs.

Citation Information

Patent Citations

  • Remote monitoring method and system for electric appliance

    CN117216481A

  • Pollution source online monitoring data quality monitoring method

    CN118503667A