Regional environmental pollution state evaluation method and device based on Internet of Things, terminal equipment and storage medium
By acquiring multi-dimensional environmental data through IoT technology and utilizing feature importance analysis and neural network models, the problem of strong subjectivity in manual assessment has been solved, enabling accurate assessment and timely early warning of regional environmental pollution status, and improving the accuracy and timeliness of the assessment.
Patent Information
- Application Number
- CN202511265951.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-12-23
AI Technical Summary
In existing technologies, the assessment of regional environmental pollution status relies on manual monitoring and experience-based predictions, resulting in low accuracy and poor timeliness of the results. This makes it difficult to quickly take effective containment measures, leading to delays in prevention and control measures and increased costs.
By adopting an Internet of Things (IoT) approach, soil, water, and vegetation data are acquired, and a pollution status prediction model is constructed using a feature importance analysis model and iterative training of a neural network. Combined with historical monitoring indicators and environmental covariates, a pollution status assessment index is calculated to achieve accurate assessment.
It improves the objectivity and timeliness of the assessment, can accurately output the prediction range of soil erosion and non-point source pollution, reduce subjective bias, and provide timely pollution risk warnings.
Smart Images

Figure CN121189620A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of Internet of Things, and in particular to a regional environmental pollution state evaluation method and device based on Internet of Things, a terminal device and a storage medium. BACKGROUND
[0002] In the traditional environmental pollution prevention and control work, the regional environmental pollution state evaluation is mainly through artificial monitoring and experience prediction. Specifically, artificial monitoring is mainly through periodic on-site sampling and naked eye observation, and experience prediction is a subjective judgment of the regional environmental pollution state through the experience accumulated by the staff in long-term work.
[0003] Affected by human operation errors and other factors, the final state evaluation result is not only not accurate, but also has poor timeliness, which directly leads to significant lag of prevention and control measures, making it difficult to take effective action to curb the intensification of soil erosion and the spread of non-point source pollution in the embryonic stage, and greatly increasing the cost and difficulty of subsequent environmental pollution control. SUMMARY
[0004] The present application provides a regional environmental pollution state evaluation method and device based on Internet of Things, a terminal device and a storage medium, which can solve the problem of strong subjectivity caused by artificial evaluation of regional environmental pollution state in the prior art.
[0005] An embodiment of the present application provides a regional environmental pollution state evaluation method based on Internet of Things, comprising:
[0006] Obtaining soil data, water quality data, vegetation data and real-time environmental covariates of a region to be detected;
[0007] Inputting the soil data, the water quality data and the vegetation data into a preset feature importance analysis model to obtain the feature importance of each data output by the feature importance analysis model;
[0008] According to the soil data, the water quality data, the vegetation data, the real-time environmental covariates and the feature importance of each data, a comprehensive monitoring index parameter is calculated;
[0009] Inputting the comprehensive monitoring index parameter into a pollution state prediction model to obtain a soil and water loss prediction quantity and a non-point source pollution spread prediction range output by the pollution state prediction model; wherein the pollution state prediction model takes a first historical monitoring index parameter, a first soil and water loss quantity label data, a first non-point source pollution range label data and a first environmental covariate of the region to be detected as input data, and takes a predicted soil and water loss simulation quantity and a non-point source pollution spread simulation range as output data, and is obtained after iterative training of a preset neural network;
[0010] According to the water and soil erosion prediction quantity, the non-point source pollution diffusion prediction range, a preset reasonable water and soil erosion loss quantity, and a preset reasonable non-point source pollution diffusion range, a pollution state evaluation index of the to-be-detected region is calculated.
[0011] Further, the construction process of the pollution state prediction model comprises:
[0012] The first historical detection data set, the first water and soil erosion quantity label data, the first non-point source pollution range label data, and the first environmental covariate are obtained, wherein the first historical detection data set comprises first historical soil data, first historical water quality data, and first historical vegetation data of the to-be-detected region at each first preset sampling time;
[0013] The first historical detection data set, the first water and soil erosion quantity label data, the first non-point source pollution range label data, and the first environmental covariate are input into the feature importance analysis model, so as to obtain the feature importance of each item of data in the first historical detection data set at each first preset sampling time output by the feature importance analysis model;
[0014] According to the feature importance of each item of data in the first historical detection data set at each first preset sampling time and the first environmental covariate, a first historical monitoring index parameter corresponding to each first preset sampling time is calculated;
[0015] The first historical monitoring index parameter corresponding to each first preset sampling time, the first water and soil erosion quantity label data, the first non-point source pollution range label data, and the first environmental covariate are taken as input data of a neural network, and the predicted water and soil erosion simulation quantity and non-point source pollution diffusion simulation range are taken as output data of the neural network, so as to iteratively train the neural network;
[0016] The trained neural network is taken as the pollution state prediction model;
[0017] In each iteration training process, the water and soil erosion simulation quantity and non-point source pollution diffusion simulation range predicted by the current neural network are compared with the corresponding first water and soil erosion quantity label data and first non-point source pollution range label data, and a loss function value is calculated according to the comparison result;
[0018] It is judged whether the loss function value converges,
[0019] If yes, the iteration training is terminated, and the trained neural network is obtained,
[0020] If no, the network parameters of the current neural network are adjusted, and the neural network for the next iteration training is obtained.
[0021] Further, the construction process of the feature importance analysis model comprises:
[0022] The second historical detection data set, the second soil and water loss amount label data, the second non-point source pollution range label data and the second environmental covariate of the to-be-detected region are acquired; the second historical detection data set comprises second historical soil data, second historical water quality data and second historical vegetation data of the to-be-detected region at each second preset sampling time;
[0023] The pre-constructed random forest model is iteratively trained according to the second historical detection data set, the second soil and water loss amount label data, the second non-point source pollution range label data and the second environmental covariate until the prediction error of the random forest model meets the preset accuracy requirement or reaches the maximum training iteration number, so as to obtain the random forest model after iterative training;
[0024] The random forest model after iterative training is taken as the feature importance analysis model.
[0025] Further, the iterative training of the pre-constructed random forest model according to the second historical detection data set, the second soil and water loss amount label data, the second non-point source pollution range label data and the second environmental covariate comprises:
[0026] The second historical detection data set, the second soil and water loss amount label data, the second non-point source pollution range label data and the second environmental covariate are taken as the input data of the random forest model,
[0027] The predicted soil and water loss amount and non-point source pollution diffusion range are taken as the output data of the random forest model,
[0028] The soil and water loss amount and non-point source pollution diffusion range are taken as the target variable in the iterative training process of the random forest model,
[0029] The information gain rate of the input data to the target variable is taken as the selection basis of each decision tree split node in the random forest model, and the pre-constructed random forest model is iteratively trained.
[0030] Further, the calculation of the comprehensive monitoring index parameter according to the soil data, the water quality data, the vegetation data, the real-time environmental covariate and the feature importance of each data comprises:
[0031] The real-time environmental covariate and the feature importance of each data are combined to give the soil data, the water quality data and the vegetation data corresponding weight values, so as to obtain the soil data, the water quality data and the vegetation data after weight assignment.
[0032] The comprehensive monitoring index parameters are calculated based on the soil data, water quality data, and vegetation data after assigning weight values.
[0033] Furthermore, before obtaining soil, water quality, and vegetation data for the area to be tested, the following steps are also included:
[0034] Obtain the original soil data set, original water quality data set, and original vegetation data set of the area to be tested;
[0035] For each data set, outliers are removed from the data set using the box plot method to obtain the data set after removing outliers.
[0036] Using the data that was not removed from the data set, the missing data in the data set after removing abnormal data is filled to obtain the preprocessed data set.
[0037] Based on the preprocessed data set, soil data, water quality data, and vegetation data for the area to be tested are determined.
[0038] Furthermore, after calculating the pollution status assessment index of the area to be monitored, the following steps are also included:
[0039] The pollution status assessment index is sent in real time to the terminal of the staff managing the area to be tested.
[0040] An embodiment of the present invention also provides a regional environmental pollution status assessment device based on the Internet of Things, comprising:
[0041] The data acquisition module is used to acquire soil data, water quality data, vegetation data, and real-time environmental covariates of the area to be detected.
[0042] The feature importance calculation module is used to input the soil data, the water quality data, and the vegetation data into a preset feature importance analysis model to obtain the feature importance of each data item output by the feature importance analysis model.
[0043] The comprehensive monitoring index parameter calculation module is used to calculate the comprehensive monitoring index parameters based on the soil data, the water quality data, the vegetation data, the real-time environmental covariates, and the characteristic importance of each data.
[0044] The prediction module is used to input the comprehensive monitoring index parameters into the pollution state prediction model to obtain the predicted amount of soil erosion and the predicted range of non-point source pollution diffusion output by the pollution state prediction model; wherein, the pollution state prediction model is obtained by iteratively training a preset neural network with the first historical monitoring index parameters, the first soil erosion label data, the first non-point source pollution range label data and the first environmental covariate of the area to be monitored as input data, and the predicted simulated amount of soil erosion and the simulated range of non-point source pollution diffusion as output data;
[0045] The pollution status assessment index calculation module is used to calculate the pollution status assessment index of the area to be monitored based on the predicted amount of soil erosion, the predicted range of non-point source pollution diffusion, the preset reasonable amount of soil erosion, and the preset reasonable range of non-point source pollution diffusion.
[0046] This application also provides a terminal device, including:
[0047] One or more processors;
[0048] A memory, coupled to the processor, for storing one or more programs;
[0049] When the one or more programs are executed by the one or more processors, the one or more processors implement the IoT-based regional environmental pollution status assessment method as described in the above embodiments of the invention.
[0050] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the IoT-based regional environmental pollution status assessment method as described in the above embodiments.
[0051] The following benefits can be obtained by implementing the present invention:
[0052] This invention provides a method, apparatus, terminal device, and storage medium for assessing regional environmental pollution status based on the Internet of Things (IoT). The method introduces a feature importance analysis model to quantify and differentiate the contributions of soil, water, and vegetation data, avoiding the drawback of traditional assessments that include all data equally. This makes the subsequently calculated comprehensive monitoring index parameters more closely reflect actual pollution scenarios. Secondly, the pollution status prediction model is built based on iterative training of a neural network, incorporating historical monitoring index parameters and environmental covariates during training. This allows for the accurate output of predicted soil erosion and non-point source pollution diffusion range by combining historical patterns with dynamic environmental changes. Compared to traditional manual experience-based predictions, it possesses stronger temporal dynamic capture capabilities and generalization. Furthermore, based on the predicted soil erosion, the predicted non-point source pollution diffusion range, a preset reasonable soil erosion amount, and a preset reasonable non-point source pollution diffusion range, a pollution status assessment index for the area to be monitored is calculated. This avoids the subjective bias of traditional manual assessments and directly reflects the degree to which pollution in the monitored area exceeds the reasonable range, making the assessment results more objective. Attached Figure Description
[0053] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0054] Figure 1 This is a flowchart illustrating a method for assessing regional environmental pollution status based on the Internet of Things, provided in a certain embodiment of this application.
[0055] Figure 2 This is a schematic diagram of the structure of a regional environmental pollution status assessment device based on the Internet of Things provided in a certain embodiment of this application;
[0056] Figure 3 This is a schematic diagram of the structure of a terminal device provided in a certain embodiment of this application. Detailed Implementation
[0057] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.
[0059] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.
[0060] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0061] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0062] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), similarly, "multiple sets" refers to two or more (including two sets), and "multiple pieces" refers to two or more (including two pieces).
[0063] In the description of the embodiments of this application, unless otherwise expressly specified and limited, technical terms such as "installation," "connection," "joining," and "fixing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. For those skilled in the art, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.
[0064] See Figure 1To address the issue of high subjectivity in existing manual assessments of regional environmental pollution status, an embodiment of the present invention provides an Internet of Things-based method for assessing regional environmental pollution status, comprising:
[0065] S1. Obtain soil data, water quality data, vegetation data, and real-time environmental covariates of the area to be detected;
[0066] To illustrate, in order to accurately assess the environmental pollution status of the area to be tested and to provide multi-dimensional and highly correlated basic data support for subsequent model analysis, it is necessary to obtain soil data, water quality data, vegetation data and real-time environmental covariates of the area to be tested.
[0067] Specifically, soil data includes soil texture, organic matter content, pH value, erosion depth, and other data related to the soil erosion sensitivity and pollutant carrying capacity of the area to be tested; water quality data includes nitrogen and phosphorus concentrations, chemical oxygen demand, dissolved oxygen content, and other data related to the water pollution level and self-purification capacity of the area to be tested; and vegetation data includes vegetation type, coverage, biomass, and other data related to the soil and water loss inhibition capacity and ecological buffering function of the area to be tested.
[0068] In a preferred embodiment, before acquiring soil data, water quality data, and vegetation data of the area to be tested, the method further includes:
[0069] Obtain the original soil data set, original water quality data set, and original vegetation data set of the area to be tested;
[0070] For each data set, outliers are removed from the data set using the box plot method to obtain the data set after removing outliers.
[0071] Using the data that was not removed from the data set, the missing data in the data set after removing abnormal data is filled to obtain the preprocessed data set.
[0072] Based on the preprocessed data set, determine the soil data, water quality data, and vegetation data for the area to be tested;
[0073] As an illustration, because the model has high requirements for the quality of input data, the soil, water and vegetation data of the area to be detected need to meet certain standards. However, the original soil, water and vegetation data sets of the area to be detected have problems such as uneven data distribution, errors or outliers. Therefore, it is necessary to preprocess the original soil, water and vegetation data sets of the area to be detected.
[0074] Specifically, for each data set, outliers are removed using a box plot method to obtain the data set after outlier removal. For each data set (including soil, water, and vegetation data sets), the lower quartile Q1 and upper quartile Q3 of the data set are first calculated. Then, the outlier threshold range is determined according to the formula [Q1 - 1.5 × (Q3 - Q1), Q3 + 1.5 × (Q3 - Q1)]. Finally, the data set is traversed, and data that are less than the lower limit of the outlier threshold range or greater than the upper limit of the outlier threshold range are identified as outliers and removed. The remaining data form the data set after outlier removal.
[0075] As an illustration, after obtaining the data group after removing abnormal data, there will be some missing data in the corresponding positions of the data group. Therefore, it is necessary to use the data that was not removed from the data group to fill in the missing data in the data group after removing abnormal data, so as to obtain the preprocessed data group.
[0076] Specifically, for the indicator columns (soil, water quality, and vegetation) with missing data in the data set after removing outliers, all non-missing data in these indicator columns that were not removed are collected, and the average value of these non-missing data is calculated. Then, the calculated average value is used to fill in the missing data positions of the indicator column, completing the missing data filling process. After filling, the min-max standardization formula is used to standardize the soil, water quality, and vegetation data respectively. The min-max standardization formula is as follows:
[0077] X'=(XX min ) / (X max -X min );
[0078] In the formula, X min X represents the minimum value of a data set (soil, water quality, or vegetation) after removing outliers; max X represents the maximum value of a data set (soil, water quality, or vegetation) after removing outliers; X represents a single original data point in a data set (soil, water quality, or vegetation) after removing outliers; X' represents the standardized data, which maps the data to the range of 0-1.
[0079] S2. Input the soil data, water quality data, and vegetation data into a preset feature importance analysis model to obtain the feature importance of each data item output by the feature importance analysis model;
[0080] To illustrate, in order to clarify the contribution of each environmental data point to the pollution assessment, it is necessary to input the soil data, the water quality data, and the vegetation data into a preset feature importance analysis model to obtain the feature importance of each data point output by the feature importance analysis model.
[0081] In a preferred embodiment, the process of constructing the feature importance analysis model includes:
[0082] Acquire the second historical monitoring dataset, second soil erosion label data, second non-point source pollution range label data, and second environmental covariate of the area to be monitored; wherein, the second historical monitoring dataset includes the second historical soil data, second historical water quality data, and second historical vegetation data of the area to be monitored at each second preset sampling time;
[0083] Based on the second historical monitoring dataset, the second soil erosion label data, the second non-point source pollution range label data, and the second environmental covariate, the pre-constructed random forest model is iteratively trained until the prediction error of the random forest model meets the preset accuracy requirements or reaches the maximum number of training iterations, thus obtaining the iteratively trained random forest model.
[0084] The iteratively trained random forest model was used as the feature importance analysis model.
[0085] Specifically, the second environmental covariate includes climate factors (such as rainfall and temperature), topographic factors (such as slope and slope length), land use data, hydrological data, etc., to reflect external environmental factors that may affect soil erosion and non-point source pollution; the second soil erosion label data is used as a true reference for the model to predict soil erosion, the second non-point source pollution range label data is used as a true reference for the model to predict the range of non-point source pollution, and the data types of the second historical soil data, the second historical water quality data, and the second historical vegetation data are consistent with the data types of the soil data, water quality data, and vegetation data of the area to be detected in step S1, and are used to provide the model with training input data that matches the feature types of the actual scene to be detected;
[0086] Indicatively, the pre-constructed random forest model is iteratively trained based on the second historical monitoring dataset, the second soil erosion label data, the second non-point source pollution range label data, and the second environmental covariate.
[0087] In a preferred embodiment, the step of iteratively training the pre-constructed random forest model based on the second historical monitoring dataset, the second soil erosion label data, the second non-point source pollution range label data, and the second environmental covariate includes:
[0088] The second historical monitoring dataset, the second soil erosion label data, the second non-point source pollution range label data, and the second environmental covariate are used as input data for the random forest model.
[0089] The predicted soil erosion and non-point source pollution spread range are used as the output data of the random forest model.
[0090] Soil erosion and the extent of non-point source pollution diffusion are used as target variables in the iterative training process of the random forest model.
[0091] The information gain rate of the input data to the target variable is used as the basis for selecting each decision tree split node in the random forest model, and the pre-constructed random forest model is iteratively trained.
[0092] Specifically, for each training iteration, input features are extracted from the second historical detection dataset and the second environmental covariate. Combined with the second soil erosion label data and the second non-point source pollution range label data, the combined value of soil erosion and non-point source pollution diffusion range is used as the target variable. Then, for each decision tree in the random forest model, during the construction process, the information gain ratio GR(A,T) of each input feature (such as soil texture, nitrogen and phosphorus concentration, vegetation cover, rainfall, slope, etc.) to the target variable is calculated. The information gain ratio is obtained by the following formula:
[0093]
[0094] In the formula, G(A,T) is the information gain; SIf(A,T) is the split information; GR(A,T) is the information gain ratio of the feature to the target variable;
[0095] Information gain G(A,T) measures the reduction in uncertainty of the target variable after feature splitting, while split information SIf(A,T) reflects the dispersion of feature values. The feature with the largest information gain ratio is selected as the splitting node of the current decision tree, and the data is further split until the decision tree is fully grown or the preset stopping condition is met. At the same time, a 10-fold cross-validation method is used to divide the input training data into 10 parts, and 9 parts are used to train the model and 1 part to validate the model in turn. This optimizes the model parameters of the random forest model, allowing the random forest model to continuously improve the prediction accuracy of soil erosion and non-point source pollution diffusion range during iterative training until the preset accuracy requirement is met or the maximum number of training iterations is reached. The pre-built random forest model is iteratively trained until the prediction error of the random forest model meets the preset accuracy requirement or the maximum number of training iterations is reached, resulting in the iteratively trained random forest model.
[0096] S3. Based on the soil data, water quality data, vegetation data, real-time environmental covariates, and the characteristic importance of each data point, calculate the comprehensive monitoring index parameters;
[0097] In a schematic way, the comprehensive monitoring index parameters are calculated by combining the feature importance of each data output by the feature importance analysis model described in step S2;
[0098] In a preferred embodiment, the step of calculating the comprehensive monitoring index parameters based on the soil data, water quality data, vegetation data, real-time environmental covariates, and the characteristic importance of each data point includes:
[0099] By combining the real-time environmental covariates and the characteristic importance of each data, corresponding weight values are assigned to the soil data, the water quality data, and the vegetation data, resulting in soil data, water quality data, and vegetation data with assigned weight values.
[0100] The comprehensive monitoring index parameters are calculated based on the soil data, water quality data, and vegetation data after assigning weight values.
[0101] As an illustration, after obtaining the feature importance of each data point, it is necessary to adaptively adjust the feature importance of each data point in conjunction with the aforementioned real-time environmental covariates;
[0102] Specifically, real-time environmental covariates include real-time meteorological conditions (such as hourly precipitation, daily average temperature, and light intensity), real-time seasonal and agricultural cycle stages (whether it is the rainy season, agricultural fertilization period, peak vegetation growth season, etc.), and real-time regional characteristics (whether it is a predominantly arable land area, a steep mountainous area, or an ecological protection area, etc.).
[0103] Based on different real-time environmental covariate triggering conditions, the importance of features is adaptively adjusted, such as:
[0104] (1) When hourly precipitation is >50 mm, increase the weight of soil data to 50% and decrease the weight of water quality data to 30%; (2) When the average daily temperature is >30℃ and lasts for more than 3 days, increase the weight of vegetation data to 35% and the weight of water quality data to 45%; (3) When the light intensity is <30000 lux and lasts for 5 days, decrease the weight of vegetation data to 15% and increase the weight of soil data to 40%; (4) During the rainy season, increase the weight of soil data to 40% and increase the weight of water quality data to 45%; (5) During the agricultural fertilization period, increase the weight of water quality data such as total nitrogen and total phosphorus to 50% and decrease the weight of soil data to 30%. 30%; (6) During the peak growing season, the weight of vegetation data is increased to 35%, and the weight of soil data is reduced to 30%; for areas dominated by arable land, the weight of water quality data is increased by 5%-10%; (7) For steep mountainous areas (slope > 25°), the weight of soil data (such as shear strength related) is increased by 10%-15%; (8) For ecological protection areas, the weight of vegetation data is increased by 10% by default; (9) When the system issues an early warning, if subsequent monitoring finds that the actual risk deviates from the early warning by more than 20%, the model will automatically backtrack and adjust the weight of related data. For example, if the early warning underestimates the risk of non-point source pollution, the weight of total nitrogen and total phosphorus will be automatically increased by 5%-8% in the next assessment;
[0105] It should be noted that the triggering conditions for real-time environmental covariates can be adaptively adjusted according to the actual situation, and no specific triggering conditions for real-time environmental covariates are limited here;
[0106] Specifically, after adaptively adjusting the importance of each data feature, the comprehensive monitoring index parameters are calculated based on the adjusted weight values of each data point. The specific calculation formula is as follows:
[0107] M(t) = w s ·S(t)+w w ·W(t)+w v ·V(t);
[0108] In the formula, M(t) represents the comprehensive monitoring index parameter; S(t) is the standardized value of soil data at time t, ranging from 0 to 1; W(t) is the standardized value of water quality data at time t; V(t) is the standardized value of vegetation data at time t; w s Weights for soil data; w w Weights for water quality data; w v The weights for vegetation data.
[0109] S4. Input the comprehensive monitoring index parameters into the pollution state prediction model to obtain the predicted amount of soil erosion and the predicted range of non-point source pollution diffusion output by the pollution state prediction model; wherein, the pollution state prediction model is obtained by iteratively training a preset neural network with the first historical monitoring index parameters, the first soil erosion label data, the first non-point source pollution range label data and the first environmental covariate of the area to be monitored as input data, and the predicted simulated amount of soil erosion and the simulated range of non-point source pollution diffusion as output data;
[0110] Indicatively, after obtaining the comprehensive monitoring index parameters, these parameters need to be input into the pollution state prediction model to obtain the predicted amount of soil erosion and the predicted range of non-point source pollution diffusion output by the pollution state prediction model. It should be noted that the predicted amount of soil erosion and the predicted range of non-point source pollution diffusion output by the pollution state prediction model are the dynamic change trends and risk distributions within the next day. By predicting the development trend of soil erosion and non-point source pollution in the short term, the model solves the problems of lagging prediction of environmental pollution risks and lack of timely early warning basis in the existing technology.
[0111] In a preferred embodiment, the process of constructing the pollution state prediction model includes:
[0112] Acquire the first historical monitoring dataset, the first soil erosion amount label data, the first non-point source pollution range label data, and the first environmental covariate; wherein, the first historical monitoring dataset includes the first historical soil data, the first historical water quality data, and the first historical vegetation data of the area to be monitored at each first preset sampling time;
[0113] The first historical monitoring dataset, the first soil erosion label data, the first non-point source pollution range label data, and the first environmental covariate are input into the feature importance analysis model to obtain the feature importance of each data item in the first historical monitoring dataset at each first preset sampling time as output by the feature importance analysis model.
[0114] Based on the feature importance of each data item in the first historical detection dataset at each first preset sampling time and the first environmental covariate, the parameters of the first historical monitoring index corresponding to each first preset sampling time are calculated.
[0115] The neural network is iteratively trained using the first historical monitoring index parameters corresponding to each first preset sampling time, the first soil erosion label data, the first non-point source pollution range label data, and the first environmental covariate as input data, and the predicted soil erosion simulation amount and non-point source pollution diffusion simulation range as output data.
[0116] The trained neural network is used as the pollution state prediction model;
[0117] In each iteration of training, the simulated amount of soil erosion and the simulated range of non-point source pollution spread predicted by the current neural network are compared with the corresponding first soil erosion amount label data and first non-point source pollution range label data, and the loss function value is calculated based on the comparison results.
[0118] Determine whether the loss function value has converged.
[0119] If so, then terminate the iterative training and obtain the trained neural network.
[0120] If not, the network parameters of the current neural network are adjusted to obtain the neural network for the next iteration of training;
[0121] Schematic, the first historical monitoring dataset includes first historical soil data, first historical water quality data, and first historical vegetation data for the area to be monitored at each first preset sampling time. First soil erosion label data and first non-point source pollution range label data are used, and a first environmental covariate is used. The data types in the first historical monitoring dataset and the first environmental covariate are consistent with the data types of the second historical monitoring dataset and the second environmental covariate.
[0122] Specifically, the first historical monitoring dataset, the first soil erosion label data, the first non-point source pollution range label data, and the first environmental covariate are input into the feature importance analysis model to obtain the feature importance of each data item in the first historical monitoring dataset at each first preset sampling time. Based on the feature importance of each data item in the first historical monitoring dataset at each first preset sampling time and the first environmental covariate, the first historical monitoring index parameters corresponding to each first preset sampling time are calculated. Similarly, in the process of calculating the first historical monitoring index parameters, the weight values need to be adaptively adjusted in combination with the real-time environmental covariate triggering conditions.
[0123] During the model training phase, the parameters of the first historical monitoring indicators corresponding to each first preset sampling time, along with the first soil erosion label data, the first non-point source pollution range label data, and the first environmental covariate, are used as inputs to the neural network. Simultaneously, the simulated soil erosion and non-point source pollution diffusion range output by the neural network are set as training targets. Iterative training is then conducted. Once training is complete, the neural network is used as a pollution state prediction model. In each round of iterative training, the simulated soil erosion and non-point source pollution diffusion range output by the current neural network are compared with the corresponding first soil erosion label data and first non-point source pollution range label data, respectively. The loss function value is then calculated based on the deviation generated by the comparison. It is then determined whether the loss function value has reached convergence. If convergence has occurred, iterative training stops, and the resulting neural network is the trained model. If convergence has not occurred, the network parameters of the current neural network are adjusted, and the adjusted network is used for the next round of iterative training until the convergence condition is met.
[0124] Specifically, the comprehensive monitoring index parameters are input into the pollution state prediction model, and the expression of the pollution state prediction model is as follows:
[0125]
[0126] In the formula, The predicted value at time (t+k); σ is the sigmoid activation function, which maps the output to the range of 0-1, W x W represents the weight parameters input to the model at the current time, and M(t) represents the comprehensive monitoring index parameters at time t. h is the weight parameter of the hidden layer; h(t) is the hidden layer state at time t; b is the bias term, which is obtained through model training and optimization.
[0127] S5. Based on the predicted amount of soil erosion, the predicted range of non-point source pollution diffusion, the preset reasonable amount of soil erosion, and the preset reasonable range of non-point source pollution diffusion, calculate the pollution status assessment index of the area to be monitored.
[0128] Specifically, the predicted amount of soil erosion is compared with the reasonable amount of soil erosion. If the predicted amount of soil erosion is not greater than the reasonable amount of soil erosion, it indicates that the predicted amount is within the normal range, and the first excess value D1 of this indicator is recorded as 0. If the predicted amount of soil erosion is greater than the reasonable amount of soil erosion, the specific formula for calculating the first excess value D1 is as follows:
[0129]
[0130] Similarly, comparing the predicted range of non-point source pollution diffusion with the reasonable range of non-point source pollution diffusion, if the predicted range of non-point source pollution diffusion is not greater than the reasonable amount of soil and water loss, it indicates that the pollution diffusion is at a normal level, and the second excess value D2 of this indicator is recorded as 0. If the predicted range of non-point source pollution diffusion is greater than the reasonable amount of soil and water loss, the specific formula for calculating the second excess value D2 is as follows:
[0131]
[0132] The first excess value D1 and the second excess value D2 are weighted and summed to obtain the comprehensive excess value D. 超出 ;
[0133] The system retrieves a pre-defined historical fault record database and uses the statistical value F of the frequency of abnormal soil erosion and non-point source pollution (i.e., indicators exceeding the normal range) in the current monitored area from the historical fault record database. 历史 Combined with the comprehensive excess value D 超出 The final pollution status assessment index P of the area to be tested is determined using the following formula:
[0134] P = α·D 超出 +(1-α)F 历史 ;
[0135] In the formula, α represents the weight of the overall influence exceeding the degree value, and its value is between 0 and 1, which can be adjusted according to the actual situation.
[0136] In a preferred embodiment, after calculating the pollution status assessment index of the area to be detected, the method further includes:
[0137] The pollution status assessment index is sent in real time to the terminal of the staff managing the area to be tested;
[0138] Specifically, the pollution status assessment index is sent in real time to the terminals of staff managing the area to be tested, so that staff can promptly grasp any abnormalities in the pollution status of the area. This provides a basis for decision-making in subsequent prevention and control work using the intelligent control module. It also allows staff to know immediately whether to initiate the optimization process after receiving feedback on the effectiveness of the plan from the effect evaluation module, thereby managing and controlling regional environmental pollution more efficiently.
[0139] See Figure 2 This invention provides an Internet of Things-based regional environmental pollution status assessment device, comprising:
[0140] The data acquisition module is used to acquire soil data, water quality data, vegetation data, and real-time environmental covariates of the area to be detected.
[0141] The feature importance calculation module is used to input the soil data, the water quality data, and the vegetation data into a preset feature importance analysis model to obtain the feature importance of each data item output by the feature importance analysis model.
[0142] The comprehensive monitoring index parameter calculation module is used to calculate the comprehensive monitoring index parameters based on the soil data, the water quality data, the vegetation data, the real-time environmental covariates, and the characteristic importance of each data.
[0143] The prediction module is used to input the comprehensive monitoring index parameters into the pollution state prediction model to obtain the predicted amount of soil erosion and the predicted range of non-point source pollution diffusion output by the pollution state prediction model; wherein, the pollution state prediction model is obtained by iteratively training a preset neural network with the first historical monitoring index parameters, the first soil erosion label data, the first non-point source pollution range label data and the first environmental covariate of the area to be monitored as input data, and the predicted simulated amount of soil erosion and the simulated range of non-point source pollution diffusion as output data;
[0144] The pollution status assessment index calculation module is used to calculate the pollution status assessment index of the area to be monitored based on the predicted amount of soil erosion, the predicted range of non-point source pollution diffusion, the preset reasonable amount of soil erosion, and the preset reasonable range of non-point source pollution diffusion.
[0145] See Figure 3 One embodiment of this application also provides a terminal device, including:
[0146] One or more processors;
[0147] A memory, coupled to the processor, for storing one or more programs;
[0148] When the one or more programs are executed by the one or more processors, the one or more processors implement the IoT-based regional environmental pollution status assessment method as described above.
[0149] The processor controls the overall operation of the terminal device to complete all or part of the steps of the IoT-based regional environmental pollution status assessment method described above. The memory stores various types of data to support the operation of the terminal device. This data may include, for example, instructions for any application or method operating on the terminal device, as well as application-related data. The memory can be implemented using any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0150] In an exemplary embodiment, the terminal device may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the IoT-based regional environmental pollution status assessment method as described in any of the above embodiments, and achieve the same technical effects as the above methods.
[0151] In another exemplary embodiment, a computer-readable storage medium including a computer program is also provided. When executed by a processor, the computer program implements the steps of the IoT-based regional environmental pollution status assessment method as described in any of the foregoing embodiments. For example, the computer-readable storage medium may be the aforementioned memory including the computer program, which may be executed by a processor of a terminal device to complete the IoT-based regional environmental pollution status assessment method as described in any of the foregoing embodiments and achieve the same technical effects as the aforementioned method.
[0152] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for assessing regional environmental pollution status based on the Internet of Things, characterized in that, include: Acquire soil data, water quality data, vegetation data, and real-time environmental covariates of the area to be tested; The soil data, water quality data, and vegetation data are input into a preset feature importance analysis model to obtain the feature importance of each data item output by the feature importance analysis model. Based on the soil data, water quality data, vegetation data, real-time environmental covariates, and the characteristic importance of each data point, comprehensive monitoring index parameters are calculated. The comprehensive monitoring index parameters are input into the pollution state prediction model to obtain the predicted amount of soil erosion and the predicted range of non-point source pollution diffusion output by the pollution state prediction model; wherein, the pollution state prediction model is obtained by iteratively training a preset neural network with the first historical monitoring index parameters, the first soil erosion label data, the first non-point source pollution range label data and the first environmental covariate of the area to be monitored as input data, and the predicted simulated amount of soil erosion and the simulated range of non-point source pollution diffusion as output data; Based on the predicted amount of soil erosion, the predicted range of non-point source pollution diffusion, the preset reasonable amount of soil erosion, and the preset reasonable range of non-point source pollution diffusion, the pollution status assessment index of the area to be monitored is calculated.
2. The method for assessing regional environmental pollution status based on the Internet of Things as described in claim 1, characterized in that, The process of constructing the pollution state prediction model includes: Acquire the first historical monitoring dataset, the first soil erosion amount label data, the first non-point source pollution range label data, and the first environmental covariate; wherein, the first historical monitoring dataset includes the first historical soil data, the first historical water quality data, and the first historical vegetation data of the area to be monitored at each first preset sampling time; The first historical monitoring dataset, the first soil erosion label data, the first non-point source pollution range label data, and the first environmental covariate are input into the feature importance analysis model to obtain the feature importance of each data item in the first historical monitoring dataset at each first preset sampling time as output by the feature importance analysis model. Based on the feature importance of each data item in the first historical detection dataset at each first preset sampling time and the first environmental covariate, the parameters of the first historical monitoring index corresponding to each first preset sampling time are calculated. The neural network is iteratively trained using the first historical monitoring index parameters corresponding to each first preset sampling time, the first soil erosion label data, the first non-point source pollution range label data, and the first environmental covariate as input data, and the predicted soil erosion simulation amount and non-point source pollution diffusion simulation range as output data. The trained neural network is used as the pollution state prediction model; In each iteration of training, the simulated amount of soil erosion and the simulated range of non-point source pollution spread predicted by the current neural network are compared with the corresponding first soil erosion amount label data and first non-point source pollution range label data, and the loss function value is calculated based on the comparison results. Determine whether the loss function value has converged. If so, then terminate the iterative training and obtain the trained neural network. If not, the network parameters of the current neural network are adjusted to obtain the neural network for the next iteration of training.
3. The method for assessing regional environmental pollution status based on the Internet of Things as described in claim 1, characterized in that, The construction process of the feature importance analysis model includes: Acquire the second historical monitoring dataset, second soil erosion label data, second non-point source pollution range label data, and second environmental covariate of the area to be monitored; wherein, the second historical monitoring dataset includes the second historical soil data, second historical water quality data, and second historical vegetation data of the area to be monitored at each second preset sampling time; Based on the second historical monitoring dataset, the second soil erosion label data, the second non-point source pollution range label data, and the second environmental covariate, the pre-constructed random forest model is iteratively trained until the prediction error of the random forest model meets the preset accuracy requirements or reaches the maximum number of training iterations, thus obtaining the iteratively trained random forest model. The random forest model, which has been trained iteratively, is used as the feature importance analysis model.
4. The method for assessing regional environmental pollution status based on the Internet of Things as described in claim 3, characterized in that, The step of iteratively training a pre-constructed random forest model based on the second historical monitoring dataset, the second soil erosion label data, the second non-point source pollution range label data, and the second environmental covariate includes: The second historical monitoring dataset, the second soil erosion label data, the second non-point source pollution range label data, and the second environmental covariate are used as input data for the random forest model. The predicted soil erosion and non-point source pollution spread range are used as the output data of the random forest model. Soil erosion and the extent of non-point source pollution diffusion are used as target variables in the iterative training process of the random forest model. The information gain rate of the input data on the target variable is used as the basis for selecting each decision tree split node in the random forest model, and the pre-constructed random forest model is iteratively trained.
5. The method for assessing regional environmental pollution status based on the Internet of Things as described in claim 1, characterized in that, The comprehensive monitoring index parameters are calculated based on the soil data, water quality data, vegetation data, real-time environmental covariates, and the characteristic importance of each data point, including: By combining the real-time environmental covariates and the characteristic importance of each data, corresponding weight values are assigned to the soil data, the water quality data, and the vegetation data, resulting in soil data, water quality data, and vegetation data with assigned weight values. The comprehensive monitoring index parameters are calculated based on the soil data, water quality data, and vegetation data after assigning weight values.
6. The method for assessing regional environmental pollution status based on the Internet of Things as described in claim 1, characterized in that, Before acquiring soil, water, and vegetation data for the area to be tested, the following steps are also included: Obtain the original soil data set, original water quality data set, and original vegetation data set of the area to be tested; For each data set, outliers are removed from the data set using the box plot method to obtain the data set after removing outliers. Using the data that was not removed from the data set, the missing data in the data set after removing abnormal data is filled to obtain the preprocessed data set. Based on the preprocessed data set, soil data, water quality data, and vegetation data for the area to be tested are determined.
7. The method for assessing regional environmental pollution status based on the Internet of Things as described in claim 1, characterized in that, After calculating the pollution status assessment index of the area to be tested, the following steps are also included: The pollution status assessment index is sent in real time to the terminal of the staff managing the area to be tested.
8. A regional environmental pollution status assessment device based on the Internet of Things (IoT), applicable to the regional environmental pollution status assessment method based on the IoT as described in claims 1-7, characterized in that, include: The data acquisition module is used to acquire soil data, water quality data, vegetation data, and real-time environmental covariates of the area to be detected. The feature importance calculation module is used to input the soil data, the water quality data, and the vegetation data into a preset feature importance analysis model to obtain the feature importance of each data item output by the feature importance analysis model. The comprehensive monitoring index parameter calculation module is used to calculate the comprehensive monitoring index parameters based on the soil data, the water quality data, the vegetation data, the real-time environmental covariates, and the characteristic importance of each data. The prediction module is used to input the comprehensive monitoring index parameters into the pollution state prediction model to obtain the predicted amount of soil erosion and the predicted range of non-point source pollution diffusion output by the pollution state prediction model; wherein, the pollution state prediction model is obtained by iteratively training a preset neural network with the first historical monitoring index parameters, the first soil erosion label data, the first non-point source pollution range label data and the first environmental covariate of the area to be monitored as input data, and the predicted simulated amount of soil erosion and the simulated range of non-point source pollution diffusion as output data; The pollution status assessment index calculation module is used to calculate the pollution status assessment index of the area to be monitored based on the predicted amount of soil erosion, the predicted range of non-point source pollution diffusion, the preset reasonable amount of soil erosion, and the preset reasonable range of non-point source pollution diffusion.
9. A terminal device, characterized in that, include: One or more processors; A memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the IoT-based regional environmental pollution status assessment method as described in any one of claims 1-7.
10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the Internet of Things-based regional environmental pollution status assessment method as described in any one of claims 1-7.