Enterprise land soil pollution risk identification method and system based on big data
Through a big data-based method, the Internet of Things and hyperspectral imaging technology is used, combined with multi-gated hybrid expert network and SHAP value analysis, the problems of synergistic effects and dynamic changes identification in enterprise land are solved, efficient pollution identification and risk warning are achieved, and identification accuracy and interpretability are improved.
Patent Information
- Application Number
- CN202510157164.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology is difficult to effectively identify the synergistic effects between multiple pollutants in enterprise land, lacks real-time response to dynamic changes in pollution, and the model of intelligent identification methods is poorly interpretable, making it difficult to provide a clear theoretical basis for pollution prevention and control decisions.
The soil pollution risk identification method for enterprise land use is adopted based on big data, and multi-source data is collected through IoT sensors and hyperspectral imaging equipment to build a multi-level characteristic index system. A multi-gated hybrid expert network is used to build a multi-task learning model, and combined with SHAP value analysis and decision tree agent model to achieve collaborative identification of multiple pollutants and dynamic risk warning.
It has realized the coordinated identification of multiple pollutants in enterprise land and dynamic risk warning, improved the identification accuracy and model interpretability, and provided a scientific basis for pollution prevention and control decisions.
Smart Images

Figure CN120031382A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular to a method and system for identifying soil pollution risks in enterprise land based on big data. Background Art
[0002] With the acceleration of industrialization, the problem of soil pollution in enterprise land has become increasingly prominent. Existing soil pollution identification methods mainly include two categories: one is a single pollutant detection method based on traditional physical and chemical indicators, which determines the pollutant content through sampling and analysis; the other is a pollution assessment method based on statistical models, which uses mathematical statistics to analyze the distribution law of pollution. These methods have achieved certain results in practical applications, especially in the identification of fixed pollution sources and the monitoring of single pollutants. At the same time, with the development of Internet of Things technology and artificial intelligence algorithms, some studies have begun to try to apply sensor networks and machine learning methods to soil pollution monitoring, providing more technical means and data support.
[0003] However, existing technical methods have obvious shortcomings: first, traditional single pollutant detection methods are difficult to cope with the complex synergistic effects of multiple pollutants and cannot effectively identify the mutual influence between different pollutants; second, methods based on statistical models mainly rely on historical data and lack the ability to respond to dynamic changes in pollution in real time; third, although existing intelligent identification methods have introduced machine learning technology, the model's interpretability is poor and it is difficult to provide a clear theoretical basis for pollution prevention and control decisions. Especially in scenarios such as corporate land, where pollution sources are complex and pollution types are diverse, existing methods are difficult to achieve coordinated identification and dynamic early warning of multiple pollutants. Summary of the invention
[0004] The present application provides a method and system for identifying soil pollution risks in enterprise land based on big data. The present application can simultaneously process multiple pollutants such as heavy metals, VOCs and SVOCs, and has a soil pollution intelligent identification method with model interpretability, thereby realizing the coordinated identification of multiple pollutants in enterprise land and dynamic risk warning.
[0005] In the first aspect, the present application provides a method for identifying soil pollution risks on enterprise land based on big data, and the method for identifying soil pollution risks on enterprise land based on big data comprises: collecting basic site environment data through Internet of Things sensors, and collecting spectral feature data and enterprise production process data through hyperspectral imaging, preprocessing the basic site environment data, spectral feature data and enterprise production process data to obtain a preprocessed data set; constructing a multi-level feature index including a basic feature layer, a pollution source feature layer and a pollution pathway feature layer according to the preprocessed data set, generating interactive features through feature transformation and feature combination processing, and screening features through correlation analysis and variance analysis to obtain an optimized feature index system; based on the optimized feature index system, constructing a multi-task learning model architecture including a shared underlying network, a multi-expert network and a task-specific network using a multi-gated hybrid expert network, and setting A gating mechanism and an attention mechanism are used to extract the characteristics of heavy metal pollution, VOCs pollution and SVOCs pollution to obtain a multi-pollutant identification model; based on the multi-pollutant identification model, a pre-training method is used to initialize the shared underlying network, and the model is trained through a curriculum learning strategy and an adversarial training method, and the model parameters are optimized through a dynamically weighted multi-task loss function to obtain an optimized pollution identification model; a SHAP value analysis method is used to perform global and local interpretative analysis on the optimized pollution identification model, and the key driving factors are determined through feature importance calculation. An interpretable rule set is constructed through a decision tree proxy model to obtain a model interpretation result; based on the model interpretation result, an early warning system including a risk assessment module and a prevention and control recommendation module is constructed, and data processing and analysis functions are realized through a microservice architecture, and graded early warning information is generated through a preset threshold system, and soil pollution risk identification results are output.
[0006] In a second aspect, the present application provides a system for identifying the risk of soil pollution in enterprise land based on big data, and the system for identifying the risk of soil pollution in enterprise land based on big data includes: An acquisition module is used to collect basic site environment data through IoT sensors, and to collect spectral feature data and enterprise production process data through hyperspectral imaging, and to preprocess the basic site environment data, spectral feature data and enterprise production process data to obtain a preprocessed data set; A transformation module is used to construct a multi-level feature index including a basic feature layer, a pollution source feature layer and a pollution pathway feature layer according to the preprocessed data set, generate interactive features through feature transformation and feature combination processing, perform feature screening through correlation analysis and variance analysis, and obtain an optimized feature index system; An extraction module is used to construct a multi-task learning model architecture including a shared underlying network, a multi-expert network and a task-specific network using a multi-gated hybrid expert network based on the optimized feature index system, extract the characteristics of heavy metal pollution, VOCs pollution and SVOCs pollution by setting a gating mechanism and an attention mechanism, and obtain a multi-pollutant recognition model; A training module, for initializing a shared underlying network using a pre-training method based on the multi-pollutant recognition model, training the model using a curriculum learning strategy and an adversarial training method, optimizing model parameters via a dynamically weighted multi-task loss function, and obtaining an optimized pollution recognition model; An analysis module is used to perform global and local interpretative analysis on the optimized pollution identification model using a SHAP value analysis method, determine key driving factors by feature importance calculation, construct an interpretable rule set via a decision tree proxy model, and obtain a model interpretation result; The early warning module is used to construct an early warning system including a risk assessment module and a prevention and control suggestion module based on the results of the model interpretation, realize data processing and analysis functions through the microservice architecture, generate graded early warning information through a preset threshold system, and output soil pollution risk identification results.
[0007] In the technical solution provided by the present application, through the multi-source data collection mechanism of IoT sensors and hyperspectral imaging equipment, the comprehensive acquisition of basic site environmental data, spectral feature data and enterprise production process data is realized, providing a rich data foundation for pollution identification; a multi-level feature index system construction method is adopted to organically combine the basic feature layer, pollution source feature layer and pollution pathway feature layer, and the feature expression ability is improved through feature transformation and combination processing; the model architecture design based on multi-gated hybrid expert network makes full use of the feature extraction ability of the shared underlying network and the task specialization advantages of the expert network to realize the coordinated identification of heavy metal pollution, VOCs pollution and SVOCs pollution; through the combination of pre-training method and curriculum learning strategy, the adaptability of the model to complex pollution scenes is enhanced, and the introduction of adversarial training method improves the robustness of the model; the application of SHAP value analysis method makes the model have good interpretability, which can clearly show the contribution of various features to the pollution identification results; the early warning system design based on microservice architecture realizes the automation of the whole process from data processing to risk warning, and the establishment of a hierarchical early warning mechanism provides a scientific basis for pollution prevention and control. The application of this method in the field of soil pollution risk identification for enterprise land particularly demonstrates the advantages of artificial intelligence algorithms in solving complex environmental problems. The design of the multi-gated hybrid expert network not only improves the accuracy of identifying multiple types of pollutants, but also enhances the model's perception of different pollution characteristics through the introduction of attention mechanisms and gating mechanisms, while maintaining a high computational efficiency. In addition, by integrating domain knowledge and data-driven methods, this method ensures the interpretability and reliability of the identification results while ensuring model performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0009] Figure 1 This is a schematic diagram of an embodiment of a method for identifying soil pollution risks of enterprise land based on big data in an embodiment of the present application; Figure 2 A schematic diagram of the process of initializing a shared underlying network in an embodiment of the present application; Figure 3 This is a schematic diagram of the risk classification standard in the embodiment of this application; Figure 4 This is a schematic diagram of an embodiment of a soil pollution risk identification system for enterprise land based on big data in an embodiment of the present application. DETAILED DESCRIPTION
[0010] The embodiment of the present application provides a method and system for identifying the risk of soil pollution in enterprise land based on big data. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0011] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In the embodiment of the present application, an embodiment of the method for identifying the risk of soil pollution in enterprise land based on big data includes: Step S101, collecting basic site environment data through Internet of Things sensors, and collecting spectral feature data and enterprise production process data through hyperspectral imaging, preprocessing the basic site environment data, spectral feature data and enterprise production process data to obtain a preprocessed data set; Step S102: construct a multi-level feature index including a basic feature layer, a pollution source feature layer, and a pollution pathway feature layer according to the preprocessed data set, generate interactive features through feature transformation and feature combination processing, perform feature screening through correlation analysis and variance analysis, and obtain an optimized feature index system; Step S103: Based on the optimized feature index system, a multi-task learning model architecture including a shared underlying network, a multi-expert network and a task-specific network is constructed using a multi-gated hybrid expert network. The features of heavy metal pollution, VOCs pollution and SVOCs pollution are extracted by setting a gating mechanism and an attention mechanism to obtain a multi-pollutant recognition model. Step S104: Based on the multi-pollutant identification model, a pre-training method is used to initialize the shared underlying network, and the model is trained through a curriculum learning strategy and an adversarial training method. The model parameters are optimized through a dynamically weighted multi-task loss function to obtain an optimized pollution identification model; Step S105: using the SHAP value analysis method to perform global and local interpretability analysis on the optimized pollution identification model, determining key driving factors through feature importance calculation, constructing an interpretable rule set through a decision tree proxy model, and obtaining a model interpretation result; Step S106: Based on the model interpretation results, an early warning system including a risk assessment module and a prevention and control recommendation module is constructed, data processing and analysis functions are realized through a microservice architecture, graded early warning information is generated through a preset threshold system, and soil pollution risk identification results are output.
[0012] It is understandable that the execution subject of this application can be a big data-based enterprise land soil pollution risk identification system, or a terminal or a server, which is not limited here. The embodiment of this application is described by taking the server as the execution subject as an example.
[0013] Specifically, the basic data of the site environment is collected by deploying an IoT sensor network. The IoT sensors include soil pH sensors, moisture sensors, and temperature sensors. Each sensor records the corresponding indicator value at a sampling interval of 5 minutes. At the same time, drones equipped with hyperspectral imaging equipment conduct regular aerial surveys to obtain spectral feature data of the surface soil of the site. The spectral data contains continuous reflectivity information from visible light to near-infrared bands. The production process data of the enterprise includes raw material usage records, waste disposal records, and process parameter records. The collected multi-source data is preprocessed, and the sliding window median filter is used to process the outliers in the sensor data. The window size is set to 60 minutes. The hyperspectral data is corrected for atmospheric scattering to eliminate atmospheric effects, and geometric correction is performed to ensure accurate spatial position. Enterprise information data is structured according to unified coding specifications. All types of data are aligned to the hourly scale in the time dimension and unified to the site grid unit in the spatial dimension.
[0014] A multi-level feature index system is constructed based on the preprocessed data. The basic feature layer includes basic physical and chemical indicators such as soil pH, organic matter content, and site utilization intensity. The pollution source feature layer includes pollution information such as the raw material composition, process type, and pollutant emission of the enterprise. The pollution pathway feature layer includes environmental factors such as topography, hydrogeology, and land use. The numerical features are logarithmically transformed and standardized, and the categorical features are uniquely encoded to generate standardized feature vectors. Interactive features are constructed through feature combination, such as the product of pH value and organic matter content reflects the soil buffering capacity. A multi-task learning architecture is constructed using a multi-gated hybrid expert network, and the underlying network is shared to extract common feature representations. The expert network focuses on the characteristic patterns of three types of pollutants: heavy metals, VOCs, and SVOCs. The weight contribution of different expert networks is dynamically adjusted through the gating mechanism, and the attention mechanism enhances the perception of key features. The model training adopts a three-stage strategy: pre-training to initialize parameters, curriculum learning to gradually increase task difficulty, and adversarial training to improve robustness. The dynamically weighted multi-task loss function balances the optimization objectives of different pollutant identification tasks.
[0015] The SHAP value analysis method is used to interpret the model prediction results, calculate the contribution of features to different pollutant identification tasks, and identify key driving factors. For example, pH value has a significant impact on the migration and transformation of heavy metals, and organic matter content is closely related to the adsorption and desorption of organic pollutants. The decision tree proxy model is used to transform complex deep learning rules into an interpretable set of discriminant conditions.
[0016] Based on the model interpretation results, an early warning system is built. The risk assessment module determines the risk level according to the pollutant concentration level and characteristic importance, and the prevention and control suggestion module generates prevention and control measures for sites with different risk levels. The system adopts a microservice architecture to ensure the loose coupling and scalability of each functional module, realizes data flow through message queues, and triggers the early warning mechanism when it is detected that the risk level exceeds the threshold.
[0017] For example, in a chemical enterprise's land, the installed pH sensors detected that the pH value in the local area dropped from 7.2 to 5.8, and the hyperspectral data showed that the near-infrared reflectivity of the area changed significantly. Combined with the company's production records, it was found that the area was close to acidic waste liquid storage facilities. Feature analysis showed that pH value, waste liquid concentration, and soil organic matter content were the main influencing factors. The system issued early warning information in a timely manner, and it was recommended to strengthen the anti-leakage inspection and acid-base balance adjustment measures of waste liquid storage facilities. This application makes full use of the complementary advantages of multi-source data, reveals the causes of pollution through explainable machine learning methods, and provides technical support for the intelligent identification and risk control of soil pollution in enterprise land.
[0018] During the processing, there is a strict logical relationship between the data. Sensor data reflects the current status of the site, hyperspectral data provides spatial distribution information, and enterprise record data reveals potential pollution sources. Feature engineering transforms these raw data into indicators that reflect pollution characteristics, the multi-task learning framework realizes the coordinated identification of different types of pollutants, and model explanatory analysis transforms data-driven results into understandable pollution prevention and control recommendations.
[0019] In the embodiment of the present application, through the multi-source data collection mechanism of IoT sensors and hyperspectral imaging equipment, the comprehensive acquisition of site environment basic data, spectral feature data and enterprise production process data is realized, providing a rich data foundation for pollution identification; a multi-level feature index system construction method is adopted to organically combine the basic feature layer, pollution source feature layer and pollution pathway feature layer, and the feature expression ability is improved through feature transformation and combination processing; the model architecture design based on multi-gated hybrid expert network makes full use of the feature extraction ability of the shared underlying network and the task specialization advantages of the expert network to realize the coordinated identification of heavy metal pollution, VOCs pollution and SVOCs pollution; through the combination of pre-training method and curriculum learning strategy, the adaptability of the model to complex pollution scenes is enhanced, and the introduction of adversarial training method improves the robustness of the model; the application of SHAP value analysis method makes the model have good interpretability, which can clearly show the contribution of various features to the pollution identification results; the early warning system design based on microservice architecture realizes the automation of the whole process from data processing to risk warning, and the establishment of a hierarchical early warning mechanism provides a scientific basis for pollution prevention and control. The application of this method in the field of soil pollution risk identification for enterprise land particularly demonstrates the advantages of artificial intelligence algorithms in solving complex environmental problems. The design of the multi-gated hybrid expert network not only improves the accuracy of identifying multiple types of pollutants, but also enhances the model's perception of different pollution characteristics through the introduction of attention mechanisms and gating mechanisms, while maintaining a high computational efficiency. In addition, by integrating domain knowledge and data-driven methods, this method ensures the interpretability and reliability of the identification results while ensuring model performance.
[0020] In a specific embodiment, the process of executing step S101 may specifically include the following steps: (1) The basic site environmental data of soil pH, water content, and temperature are collected by IoT sensors, the spectral characteristic data of the site surface soil reflectance is obtained by hyperspectral imaging equipment, and the enterprise production process data of production materials, process flow, and waste disposal are obtained from the enterprise database; (2) Use a sliding median filter with a time window size of T to remove outliers from the basic data of the site environment, perform atmospheric scattering correction and geometric position correction on the spectral feature data, and perform structured coding on the enterprise production process data according to standard coding specifications; (3) Resample the corrected site environment basic data according to the time series, spatially align the corrected spectral feature data with the site environment basic data, and temporally align the encoded enterprise production process data with the site environment basic data; (4) For the missing values in the site environment basic data after resampling, the time series interpolation algorithm is used to fill them. For the missing values in the spectral feature data after registration, the spatial interpolation algorithm is used to fill them. For the missing values in the enterprise production process data after time alignment, the nearest neighbor filling method is used to fill them. (5) Construct a data quality assessment system that includes data integrity indicators, data consistency indicators, and data accuracy indicators, and perform quality scoring on the filled site environment basic data, spectral characteristic data, and enterprise production process data; (6) The basic site environment data, spectral feature data, and enterprise production process data whose quality scores exceed the preset threshold are merged and integrated to generate a preprocessed data set.
[0021] Specifically, the IoT sensor network contains multiple sensor nodes, each of which is equipped with a soil pH sensor, a moisture sensor, and a temperature sensor, with a sampling frequency of once every 5 minutes. UAVs equipped with hyperspectral imaging equipment conduct regular aerial surveys to obtain continuous spectral reflectance data in the 400-2500nm band with a spectral resolution of 10nm. At the same time, data such as raw material usage lists, production process parameters, and waste disposal records are extracted from the enterprise production management system. The sliding median filter algorithm is used to remove outliers for the processing of site environmental basic data, and the time window size T is set to 60 minutes. The sliding window moves sequentially in the time series, and the median is calculated for the data in each window to replace the outliers. For spectral feature data, atmospheric scattering correction is performed to eliminate the influence of factors such as water vapor and aerosols in the atmosphere, and then geometric correction is performed based on ground control points to ensure accurate spatial position. The structured processing of enterprise production process data includes unified coding of raw material types, process parameters, and waste types, and establishing a standard data format.
[0022] Data temporal and spatial alignment is a key step to ensure the comparability of multi-source data. The basic data of the site environment are resampled at intervals of 1 hour, and the representative value is obtained by calculating the average value of all sampling points within each hour. The spatial registration of spectral feature data is based on the geographic coordinate system, and the spectral images acquired at different times are aligned to a unified spatial reference system. The enterprise production process data are time-stamped according to the production batch and aligned with the timestamp of the environmental monitoring data. Different filling strategies are used for the problem of missing data. The basic data of the site environment adopts a time series interpolation algorithm, based on the temporal continuity characteristics of the data, and uses the valid values of the previous and next moments for interpolation calculations. The missing pixels in the spectral feature data use a spatial interpolation algorithm, and the spectral information of the surrounding pixels is used for spatial interpolation. The nearest neighbor filling method is used for the enterprise production process data to fill the missing values with the closest valid records.
[0023] The data quality assessment system is evaluated from three dimensions: completeness, consistency, and accuracy. The completeness index calculates the proportion of valid records of data, the consistency index evaluates the degree of internal correlation of data, and the accuracy index measures the degree of deviation of data from reference values. The quality score is calculated for each type of data, and the score threshold is set to filter high-quality data. The various types of data that have passed the quality assessment are merged and integrated to form a comprehensive data set containing spatiotemporal information, environmental parameters, and production records. The data are associated with each other through timestamps and spatial coordinates to facilitate subsequent feature extraction and analysis modeling.
[0024] For example, 100 sensor nodes were deployed in the factory area, and monitoring data was collected for 30 consecutive days. The soil pH data showed abnormal jumps due to sensor failure. These abnormal values were effectively identified and corrected through sliding median filtering. Hyperspectral imaging acquired 5 images of the surface soil of the site. After correction and registration, it was found that some areas had missing data due to weather reasons. The spectral information was supplemented by a spatial interpolation algorithm. The production records provided by the company showed the fluctuations in process parameters during the production of a batch of products. After time alignment, these data established a clear correspondence with the environmental monitoring data. In the final preprocessed data set, the time resolution of the site environment basic data is 1 hour and the spatial resolution is 10 meters. It contains multi-dimensional information such as soil physical and chemical parameters, spectral characteristics, and production processes.
[0025] In a specific embodiment, the process of executing step S102 may specifically include the following steps: (1) The soil pH value, organic matter content, site utilization intensity, utilization years, and meteorological and hydrological parameters in the pre-processed data set are divided into the basic characteristic layer; the enterprise raw material composition, process type, pollutant emission, treatment process, and surrounding pollution source distribution data are divided into the pollution source characteristic layer; the topography, hydrogeology, land use pattern, and population distribution data are divided into the pollution pathway characteristic layer; (2) Perform logarithmic transformation and standardization on the numerical features in the basic feature layer, perform one-hot encoding on the categorical features, and generate a basic feature vector; perform normalization on the numerical features in the pollution source feature layer, perform label encoding on the categorical features, and generate a pollution source feature vector; perform Z-score standardization on the numerical features in the pollution pathway feature layer, perform serial number encoding on the categorical features, and generate a pollution pathway feature vector; (3) Combine the basic feature vector with the pollution source feature vector according to the time dimension, and combine the combined result with the pollution pathway feature vector according to the spatial dimension to generate interactive features; (4) Calculate the Pearson correlation coefficient between each pair of features in the interaction features, and eliminate redundant features with a correlation coefficient greater than the threshold; perform a one-way ANOVA on the remaining features, calculate the F statistic of each feature pair with respect to the pollutant concentration, and eliminate ineffective features with an F statistic less than the threshold; (5) Combine each feature in the interaction features to construct feature interaction terms including first-order feature combinations and second-order feature combinations, and calculate the mutual information value between the feature interaction terms and the pollutant concentration; (6) Integrate the feature combinations with a mutual information value higher than the threshold in the feature interaction terms and the remaining features to form an optimized feature index system.
[0026] Specifically, the preprocessed dataset is divided into three levels according to feature attributes. The basic feature layer contains indicators reflecting the background characteristics of the site: the soil pH value reflects the acid-base environment, the organic matter content characterizes the soil adsorption capacity, the site utilization intensity describes the degree of land development, the utilization years record the usage duration, and the meteorological and hydrological parameters include environmental factors such as rainfall and groundwater level. The pollution source feature layer contains indicators related to enterprise production: the raw material composition records the types and amounts of raw materials, the process type describes the production process, the pollutant emissions statistics record the emissions of wastewater and waste gas, the treatment process records the pollution control methods, and the surrounding pollution source distribution data record the locations of other pollution sources. The pollution pathway feature layer contains indicators related to migration and diffusion: the topography and geomorphology describe the surface morphology, the hydrogeology records the groundwater flow direction, the land use pattern reflects the exposure pathway, and the population distribution data characterizes the receptor situation. Three methods are used for data standardization processing: perform logarithmic transformation on the numerical features in the basic feature layer to eliminate the dimension difference, and then standardize them to the interval [0, 1]; use the maximum-minimum normalization for the numerical features in the pollution source feature layer; use the Z-score normalization for the numerical features in the pollution pathway feature layer to make the data follow the standard normal distribution. Three methods are also used for categorical feature encoding: use one-hot encoding for the basic feature layer to convert the category into a binary vector; use label encoding for the pollution source feature layer to assign an integer label to each category; use ordinal encoding for the pollution pathway feature layer to assign values in a specific order.
[0027] The feature combination process is divided into two steps: in the time dimension, pair the basic feature vector and the pollution source feature vector according to the timestamp to form a time-series combined feature; then in the space dimension, correspond the time-series combined feature and the pollution pathway feature vector according to the spatial coordinates to form the complete interaction features.
[0028] For feature correlation analysis, the following method is used: Pearson correlation coefficient calculation formula: , where, represents the correlation coefficient between features x and y, and They represent the eigenvalues of the i-th sample, and They represent the means of features x and y respectively, and n is the number of samples.
[0029] F statistic calculation formula: , in, represents the F statistic, MSB represents the mean square between groups, and MSW represents the mean square within groups. represents the mean of the jth group, represents the overall mean, represents the value of the i-th sample in the j-th group, k represents the number of groups, represents the number of samples in the jth group, and N represents the total number of samples.
[0030] Mutual information value calculation formula: , in, represents the mutual information value of features X and Y, represents the joint probability distribution, p(x) and p(y) represent the marginal probability distribution, and They represent the feature value space respectively.
[0031] Taking the enterprise land of a chemical park as an example, the monitoring data was processed through the above-mentioned feature engineering method: the soil pH value showed significant temporal volatility, and the distribution became more even after logarithmic transformation; the raw material composition data contained multiple chemical categories, which were converted into numerical features through label coding; the original values of the topographic data varied greatly, and were easier to compare after Z-score standardization. During the feature screening process, the soil pH value showed a strong correlation with the heavy metal content, while some meteorological parameters were eliminated due to the small F statistic. Feature interaction analysis found that the interaction term between pH value and organic matter content had a strong correlation with pollutant mobility. The optimized feature indicator system contains feature combinations that are both independent of each other and have significant explanatory power.
[0032] In a specific embodiment, the process of executing step S103 may specifically include the following steps: (1) Input the optimized feature index system into the shared underlying network consisting of a three-layer feedforward neural network, perform nonlinear transformation on the input features, and generate shared feature representations; (2) The shared feature representation is input into the heavy metal expert network, VOCs expert network, and SVOCs expert network respectively to separate and extract the features of different pollutants and form an expert feature vector; (3) Calculate the expert network weight coefficient of the expert feature vector through the soft gating function, and perform weighted fusion on the expert feature vector using the weight coefficient to obtain the fused feature vector; (4) Input the fused feature vector into the heavy metal task network, VOCs task network, and SVOCs task network, calculate the feature importance distribution of each pollutant type through the attention weight, and generate the task feature vector; (5) Normalize the task feature vector to generate the identification results of heavy metal pollution characteristics, VOCs pollution characteristics, and SVOCs pollution characteristics, and form a multi-pollutant identification model; (6) The output results of the multi-pollutant recognition model are compared with the true labels through the cross-entropy loss function, and the loss value of each task network is calculated.
[0033] Specifically, the optimized feature index system is input into the shared underlying network for processing. The shared underlying network adopts a three-layer feedforward neural network structure, and each layer is composed of multiple neurons. The feedforward neural network is a network structure that transmits information unidirectionally from the input layer to the output layer. The middle hidden layer transforms the input features through a nonlinear activation function to enhance the expressive ability of the features. The first layer receives the original feature input, and the second and third layers gradually extract higher-level abstract features to finally generate a shared feature representation. The obtained shared feature representation is input into three expert networks for processing. The heavy metal expert network specializes in processing feature patterns related to heavy metal elements, such as pH value, organic matter and other factors that affect the migration and transformation of heavy metals; the VOCs expert network focuses on the feature patterns of volatile organic compounds, such as temperature, pressure and other parameters that affect the volatilization and diffusion of VOCs; the SVOCs expert network focuses on the feature patterns of semi-volatile organic compounds, such as adsorption coefficient, solubility and other indicators that affect the behavior of SVOCs in soil. After training, each expert network can extract feature vectors related to its respective tasks.
[0034] The feature vectors extracted by different expert networks need to be fused through a soft gating mechanism. The soft gating function uses the softmax function to calculate the weight coefficient of each expert network. The softmax function maps any real number to the interval (0,1) and the sum is 1, which ensures the rationality of the weight coefficient. According to the properties of the current input features, the weight contributions of different expert networks are dynamically adjusted to achieve adaptive fusion of features. The fused feature vectors are input into three task-specific networks: heavy metal task network, VOCs task network, and SVOCs task network. Each task network contains an attention mechanism layer, which is used to calculate the importance weights of different features for the identification of this type of pollutants. The attention mechanism learns the correlation strength between features and tasks through training. Important features obtain higher attention weights, and secondary features obtain lower weights.
[0035] The task feature vector is normalized to the maximum and minimum values, and the values are mapped to the standard range. The normalized feature vector passes through the fully connected layer and the softmax layer, and the recognition results of heavy metal pollution, VOCs pollution and SVOCs pollution are output respectively, including the category probability distribution of each pollutant. The cross entropy loss function is used to evaluate the difference between the model output and the true label. The cross entropy loss function measures the degree of difference between the two probability distributions and provides an optimization target for model training.
[0036] Taking a site of an electroplating enterprise as an example, in the soil monitoring data processing, data containing features such as pH value, heavy metal content, and VOCs concentration are input into a shared underlying network. After the data undergoes three layers of nonlinear transformation, high-level features reflecting the spatial distribution and temporal changes of pollutants are extracted. These features are simultaneously input into three expert networks, and each network extracts characteristic patterns related to specific pollutants according to its own parameter settings. When it is detected that the pH value of a certain area is abnormally reduced and the heavy metal content is increased, the weight coefficient of the heavy metal expert network will automatically increase, highlighting the characteristic expression of heavy metal pollution. The final model output shows that there is a risk of excessive heavy metals such as nickel and chromium in the area, while the risk levels of VOCs and SVOCs are low, which is consistent with the actual situation. The entire recognition process reflects the advantages of the multi-task learning framework in the assessment of complex contaminated sites, and improves the accuracy of recognition through feature sharing and task decomposition.
[0037] Soft gating function formula: , in, represents the weight coefficient of the i-th expert network, represents the weight matrix, h represents the input feature, represents the bias term, and M represents the number of expert networks.
[0038] Attention mechanism calculation formula: , in, represents the attention weight of the k-th feature, represents the feature vector, Q represents the query matrix, f represents the input feature, c represents the context vector, and N represents the number of features.
[0039] Cross entropy loss function formula: , Among them, L represents the loss value, represents the true label of the jth task in the i-th category, represents the predicted value, C represents the number of pollutant categories, and T represents the number of tasks.
[0040] In a specific embodiment, the process of executing step S104 may specifically include the following steps: (1) Input the historical soil pollution monitoring data into the shared underlying network, pre-train the network parameters through unsupervised learning, and generate initialization parameters; (2) The training samples were stratified according to the pollutant concentration thresholds of heavy metals, VOCs, and SVOCs. A multi-level pollution intensity sample set was formed by concentration interval division. Data was input and parameters were adjusted for the multi-pollutant identification model in order of pollution intensity from low to high to obtain preliminary training results. (3) Add adversarial perturbation samples to the preliminary training results, optimize parameters through adversarial training methods, and generate adversarial training results; (4) Calculate the loss values of the three tasks of heavy metal pollution, VOCs pollution, and SVOCs pollution in the adversarial training results, and weight them based on the task difficulty coefficient to obtain the multi-task loss value; (5) Back-propagate the multi-task loss values to the shared underlying network, multi-expert network, and task-specific network, update the network connection weights, and generate optimization parameters; (6) The optimized parameters are updated to the multi-pollutant identification model to form an optimized pollution identification model.
[0041] Specifically, Figure 2As shown, it is a schematic diagram of the process of initializing the shared underlying network in the embodiment of the present application. The historical monitoring data is input into the autoencoder for pre-training to generate network initialization parameters; the samples are layered according to the pollutant concentration threshold to form a multi-level sample set, and the training is carried out in the order of pollution intensity to obtain preliminary training results; adversarial samples are generated based on the preliminary training results, and adversarial training is carried out together with the original samples to obtain training results; the loss values of the three tasks of heavy metal pollution, VOCs pollution, and SVOCs pollution in the training results are calculated respectively, and the multi-task loss values are obtained by weighting the task difficulty coefficient; the multi-task loss values are updated to the network through back propagation, and finally an optimized pollution identification model is formed. The shared underlying network is pre-trained unsupervised using historical monitoring data. The historical data containing various pollutants such as heavy metals, VOCs, and SVOCs are used as input, and the intrinsic feature representation of the data is learned by the autoencoder. The autoencoder is an unsupervised learning method that compresses the input data and then reconstructs it. The feature representation of the data is learned by minimizing the reconstruction error, and then the initialization parameters of the network are obtained. The layered processing of the training samples adopts a method based on the pollutant concentration threshold. According to the soil environmental quality standards, heavy metal pollution samples are divided into three levels: mild, moderate, and severe according to their content levels; VOCs pollution samples are divided into three intervals: low, medium, and high according to the concentration of volatile organic compounds; SVOCs pollution samples are divided into different pollution levels according to the content of semi-volatile organic compounds. The divided sample sets are input into the model for training in order of pollution intensity from low to high, first learning simple pollution patterns, and then gradually transitioning to complex pollution scenarios.
[0042] Adversarial training enhances the robustness of the model by adding perturbed samples. Adding small perturbations based on gradient information to existing training samples generates adversarial samples. Adversarial samples are similar to original samples in feature space, but may cause significant changes in model prediction results. By training original samples and adversarial samples at the same time, the model's resistance to data perturbations is enhanced. In the multi-task learning framework, it is necessary to balance the training objectives of different pollutant identification tasks. The cross entropy loss is calculated for the three tasks of heavy metal pollution, VOCs pollution, and SVOCs pollution, and weighted according to the difficulty coefficient of the task. The task difficulty coefficient reflects the complexity of identifying different types of pollutants. More difficult tasks are given larger weights to balance the learning process.
[0043] The back propagation of the loss value uses the chain rule to pass the gradient information to each network layer in turn. Starting from the output layer, the gradient is back propagated layer by layer to the task-specific network, the multi-expert network, and the shared underlying network. The network weights of each layer are updated according to the gradient information, and the update amplitude is controlled by the learning rate. The back propagation process adjusts the network parameters in the direction of reducing the loss value. The optimized parameters are applied to the multi-pollutant recognition model to complete the model optimization process. The optimized model has better generalization and anti-interference capabilities while maintaining recognition accuracy.
[0044] For example, soil monitoring data from the past five years was used to construct a training set, which included a variety of pollutant indicators such as heavy metals (copper, zinc, lead, cadmium), VOCs (benzene, toluene, xylene) and SVOCs (polycyclic aromatic hydrocarbons, phthalates). During the model training process, the data was preprocessed and feature extracted by the autoencoder to obtain the initial feature representation reflecting the distribution law of pollutants. Then the samples were divided into different levels according to the degree of pollution, starting with low-concentration samples of a single pollutant, and gradually introducing complex samples of coexistence of multiple pollutants. During the training process, different task weights were set according to the characteristics of different pollutants, such as giving higher weights to VOCs with strong mobility. By adding adversarial samples to simulate the uncertainty in the field sampling and analysis process, the adaptability of the model was improved. The final optimized model can accurately identify the distribution characteristics and interaction relationships of different types of pollutants.
[0045] Pre-training loss function: , in, represents the pre-training loss, represents the input features, represents the reconstruction feature, represents the reconstruction weight, represents the network parameters, represents the regularization coefficient, D represents the feature dimension, and P represents the number of parameters.
[0046] Adversarial perturbation generation formula: , in, represents the adversarial sample, x represents the original sample, represents the disturbance amplitude, represents the task loss function, and y represents the label.
[0047] Multi-task loss function: , in, represents the total loss, represents the loss of the kth task, represents the task weight, represents the regularization term, represents the regularization coefficient.
[0048] In a specific embodiment, the process of executing step S105 may specifically include the following steps: (1) Calculate the SHAP value for each feature data input into the optimized pollution identification model, and obtain the global interpretability evaluation index through feature cumulative contribution analysis; (2) The global explanatory evaluation indicators were grouped according to the pollutant type, and the characteristic SHAP values of heavy metal pollution, VOCs pollution, and SVOCs pollution were ranked to generate a characteristic importance ranking table; (3) Based on the feature importance ranking table, a feature subset with a contribution greater than a threshold is selected, and the local SHAP value is calculated for the contamination identification sample to form a local explanatory feature; (4) Input local explanatory features into the decision tree to divide the feature space, extract feature splitting rules, and construct a feature discrimination condition set; (5) Combine the characteristic discrimination condition sets to generate identification rule chains for heavy metal pollution, VOCs pollution, and SVOCs pollution, forming an interpretable rule set; (6) Combine the interpretable rule set with the feature importance ranking table to generate the model interpretation results.
[0049] Specifically, the SHAP value needs to be calculated for the explanatory analysis of the optimized pollution identification model. The SHAP value is a feature importance measurement method based on game theory, which explains the model decision by calculating the marginal contribution of each feature to the model prediction result. For each input feature, its contribution to the prediction result under different feature combinations is calculated separately, and the global importance evaluation index of the feature is accumulated. The global explanatory evaluation index is grouped into three categories: heavy metal pollution, VOCs pollution, and SVOCs pollution. The relevant features of each type of pollutant are sorted according to the size of the SHAP value to form a feature importance ranking table. For heavy metal pollution, the SHAP values of physical and chemical indicators such as soil pH and organic matter content are relatively high; for VOCs pollution, the SHAP values of environmental parameters such as temperature and pressure are significant; for SVOCs pollution, the SHAP values of features such as soil texture and adsorption coefficient are prominent.
[0050] Based on the feature importance ranking table, a feature subset with a SHAP value greater than the set threshold is selected for local explanatory analysis. The local SHAP value calculation is for a single prediction sample to evaluate the contribution of each feature to the specific prediction result. The local explanatory features reflect the feature combination that the model focuses on when processing a specific sample. The local explanatory features are input into the decision tree model, and the discrimination rules are extracted by recursively dividing the feature space. Each node of the decision tree represents a feature discrimination condition, such as "pH value is less than 5.5" or "benzene concentration is greater than 0.2 mg / kg". By traversing all paths of the decision tree, a complete set of feature discrimination conditions is extracted.
[0051] The characteristic discrimination condition set is combined to form a pollution identification rule chain. The rule chain describes the identification logic of different types of pollutants, such as "when the pH value is lower than 5.5 and the organic matter content is greater than 2%, the risk of heavy metal exceeding the standard is significant", "when the groundwater level is high and the VOCs concentration exceeds the threshold, the risk of pollution diffusion increases", etc. These rules express the decision logic of the model in the form of if-then.
[0052] The interpretable rule set is integrated with the feature importance ranking table to form a complete model interpretation result. The interpretation result includes both the global importance assessment of the feature for pollution identification and the decision-making rules in specific scenarios, providing an understandable basis for pollution risk assessment. Taking the site pollution survey of a chemical enterprise as an example, the SHAP value analysis found that: in the identification of heavy metal pollution, the SHAP values of pH value and organic matter content ranked high, indicating that these two factors have a significant impact on the migration and transformation of heavy metals; in the identification of VOCs pollution, the SHAP values of temperature change and soil moisture content are relatively high, reflecting that these environmental factors play an important role in the diffusion of volatile organic compounds; in the identification of SVOCs pollution, the SHAP values of soil organic matter and clay content are prominent, indicating that these features are closely related to the adsorption and desorption of semi-volatile organic compounds.
[0053] In the analysis of specific samples, it was found that the soil pH in a certain area was 4.8 and the organic matter content was 3.5%. The calculation of the local SHAP value showed that this combination significantly increased the activity of heavy metals such as copper and zinc. The rules generated by the decision tree analysis show that when the pH value is lower than 5.0, the bioavailability of heavy metal ions increases; when the organic matter content exceeds 3%, the complexation of heavy metals is enhanced. These rules help explain why the risk of heavy metal pollution in this area is high. Similarly, in VOCs polluted areas, when the temperature rises above 25°C and the soil moisture content is less than 15%, the volatilization rate of volatile organic compounds increases significantly. This rule explains why VOCs pollution problems in this area are more prominent during summer droughts. In this way, the prediction results of the model are converted into specific and actionable pollution prevention and control recommendations.
[0054] In a specific embodiment, the process of executing step S106 may specifically include the following steps: (1) The feature importance ranking table and interpretable rule set in the model interpretation results are used as input data and transmitted to the risk assessment module through the data interface for data aggregation to generate a risk assessment matrix; (2) The risk assessment matrix is graded according to the risk level thresholds of heavy metal pollution, VOCs pollution, and SVOCs pollution, a three-level pollution risk warning standard is established, and a risk grading standard table is formed; (3) Input the risk classification standard table into the prevention and control recommendation module, generate a corresponding list of prevention and control measures according to the characteristic discrimination condition set of different pollutant types, and construct a prevention and control response strategy; (4) Encapsulate the risk assessment matrix and the prevention and control response strategy, transmit the data through the message queue in the microservice architecture, and generate a risk warning data stream; (5) Dynamically monitor the risk level information in the risk warning data stream. When the risk level exceeds the threshold, the graded warning mechanism is triggered to form graded warning information; (6) Associate the graded warning information with the corresponding prevention and control response strategies to form soil pollution risk identification results.
[0055] Specifically, the construction of the risk assessment matrix starts with the feature importance ranking table and the interpretable rule set. The feature importance ranking table records the contribution of each indicator to pollution identification, such as the importance weights of physical and chemical indicators such as pH value and organic matter content; the interpretable rule set contains descriptive rules for the behavior of pollutants, such as "heavy metal activity increases when pH value is lower than 5.5" and other discrimination conditions. These two types of data are standardized and packaged in JSON format, transmitted to the risk assessment module through the data interface, and matrixed according to the type of pollutants and influencing factors. The classification of the risk assessment matrix is based on the national soil environmental quality standards and the technical guidelines for site risk assessment. For heavy metal pollution, the risk level is divided into mild (1-2 times the standard value), moderate (2-3 times the standard value), and severe (more than 3 times the standard value) according to the content of pollutants; for VOCs pollution, a three-level warning threshold is set based on the concentration of volatile organic compounds; for SVOCs pollution, the risk level is determined based on the content of semi-volatile organic compounds. The classification results form a standardized risk grading table, which contains the risk level determination criteria and corresponding warning thresholds for various pollutants.
[0056] The risk grading standard table is combined with the characteristic discrimination condition set to generate targeted prevention and control measures recommendations. For heavy metal pollution, prevention and control measures in high-risk areas include technical means such as pH adjustment, barrier and anti-seepage; for VOCs pollution, focus on gas phase diffusion prevention and control and volatilization inhibition measures; for SVOCs pollution, emphasis is placed on adsorption management and migration control programs. The prevention and control measures for each type of pollutant are associated with specific trigger conditions to form a complete prevention and control response strategy. The risk assessment matrix and prevention and control response strategy use a microservice architecture for data transmission. Asynchronous data processing is achieved through message queues, and each piece of data contains fields such as timestamp, spatial coordinates, pollutant concentration, and risk level. The message queue adopts a publish-subscribe model to ensure real-time transmission and processing of data streams. The data is encrypted and compressed during transmission to ensure transmission efficiency and data security. Such as Figure 3 As shown, it is a schematic diagram of the risk grading standard in the embodiment of this application, showing the risk grading standards of three types of pollutants (heavy metals, VOCs, SVOCs), and each type of pollutant is divided into three levels: light pollution, moderate pollution, and severe pollution. Among them, heavy metal pollution is graded based on the multiples of the standard value, light pollution corresponds to 1-2 times the standard value, moderate pollution corresponds to 2-3 times the standard value, and severe pollution corresponds to more than 3 times the standard value; VOCs and SVOCs pollution are graded based on screening values and control values, light pollution is between the screening value and the control value, moderate pollution is 1-3 times the control value, and severe pollution exceeds 3 times the control value.
[0057] Dynamic monitoring of risk levels is based on real-time data stream analysis. The monitoring program continuously compares the latest monitoring data with the warning threshold. When an excess condition is detected, the corresponding level of warning mechanism is immediately triggered. The warning information contains key information such as the point of excess, type of pollutant, multiple of excess, and time of occurrence. The determination of the warning level comprehensively considers multiple dimensions such as pollutant concentration, scope of impact, and duration. The graded warning information is associated and integrated with the corresponding prevention and control response strategy. Each warning information is linked to the corresponding list of prevention and control measures to form a complete risk identification result. The results include both quantitative indicators of risk assessment and specific prevention and control recommendations to provide support for site management decisions.
[0058] Taking a fine chemical enterprise site as an example, the risk identification process collected feature importance data and found that soil pH and organic matter content had significant effects on heavy metal migration behavior, and temperature and pressure played a dominant role in VOCs diffusion. The risk assessment matrix showed that the site was contaminated by excessive nickel and benzene. The soil pH in a certain area was continuously below 4.5, resulting in a significant increase in heavy metal activity; in another area, the horizontal migration of VOCs was aggravated due to the increase in groundwater levels. Based on these findings, prevention and control measures include specific plans such as soil pH adjustment and groundwater level control. Real-time monitoring data showed that the concentration of benzene fluctuated and increased during the rainfall, triggering a moderate risk warning, and corresponding prevention and control measures such as leakage inspection and wastewater collection were initiated.
[0059] In a specific embodiment, the process of performing the step of grading the risk assessment matrix according to the risk level thresholds of heavy metal pollution, VOCs pollution, and SVOCs pollution may specifically include the following steps: (1) The heavy metal pollution values in the risk assessment matrix are classified into three risk levels: light pollution, moderate pollution, and heavy pollution according to the pollutant content range, and the heavy metal pollution risk classification data is generated; (2) The VOCs pollution values in the risk assessment matrix are classified into three risk levels: light pollution, moderate pollution, and heavy pollution according to the concentration range of volatile organic compounds, and VOCs pollution risk classification data is generated; (3) The SVOCs pollution values in the risk assessment matrix are classified into three risk levels: light pollution, moderate pollution, and severe pollution according to the concentration range of semi-volatile organic compounds, and the SVOCs pollution risk classification data is generated; (4) Combine the heavy metal pollution risk classification data, VOCs pollution risk classification data, and SVOCs pollution risk classification data according to the pollutant type, mark the warning color of each risk level, and generate a risk level warning matrix; (5) Comprehensively score the risk level of each pollutant type in the risk level warning matrix, calculate the comprehensive risk index according to the risk weight, and generate a comprehensive risk assessment table; (6) Combine the risk level warning matrix with the comprehensive risk assessment table to generate a risk grading standard table.
[0060] Specifically, according to the soil environmental quality standards, the benchmark values for heavy metal pollution classification are determined by element type, such as copper, lead, cadmium, nickel, etc. For each heavy metal element, the ratio of the detected concentration to the standard value is used as the basis for classification: when the ratio is within the range of 1-2 times, it is classified as light pollution, 2-3 times is classified as moderate pollution, and more than 3 times is classified as heavy pollution. In the case of the coexistence of multiple heavy metals, the heavy metal pollution risk level of the point is determined according to the highest level of the single item.
[0061] The hierarchical treatment of VOCs pollution is based on the characteristics of volatile organic compounds. According to the risk screening values of VOCs in soil, typical VOCs components such as benzene, toluene, and xylene are classified. Mild pollution corresponds to the concentration range that exceeds the screening value but is lower than the regulatory value. Moderate pollution corresponds to the concentration between the regulatory value and three times the regulatory value. Severe pollution corresponds to the concentration exceeding three times the regulatory value. Considering the volatile characteristics of VOCs, the classification process also needs to be corrected by considering influencing factors such as environmental temperature and soil moisture content. The SVOCs pollution classification adopts a similar three-level classification method. For semi-volatile organic compounds such as polycyclic aromatic hydrocarbons and polychlorinated biphenyls, corresponding concentration classification thresholds are set based on their residual characteristics and bioavailability in soil. The influence of soil organic matter content on the adsorption of SVOCs is considered in the classification process, and the concentration threshold is appropriately increased under the condition of high organic matter content.
[0062] The classification data of the three types of pollutants are integrated, and different risk levels are marked with color identifiers: green indicates mild pollution, yellow indicates moderate pollution, and red indicates severe pollution. The color identifiers intuitively reflect the spatial distribution characteristics of the pollution degree. The risk level data of different types of pollutants are corresponded according to the spatial coordinates to form a multi-dimensional risk level warning matrix. The comprehensive risk assessment adopts a weighted scoring method. Weight coefficients are set for each type of pollutant, and the weight values are determined based on the toxicity, mobility, and cumulative properties of the pollutants. For heavy metal pollution, its persistence and cumulative properties are considered. The weight of VOCs pollution is related to its volatile diffusion characteristics. The weight of SVOCs pollution reflects its enrichment effect in the food chain. For each site, the classification scores of various pollutants are multiplied by the weight coefficients and summed to obtain the comprehensive risk index of the site.
[0063] The risk level warning matrix and the comprehensive risk assessment table are combined to form a complete risk classification standard table. The standard table not only includes the risk level determination results of individual pollutants but also reflects the overall pollution risk status of the site.
[0064] For example, this site involves multiple production processes such as electroplating and organic synthesis. Through systematic sampling and analysis, it is found that: chromium pollution is mainly distributed around the electroplating workshop, and the chromium content at some points reaches 2.5 times the standard value, which is classified as moderate pollution; benzene series pollution is concentrated in the organic synthesis area, and the highest benzene concentration at a point exceeds the regulatory value by 2.8 times, which also belongs to moderate pollution; polycyclic aromatic hydrocarbon pollution is detected in the raw material storage area, but the concentration is low, belonging to mild pollution. Through weighted superposition analysis, it shows that the comprehensive risk index of the area around the electroplating workshop is the highest, followed by the organic synthesis area.
[0065] In a specific embodiment, the process of performing the step of dynamically monitoring the risk level information in the risk warning data stream may specifically include the following steps: (1) Segment the risk warning data stream according to time windows, calculate the change rate of heavy metal pollution, VOCs pollution, and SVOCs pollution risk index within each time window, and generate risk change trend data; (2) Compare the risk change trend data with the thresholds in the risk grading standard table, mark the time points when the thresholds of light pollution, moderate pollution, and heavy pollution are exceeded, and form a time series table of risk breakthrough points; (3) Count the continuous breakthrough time periods in the risk breakthrough point time series table, record the duration of the risk level according to the pollutant type, and generate a risk continuity assessment table; (4) Group the data in the risk persistence assessment table according to pollutant type and risk level, and count the frequency of risk events of each level to form a risk frequency statistical table; (5) Set warning trigger conditions for the data in the risk frequency statistics table, mark high-frequency risk events as key warning targets, and generate a warning target list; (6) Combine the list of warning objects with the corresponding risk change trend data to form graded warning information.
[0066] Specifically, the risk warning data stream is divided into time windows. 24 hours is selected as the basic time window unit, and the risk index change rate is calculated for the heavy metal, VOCs and SVOCs pollution data in each window. The risk index change rate is obtained by dividing the difference between the pollutant concentration value in the current time window and the previous time window by the time interval, reflecting the dynamic change characteristics of the pollution situation. The calculated risk change trend data is compared and analyzed with the risk grading standard table. The standard table contains thresholds for three pollution levels: mild, moderate and severe, which correspond to the concentration limits of different pollutants. Through data comparison, the time point when the pollutant concentration exceeds the threshold of each level for the first time is marked, and the specific values and environmental parameters of the breakthrough moment are recorded to generate a time series table of risk breakthrough points.
[0067] Conduct continuous analysis of the data in the risk breakthrough point time series table. Identify the time periods of continuous exceeding of the standard, and record the start and end time and duration of each exceeding standard event. Count them separately according to the type of pollutant. Focus on the cumulative effect of heavy metal pollution, focus on the daily change pattern of VOCs pollution, and track the long-term change trend of SVOCs pollution. This information is summarized to form a risk continuity assessment table. The data in the risk continuity assessment table is classified and summarized according to the pollution type and risk level. Count the number of occurrences of each pollutant at different risk levels, and record the longest duration and cumulative duration of a single event. In this way, the occurrence patterns and time characteristics of different types of pollution events are mastered to form a risk frequency statistical table.
[0068] Set graded warning conditions for the risk frequency statistics table. Determine the warning priority based on the frequency, duration and risk level of pollution events. When risk events of a certain type of pollutant occur frequently or last for a long time, mark it as a key warning object. The warning object list records the types of pollutants that need to be focused on, the affected areas and key influencing factors.
[0069] Combined with the list of early warning objects and risk change trend data, a complete hierarchical early warning information is constructed. The early warning information includes the dynamic change characteristics of pollutants, the time pattern of breaking through the threshold, the degree of continuous exceeding the standard, and the judgment results of the early warning level.
[0070] For example, monitoring data showed that chromium pollution fluctuated significantly during the rainy season. Through a 24-hour time window analysis, it was found that the concentration of chromium rose rapidly within 4 hours after the start of rainfall, exceeding the threshold of mild pollution, and then maintained at a moderate pollution level during the continuous rainfall. Time series analysis shows that this phenomenon will be repeated after each heavy rainfall, and the duration is usually 48-72 hours. Through frequency statistics, it was found that such events occurred an average of 2-3 times per month during the rainy season, which is a high-incidence risk event. At the same time, VOCs pollution also showed regular changes during the high temperatures in summer, and short-term exceeding the standard often occurred during the period of highest daytime temperature. These findings are integrated into the graded early warning information, clarifying the risk characteristics and early warning priorities of different types of pollutants, and providing targeted decision-making basis for site pollution prevention and control.
[0071] The above describes the method for identifying the risk of soil pollution in enterprise land based on big data in the embodiment of the present application. The following describes the system for identifying the risk of soil pollution in enterprise land based on big data in the embodiment of the present application. Figure 4 In the embodiment of the present application, an embodiment of the enterprise land soil pollution risk identification system based on big data includes: The acquisition module 201 is used to collect basic site environment data through an Internet of Things sensor, and collect spectral feature data and enterprise production process data through hyperspectral imaging, and pre-process the basic site environment data, spectral feature data and enterprise production process data to obtain a pre-processed data set; A transformation module 202 is used to construct a multi-level feature index including a basic feature layer, a pollution source feature layer and a pollution pathway feature layer according to the preprocessed data set, generate interactive features through feature transformation and feature combination processing, perform feature screening through correlation analysis and variance analysis, and obtain an optimized feature index system; Extraction module 203, used to construct a multi-task learning model architecture including a shared underlying network, a multi-expert network and a task-specific network based on the optimized feature index system using a multi-gated hybrid expert network, extract the features of heavy metal pollution, VOCs pollution and SVOCs pollution by setting a gating mechanism and an attention mechanism, and obtain a multi-pollutant recognition model; A training module 204 is used to initialize the shared underlying network based on the multi-pollutant recognition model using a pre-training method, perform model training using a curriculum learning strategy and an adversarial training method, and optimize model parameters via a dynamically weighted multi-task loss function to obtain an optimized pollution recognition model; The analysis module 205 is used to perform global and local interpretative analysis on the optimized pollution identification model using the SHAP value analysis method, determine the key driving factors by feature importance calculation, construct an interpretable rule set through a decision tree proxy model, and obtain a model interpretation result; The early warning module 206 is used to construct an early warning system including a risk assessment module and a prevention and control suggestion module based on the model interpretation results, realize data processing and analysis functions through the microservice architecture, generate graded early warning information through a preset threshold system, and output soil pollution risk identification results.
[0072] Through the collaborative cooperation of the above components, the multi-source data collection mechanism of IoT sensors and hyperspectral imaging equipment has achieved comprehensive acquisition of site environmental basic data, spectral feature data and enterprise production process data, providing a rich data foundation for pollution identification; the construction method of a multi-level feature indicator system is adopted to organically combine the basic feature layer, pollution source feature layer and pollution pathway feature layer, and the feature expression ability is improved through feature transformation and combination processing; the model architecture design based on multi-gated hybrid expert network fully utilizes the feature extraction ability of the shared underlying network and the task specialization advantages of the expert network to achieve the coordinated identification of heavy metal pollution, VOCs pollution and SVOCs pollution; through the combination of pre-training methods and curriculum learning strategies, the model's adaptability to complex pollution scenarios is enhanced, and the introduction of adversarial training methods improves the robustness of the model; the application of SHAP value analysis method makes the model have good interpretability and can clearly show the contribution of various features to pollution identification results; the early warning system design based on microservice architecture realizes the automation of the entire process from data processing to risk warning, and the establishment of a hierarchical early warning mechanism provides a scientific basis for pollution prevention and control. The application of this method in the field of soil pollution risk identification for enterprise land particularly demonstrates the advantages of artificial intelligence algorithms in solving complex environmental problems. The design of the multi-gated hybrid expert network not only improves the accuracy of identifying multiple types of pollutants, but also enhances the model's perception of different pollution characteristics through the introduction of attention mechanisms and gating mechanisms, while maintaining a high computational efficiency. In addition, by integrating domain knowledge and data-driven methods, this method ensures the interpretability and reliability of the identification results while ensuring model performance.
[0073] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features therein can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for identifying soil pollution risks in enterprise land based on big data, characterized in that: The method for identifying soil pollution risks in enterprise land based on big data includes: The basic data of the site environment is collected through the Internet of Things sensor, and the spectral feature data and the enterprise production process data are collected through hyperspectral imaging, and the basic data of the site environment, the spectral feature data and the enterprise production process data are preprocessed to obtain a preprocessed data set; According to the preprocessed data set, a multi-level feature index including a basic feature layer, a pollution source feature layer and a pollution pathway feature layer is constructed, interactive features are generated through feature transformation and feature combination processing, and feature screening is performed through correlation analysis and variance analysis to obtain an optimized feature index system; According to the optimized feature index system, a multi-task learning model architecture including a shared underlying network, a multi-expert network and a task-specific network is constructed using a multi-gated hybrid expert network. The features of heavy metal pollution, VOCs pollution and SVOCs pollution are extracted by setting a gating mechanism and an attention mechanism to obtain a multi-pollutant recognition model. Based on the multi-pollutant identification model, a pre-training method is used to initialize the shared underlying network, the model is trained through a curriculum learning strategy and an adversarial training method, and the model parameters are optimized through a dynamically weighted multi-task loss function to obtain an optimized pollution identification model; The optimized pollution identification model is analyzed globally and locally by using the SHAP value analysis method, key driving factors are determined by feature importance calculation, and an interpretable rule set is constructed through a decision tree proxy model to obtain a model interpretation result; According to the interpretation results of the model, an early warning system including a risk assessment module and a prevention and control recommendation module is constructed. The data processing and analysis functions are realized through the microservice architecture, and graded early warning information is generated through a preset threshold system to output soil pollution risk identification results.
2. The method for identifying soil pollution risk of enterprise land based on big data according to claim 1 is characterized in that: The basic site environment data is collected by the Internet of Things sensor, and the spectral feature data and the enterprise production process data are collected by hyperspectral imaging, and the basic site environment data, the spectral feature data and the enterprise production process data are preprocessed to obtain a preprocessed data set, including: The Internet of Things sensors collect basic site environmental data such as soil pH, moisture content, and temperature. The hyperspectral imaging equipment obtains the spectral characteristic data of the site surface soil reflectance. The enterprise database obtains the enterprise production process data of production materials, process flow, and waste disposal. The site environment basic data is subjected to a sliding median filter with a time window size of T to remove outliers, the spectral feature data is subjected to atmospheric scattering correction and geometric position correction, and the enterprise production process data is subjected to structured coding according to standard coding specifications; Resample the corrected site environment basic data according to the time series, spatially align the corrected spectral feature data with the site environment basic data, and time-align the encoded enterprise production process data with the site environment basic data; The missing values in the resampled site environment basic data are filled by the time series interpolation algorithm, the missing values in the registered spectral feature data are filled by the spatial interpolation algorithm, and the missing values in the enterprise production process data after time alignment are filled by the nearest neighbor filling method. Construct a data quality assessment system that includes data integrity indicators, data consistency indicators, and data accuracy indicators, and perform quality scoring on the filled site environment basic data, spectral feature data, and enterprise production process data; The site environment basic data, spectral feature data and enterprise production process data whose quality scores exceed a preset threshold are merged and integrated to generate the preprocessed data set.
3. The method for identifying soil pollution risk of enterprise land based on big data according to claim 1 is characterized in that: The method constructs a multi-level feature index including a basic feature layer, a pollution source feature layer and a pollution pathway feature layer according to the preprocessed data set, generates interactive features through feature transformation and feature combination processing, performs feature screening through correlation analysis and variance analysis, and obtains an optimized feature index system, including: The soil pH value, organic matter content, site utilization intensity, utilization years, and meteorological and hydrological parameters in the pre-processed data set are divided into the basic characteristic layer, the enterprise raw material composition, process type, pollutant emission, treatment process, and surrounding pollution source distribution data are divided into the pollution source characteristic layer, and the topography, hydrogeology, land use mode, and population distribution data are divided into the pollution pathway characteristic layer; Perform logarithmic transformation and standardization on the numerical features in the basic feature layer, perform one-hot encoding on the categorical features, and generate a basic feature vector; perform normalization on the numerical features in the pollution source feature layer, perform label encoding on the categorical features, and generate a pollution source feature vector; perform Z-score standardization on the numerical features in the pollution pathway feature layer, perform serial number encoding on the categorical features, and generate a pollution pathway feature vector; Combining the basic feature vector with the pollution source feature vector according to the time dimension, and combining the combination result with the pollution pathway feature vector according to the space dimension to generate the interaction feature; The Pearson correlation coefficient is calculated between each feature pair in the interactive feature, and redundant features with correlation coefficients greater than a threshold are eliminated; a one-way ANOVA is performed on the remaining features, the F statistic of each feature pair for pollutant concentration is calculated, and invalid features with F statistics less than a threshold are eliminated; Combining the features in the interactive features, constructing a feature interaction term including a first-order feature combination and a second-order feature combination, and calculating a mutual information value between the feature interaction term and the pollutant concentration; The feature combinations whose mutual information values in the feature interaction items are higher than a threshold are integrated with the remaining features to form the optimized feature index system.
4. The method for identifying soil pollution risks of enterprise land based on big data according to claim 1 is characterized in that: According to the optimization feature index system, a multi-task learning model architecture including a shared underlying network, a multi-expert network and a task-specific network is constructed using a multi-gated hybrid expert network. By setting a gating mechanism and an attention mechanism, the characteristics of heavy metal pollution, VOCs pollution and SVOCs pollution are extracted to obtain a multi-pollutant recognition model, including: Inputting the optimized feature index system into a shared underlying network consisting of a three-layer feedforward neural network, performing nonlinear transformation on the input features, and generating a shared feature representation; The shared feature representation is respectively input into the heavy metal expert network, the VOCs expert network, and the SVOCs expert network to separate and extract the features of different pollutants to form an expert feature vector; Calculating the expert network weight coefficient for the expert feature vector through a soft gating function, and weighted fusion of the expert feature vector using the weight coefficient to obtain a fused feature vector; Inputting the fused feature vector into the heavy metal task network, the VOCs task network, and the SVOCs task network, calculating the feature importance distribution of each pollutant type through the attention weight, and generating a task feature vector; Normalizing the task feature vector to generate recognition results of heavy metal pollution features, VOCs pollution features, and SVOCs pollution features to form the multi-pollutant recognition model; The output results of the multi-pollutant recognition model are compared with the true labels through the cross entropy loss function, and the loss value of each task network is calculated.
5. The method for identifying soil pollution risks of enterprise land based on big data according to claim 1 is characterized in that: Based on the multi-pollutant identification model, a pre-training method is used to initialize the shared underlying network, a model is trained through a curriculum learning strategy and an adversarial training method, and model parameters are optimized through a dynamically weighted multi-task loss function to obtain an optimized pollution identification model, including: Inputting historical soil pollution monitoring data into the shared underlying network, pre-training network parameters through unsupervised learning, and generating initialization parameters; The training samples are stratified according to the pollutant concentration thresholds of heavy metals, VOCs, and SVOCs, and a multi-level pollution intensity sample set is formed by concentration interval division. The multi-pollutant identification model is input with data and parameter adjustment in the order of pollution intensity from low to high to obtain preliminary training results; Adding adversarial perturbation samples to the preliminary training results, optimizing parameters through an adversarial training method, and generating adversarial training results; Calculate the respective loss values of the three tasks of heavy metal pollution, VOCs pollution, and SVOCs pollution in the adversarial training results, and weight them based on the task difficulty coefficient to obtain a multi-task loss value; Back-propagating the multi-task loss values to the shared underlying network, the multi-expert network, and the task-specific network, updating the network connection weights, and generating optimization parameters; The optimization parameters are updated to the multi-pollutant identification model to form the optimized pollution identification model.
6. The method for identifying soil pollution risks of enterprise land based on big data according to claim 1 is characterized in that: The SHAP value analysis method is used to perform global and local interpretability analysis on the optimized pollution identification model, key driving factors are determined by feature importance calculation, and an interpretable rule set is constructed through a decision tree proxy model to obtain model interpretation results, including: Calculate the SHAP value for each feature data input into the optimized pollution identification model, and obtain the global interpretability evaluation index through feature cumulative contribution analysis; The global explanatory evaluation indicators are grouped according to the pollutant types, and the characteristic SHAP values of heavy metal pollution, VOCs pollution, and SVOCs pollution are ranked to generate a characteristic importance ranking table; Based on the feature importance ranking table, a feature subset whose contribution is greater than a threshold is selected, and a local SHAP value is calculated for the contamination identification sample to form a local explanatory feature; Inputting the local explanatory features into a decision tree to perform feature space division, extracting feature splitting rules, and constructing a feature discrimination condition set; Combining the characteristic discrimination condition set to generate identification rule chains of heavy metal pollution, VOCs pollution, and SVOCs pollution to form the explainable rule set; The interpretable rule set is combined with the feature importance ranking table to generate the model interpretation result.
7. The method for identifying soil pollution risks of enterprise land based on big data according to claim 1 is characterized in that: According to the model interpretation results, an early warning system including a risk assessment module and a prevention and control suggestion module is constructed, data processing and analysis functions are realized through a microservice architecture, hierarchical early warning information is generated through a preset threshold system, and soil pollution risk identification results are output, including: The feature importance ranking table and the interpretable rule set in the model interpretation result are used as input data, and transmitted to the risk assessment module through the data interface for data aggregation to generate a risk assessment matrix; The risk assessment matrix is graded according to the risk level thresholds of heavy metal pollution, VOCs pollution, and SVOCs pollution, a three-level pollution risk warning standard is established, and a risk grading standard table is formed; Input the risk grading standard table into the prevention and control recommendation module, generate a corresponding list of prevention and control measures according to the characteristic discrimination condition set of different pollutant types, and construct a prevention and control response strategy; The risk assessment matrix and the prevention and control response strategy are encapsulated and transmitted through the message queue in the microservice architecture to generate a risk warning data stream; Dynamically monitor the risk level information in the risk warning data stream, and trigger a graded warning mechanism when the risk level exceeds a threshold to form the graded warning information; The graded warning information is associated with the corresponding prevention and control response strategy to form the soil pollution risk identification result.
8. The method for identifying soil pollution risks of enterprise land based on big data according to claim 7 is characterized in that: The risk assessment matrix is graded according to the risk level thresholds of heavy metal pollution, VOCs pollution, and SVOCs pollution, a three-level pollution risk warning standard is established, and a risk grading standard table is formed, including: The heavy metal pollution values in the risk assessment matrix are graded according to the pollutant content range, divided into three risk levels of light pollution, moderate pollution, and heavy pollution, and heavy metal pollution risk classification data is generated; The VOCs pollution values in the risk assessment matrix are graded according to the concentration range of volatile organic compounds, and divided into three risk levels of light pollution, moderate pollution, and heavy pollution, to generate VOCs pollution risk classification data; The SVOCs pollution values in the risk assessment matrix are graded according to the concentration range of semi-volatile organic compounds, and divided into three risk levels of light pollution, moderate pollution, and heavy pollution, to generate SVOCs pollution risk grading data; The heavy metal pollution risk grading data, VOCs pollution risk grading data, and SVOCs pollution risk grading data are combined according to the pollutant type, and the warning color mark of each risk level is marked to generate a risk level warning matrix; Comprehensively scoring the risk level of each pollutant type in the risk level warning matrix, calculating the comprehensive risk index according to the risk weight, and generating a comprehensive risk assessment table; The risk level warning matrix is combined with the comprehensive risk assessment table to generate the risk grading standard table.
9. The method for identifying soil pollution risks of enterprise land based on big data according to claim 7 is characterized in that: The step of dynamically monitoring the risk level information in the risk warning data stream and triggering a graded warning mechanism when the risk level exceeds a threshold to form the graded warning information includes: Segment the risk warning data stream according to time windows, calculate the change rate of heavy metal pollution, VOCs pollution, and SVOCs pollution risk index within each time window, and generate risk change trend data; Compare the risk change trend data with the thresholds in the risk grading standard table, mark the time points that exceed the light pollution, moderate pollution, and heavy pollution thresholds, and form a time series table of risk breakthrough points; The continuous breakthrough time periods in the risk breakthrough point time series table are counted, the duration of the risk level is recorded according to the pollutant type, and a risk continuity assessment table is generated; The data in the risk continuity assessment table are grouped according to pollutant type and risk level, and the frequency of occurrence of risk events of each level is counted to form a risk frequency statistical table; Setting warning trigger conditions for the data in the risk frequency statistics table, marking high-frequency risk events as key warning objects, and generating a warning object list; The warning object list is combined with the corresponding risk change trend data to form the graded warning information.
10. A big data-based enterprise land soil pollution risk identification system, used to implement the enterprise land soil pollution risk identification method based on big data as described in any one of claims 1 to 9, characterized in that: The enterprise land soil pollution risk identification system based on big data includes: An acquisition module is used to collect basic site environment data through IoT sensors, and to collect spectral feature data and enterprise production process data through hyperspectral imaging, and to preprocess the basic site environment data, spectral feature data and enterprise production process data to obtain a preprocessed data set; A transformation module is used to construct a multi-level feature index including a basic feature layer, a pollution source feature layer and a pollution pathway feature layer according to the preprocessed data set, generate interactive features through feature transformation and feature combination processing, perform feature screening through correlation analysis and variance analysis, and obtain an optimized feature index system; An extraction module is used to construct a multi-task learning model architecture including a shared underlying network, a multi-expert network and a task-specific network using a multi-gated hybrid expert network based on the optimized feature index system, extract the characteristics of heavy metal pollution, VOCs pollution and SVOCs pollution by setting a gating mechanism and an attention mechanism, and obtain a multi-pollutant recognition model; A training module, for initializing a shared underlying network using a pre-training method based on the multi-pollutant recognition model, training the model using a curriculum learning strategy and an adversarial training method, optimizing model parameters via a dynamically weighted multi-task loss function, and obtaining an optimized pollution recognition model; An analysis module is used to perform global and local interpretative analysis on the optimized pollution identification model using a SHAP value analysis method, determine key driving factors by feature importance calculation, construct an interpretable rule set via a decision tree proxy model, and obtain a model interpretation result; The early warning module is used to construct an early warning system including a risk assessment module and a prevention and control suggestion module based on the results of the model interpretation, realize data processing and analysis functions through the microservice architecture, generate graded early warning information through a preset threshold system, and output soil pollution risk identification results.
Citation Information
Cited By
Multi-physical field regulation and control method and system in aluminum product processing
CN120540253A
Information generation method and device
CN120706964A
Information generation method and apparatus
CN120706964B
Analog quantity wireless acquisition system based on dynamic power consumption management
CN120825757A
Analog quantity wireless acquisition system based on dynamic power management
CN120825757B