Big Data Analysis System for Soil Environmental Damage and Risk Assessment
By building a big data analysis system for soil environmental damage and risk assessment, and using multi-sensor groups and enhanced Lasso regression modeling, the comprehensiveness and real-time inadequate soil pollution assessment in the existing technology are solved, and efficient and accurate risk assessment and automated early warning of multiple pollutants are achieved.
Patent Information
- Application Number
- CN202510157830.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-02-13
AI Technical Summary
The existing soil pollution assessment technology has shortcomings in comprehensive, real-time and multi-pollutant analysis, and it is difficult to reflect the diffusion process and its changing trends in real time, and there is a lack of effective analysis of the complex relationship between multi-pollutants and their environmental factors.
Build a big data analysis system for soil environmental damage and risk assessment, through the coordinated work of data acquisition, processing and fusion, spatial correlation analysis and risk assessment modules, use multiple sensor groups to obtain multiple pollutant data, and combine enhanced Lasso regression modeling to build a risk prediction model to achieve efficient collection, standardized processing, dynamic analysis and accurate risk prediction of multiple pollutant concentrations and soil environmental parameters.
It significantly improves the comprehensiveness, real-timeness and accuracy of soil pollution assessment, has dynamic response and automated early warning functions, and can accurately assess the risks of multiple pollutants and identify key factors, providing a basis for governance plans.
Smart Images

Figure CN119619466B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a big data analysis system for soil environmental damage and risk assessment. Background Art
[0002] The assessment and treatment of soil environmental damage have become a key area of environmental protection that has attracted global attention. However, due to the complex spatial distribution of soil pollution, diverse types of pollutants, and the dynamic characteristics of migration and diffusion, existing soil risk assessment technologies still have many limitations in practical applications. Currently, various methods and technologies have been developed in the fields of soil pollution assessment and risk management, including pollutant concentration monitoring, soil physical and chemical property analysis, and pollution diffusion simulation, etc.
[0003] Traditional soil pollution monitoring usually relies on manual sampling and laboratory analysis to determine the concentration of heavy metals or organic pollutants at specific sampling points. This method has high accuracy in quantitative detection of pollutants, but its monitoring range is limited and it cannot comprehensively reflect the spatial distribution characteristics of pollutants. In addition, due to the limited number of sampling points, the results are easily affected by abnormal data at individual points, leading to deviations in the assessment results. In recent years, with the progress of data processing technology, the integrated analysis of multi-dimensional pollutant data has gradually become a research hotspot. For example, researchers have tried to comprehensively analyze the concentrations of various pollutants in soil, soil physical and chemical parameters (such as pH, cation exchange capacity, etc.), and topographic features to reveal the correlation between pollutants and their diffusion laws. Such methods have improved the comprehensiveness of assessment to a certain extent, but due to the lack of a dynamic update mechanism, they cannot reflect the diffusion process and its changing trend of pollutants in real time. Methods based on machine learning and statistical modeling have begun to be applied to soil pollution risk prediction. Random forests, support vector machines, and Lasso regression models have been widely used to predict the concentration or risk level of pollutants in soil. Such methods can improve the accuracy of prediction to a certain extent, but the construction of the model usually relies on a large amount of historical data and lacks the ability to respond to real-time data. In addition, most traditional prediction models only consider single pollutants or environmental factors and have insufficient analysis of the complex relationships between multi-pollutants and environmental sensitivity. Summary of the Invention
[0004] The main objective of the present invention is to provide a big data analysis system for soil environmental damage and risk assessment. Through the collaborative work of data acquisition, processing and fusion, spatial correlation analysis, and risk assessment modules, it realizes the efficient collection, standardized processing, dynamic analysis, and accurate risk prediction of the concentrations of various pollutants and soil environmental parameters. Based on enhanced Lasso regression modeling and combined with environmental sensitivity, spatial distribution characteristics, and sparse features, the system constructs an intelligent risk assessment and early warning mechanism. Compared with the prior art, the present invention significantly improves the comprehensiveness, real-time performance, and accuracy of the assessment, and at the same time has functions of dynamic response, multi-dimensional data integration, and automatic early warning, providing efficient, intelligent, and reliable technical support for soil pollution control and environmental protection.
[0005] To solve the above problems, the technical solution of the present invention is implemented as follows:
[0006] A big data analysis system for soil environmental damage and risk assessment, the system includes: a data acquisition module, a data processing and fusion module, a spatial correlation analysis module, and a risk assessment module; the data acquisition module is used to set multiple sampling points in the target soil area, deploy the same sensor group at each sampling point, the sensor group includes multiple sensors, and each sensor respectively acquires data of a type of pollutant in the sampling point and sends the acquired data to the cloud big database; the data processing and fusion module is used to obtain the historical data acquired at the sampling time before the current sampling time from the cloud big database, perform standardized and fusion processing on the historical data to obtain the preprocessed data of each sampling point, perform high-dimensional sparse feature extraction on the preprocessed data, and remove redundant information layer by layer by constructing two-layer feature dictionaries to extract the sparse features of the preprocessed data; the spatial correlation analysis module is used to analyze the variation law of pollutant concentration with time and space based on the preprocessed data of each sampling point, and reveal the spatial diffusion characteristics of pollutants between sampling points by constructing a spatial weight matrix and a time variation model to obtain a spatial correlation index; the risk assessment module is used to calculate the environmental sensitivity of the target soil area according to the environmental parameters of the target soil area, combine enhanced Lasso regression modeling, associate the spatial correlation index, environmental sensitivity, and sparse features, construct a risk prediction model; and use the data vector composed of the data of the pollutants acquired at the sampling point at the current time as the input and input it into the risk prediction model to obtain the current soil environmental risk value corresponding to each sampling point.
[0007] Further, the area range of the target soil area is 50 to 500 square kilometers; in the target soil area, sampling points are set using a uniform grid, and the range of the grid side length is 10 to 50 meters; the sensor group at each sampling point conducts sampling once a day to obtain pollutant data.
[0008] Further, the data of the pollutants include: cadmium concentration, lead concentration, arsenic concentration, mercury concentration, copper concentration, zinc concentration, nickel concentration, chromium concentration, polycyclic aromatic hydrocarbon concentration, polychlorinated biphenyl concentration, nitrate concentration, sulfate concentration, chloride concentration, ammonium nitrogen concentration, and total petroleum hydrocarbon concentration; the environmental parameters are averages, including: average organic matter content, average cation exchange capacity, average pH value, average soil bulk density, average soil permeability coefficient, average soil electrical conductivity, average clay content, average silt content, and average terrain slope of the target soil area.
[0009] Further, the historical data is standardized and fused through the following formula:
[0010] ;
[0011] where is the preprocessed data of the th sampling point; is the number of pollutant types; is the original concentration of the th pollutant obtained at the previous sampling time before the current sampling time at the th sampling point; is the historical average concentration of the th pollutant; is the historical standard deviation of the th pollutant; is the historical highest concentration of the th pollutant obtained at the th sampling point; is the historical lowest concentration of the th pollutant obtained at the th sampling point.
[0012] Further, through the following formula, by constructing a two-layer feature dictionary, redundant information is removed layer by layer to extract the sparse features of the preprocessed data:
[0013] ;
[0014] By solving the minimum optimization problem represented by this formula, , and are obtained, and based on this, the sparse feature of the preprocessed data is calculated as: ; where is the preprocessing vector composed of the preprocessed data of all sampling points; where is the number of sampling points; is the first-layer feature dictionary; is the second-layer feature dictionary; is a sparse coefficient matrix, and is a preset diagonal matrix; represents the L2 norm; represents the L1 norm; is the Laplace operator; is the absolute value operator; is the F norm.
[0015] Furthermore, the spatial correlation index is calculated through the following formula :
[0016] ;
[0017] where represents the Euclidean distance between the th sampling point and the th sampling point; represents the preprocessed data of the th sampling point.
[0018] Furthermore, the environmental sensitivity of the target soil area is calculated through the following formula according to the environmental parameters of the target soil area :
[0019] ;
[0020] where is the average organic matter content; is the average cation exchange capacity; is the average pH value; is the average soil bulk density; is the average soil permeability coefficient; is the average soil conductivity; and are the average clay content and average silt content of the soil respectively; is the average value of the terrain slope; is the gamma function.
[0021] Furthermore, through the following formula, combined with enhanced Lasso regression modeling, the spatial correlation index, environmental sensitivity and sparse features are associated to construct a risk prediction model:
[0022] ;
[0023] where is the data vector composed of the pollutant data obtained at the th sampling point at the current time; This risk prediction model is a minimum value optimization problem, and by solving the risk prediction model, the risk prediction value ; To calculate the determinant value of a matrix; It represents calculating the modulus of a vector.
[0024] Furthermore, the risk assessment module compares the risk prediction value with a preset threshold. If it exceeds the set threshold, it is determined that the target soil area has risks and a warning is issued.
[0025] The big data analysis system for soil environmental damage and risk assessment of the present invention has the following beneficial effects:
[0026] The system of the present invention realizes high-precision monitoring of the concentrations of various pollutants (such as heavy metals, organic pollutants, volatile organic compounds, salts, etc.) by setting multiple sampling points in the target soil area and arranging multifunctional sensor groups at each sampling point. By introducing a dynamic sampling mechanism, the system can capture the temporal and spatial changes of pollutant concentrations in real time. Compared with the traditional single-point sampling or intermittent monitoring methods, the present invention significantly improves the comprehensiveness and continuity of data.
[0027] In addition, through the data processing and fusion module, the system standardizes and fuses the collected multi-dimensional pollutant data, solves the problem of inconsistent dimensions of different pollutant concentrations, and removes redundant information through sparse feature extraction technology to extract core features. This processing method ensures the integration quality of the data and provides high-quality input for subsequent analysis. Especially in the case where the interaction relationships of multiple pollutants are complex, the fusion processing method of the present invention can accurately reveal the coupling effects between pollutants and makes up for the deficiencies in the analysis of multiple pollutants in the prior art.
[0028] In the risk assessment process of the present invention, environmental sensitivity factors of the target area are fully integrated, including key parameters such as soil pH, organic matter content, cation exchange capacity, permeability coefficient, bulk density, and terrain slope. Through the environmental sensitivity calculation formula, the system performs non-linear modeling on these parameters to quantify the adsorption, diffusion, and accumulation capabilities of the soil to pollutants. Compared with the prior art, the present invention more comprehensively considers the regulatory effects of soil physical and chemical properties on pollutant behavior, thereby improving the scientificity of risk assessment. Especially in a complex environment with co-existing multiple pollutants, the impact of soil environmental sensitivity on risk assessment is particularly important. The present invention combines environmental sensitivity with pollutant concentration, spatial distribution characteristics, and sparse feature matrix to construct a multi-dimensional and multi-factor risk assessment framework. This comprehensive analysis method can not only accurately assess the pollution risks of the target area, but also identify the key factors that have the greatest impact on environmental sensitivity, providing a basis for the priority ranking of treatment plans. Description of the Drawings
[0029] Figure 1Schematic diagram of the system structure of the big data analysis system for soil environmental damage and risk assessment provided by the embodiments of the present invention. Detailed implementation manners
[0030] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0031] Example 1, refer to Figure 1 : A big data analysis system for soil environmental damage and risk assessment, the system includes: a data acquisition module, a data processing and fusion module, a spatial correlation analysis module, and a risk assessment module; the data acquisition module is used to set a plurality of sampling points in the target soil area, and arrange the same sensor group at each sampling point. The sensor group includes a variety of sensors, and each sensor respectively acquires data of one type of pollutant at the sampling point, and sends the acquired data to the cloud big database; the data processing and fusion module is used to obtain the historical data acquired at the sampling time before the current sampling time from the cloud big database, perform standardization and fusion processing on the historical data to obtain the preprocessed data of each sampling point, perform high-dimensional sparse feature extraction on the preprocessed data, and remove redundant information layer by layer by constructing a two-layer feature dictionary to extract the sparse features of the preprocessed data; the spatial correlation analysis module is used to analyze the variation law of pollutant concentration with time and space based on the preprocessed data of each sampling point, and reveal the spatial diffusion characteristics of pollutants between sampling points by constructing a spatial weight matrix and a time change model to obtain a spatial correlation index; the risk assessment module is used to calculate the environmental sensitivity of the target soil area according to the environmental parameters of the target soil area, and combine enhanced Lasso regression modeling to associate the spatial correlation index, environmental sensitivity with the sparse features to construct a risk prediction model; and use the data vector composed of the data of the pollutants acquired at the sampling points at the current time as input, and input it into the risk prediction model to obtain the current soil environmental risk value corresponding to each sampling point.
[0032] Specifically, the data acquisition module is the foundation and starting point of the entire system. Its core principle lies in providing accurate and comprehensive raw data support for subsequent data processing, spatial analysis, and risk assessment through efficient multi-dimensional sampling and real-time monitoring. To achieve this goal, multiple sampling points are arranged in the target soil area, and a sensor group is installed at each sampling point. The design concept of the sensor group combines diversity and precision. Each type of sensor is used to monitor a specific type of pollutant, such as heavy metals, organic pollutants, volatile organic compounds, salts, or inorganic substances. Through the collaborative work of these multi-sensors, the system can achieve high-coverage and high-precision monitoring of pollutants in the target area. Specifically, this module combines Internet of Things technology and cloud computing architecture. Each sensor group is connected to a regional gateway device through wireless transmission or wired connection, and the data is uploaded to the cloud big database through the gateway. During the selection and arrangement of sensors, the present invention pays attention to the diversity of pollutant characteristics and the complexity of the soil environment. For example, for the monitoring of heavy metal pollution, inductively coupled plasma mass spectrometry (ICP-MS) or spectroscopic sensors are used to capture the trace changes of heavy metals such as cadmium and lead in the soil with high sensitivity; for organic pollutants, portable gas chromatographs or gas sensors are selected to monitor the concentration of polycyclic aromatic hydrocarbons or organochlorine pesticides; for salts and inorganic pollutants, ion-selective electrode sensors are used to quickly detect key indicators such as nitrates and ammonium nitrogen. In addition, to capture the dynamic change characteristics of pollutants, the data acquisition module designs two modes: timed sampling and continuous monitoring. Timed sampling is suitable for pollutants with slower changes, while continuous monitoring is used to capture transient concentration changes such as volatile organic compounds.
[0033] The data acquisition module not only focuses on the collection of single-point data, but also realizes the scientificity and representativeness of spatial distribution by optimizing the layout scheme of sampling points. Specifically, based on the topography, historical pollution distribution data and potential pollution source locations of the target area, the module adopts various sampling strategies, such as regular grid sampling, random sampling or hot spot sampling, to ensure that the distribution of sampling points can comprehensively reflect the distribution characteristics of pollutants in the area. For areas with large pollutant concentration gradients, such as around industrial pollution sources or areas with intensive agricultural fertilization, the sampling point spacing will be relatively small to capture the detailed distribution of pollution. For areas with small concentration gradient changes, the sampling point spacing can be appropriately increased to improve the monitoring efficiency while reducing costs. In addition, by dynamically adjusting the sampling point position and density, the module can quickly respond to sudden pollution events, flexibly increase the number of sampling points in high-risk areas, and improve the timeliness and accuracy of data. To achieve efficient data transmission and storage, the data acquisition module constructs a distributed data transmission and cloud storage mechanism. After the sensors complete data collection, the raw data is transmitted to the cloud platform in real time through the regional gateway and undergoes preliminary processing during transmission, such as noise filtering, outlier detection and data compression. These processing steps are based on preset algorithm models. For example, the filtering algorithm for heavy metal data can effectively remove accidental measurement errors of sensors, and the outlier detection of organic pollutant concentrations marks suspicious values through statistical models. The processed data is uniformly stored in the cloud database. The design of the database supports efficient indexing and rapid retrieval of multi-dimensional data, thus providing high-quality data support for subsequent analysis modules. In addition, the data acquisition module also realizes the synchronous collection of pollutant concentrations and environmental factors by introducing environmental parameter monitoring. These environmental parameters include soil pH value, organic matter content, temperature and humidity, etc., which are closely related to the mobility and bioavailability of pollutants. During actual sampling, the monitoring of environmental parameters is completed by dedicated sensors and recorded synchronously with the pollutant concentration data to ensure that the time and space dimensions of both are completely matched. This multi-dimensional data collection method enables the system to comprehensively consider the dynamic characteristics of pollutant behavior in subsequent analysis and improve the scientificity of risk assessment.
[0034] The first step in data processing is to standardize the original pollutant concentration data from the data acquisition module. For different types of pollutants, such as heavy metals, organic pollutants, salts, etc., there are significant differences in their concentration ranges, physical units, and distribution characteristics. This difference will directly affect the accuracy of subsequent analysis. The goal of standardization is to eliminate these differences so that data of different pollutants can be compared and analyzed under a unified dimension. During the standardization process, the system uses methods of interval scaling and distribution adjustment to map the concentration values of each pollutant into the same normalized interval while retaining their original distribution characteristics to avoid losing important information due to data transformation. In addition, a time variation factor is introduced during the standardization process to quantify the dynamic changes in pollutant concentrations. This process fully considers that the concentrations of some pollutants are affected by seasonal factors and environmental condition changes, and ensures that the processed data can truly reflect the spatial characteristics of pollutants by dynamically adjusting weights. After completing the standardization, the system enters the data fusion stage. The fusion of pollutant data is one of the core tasks of this module, and its purpose is to integrate the concentration data of different sampling points, different time periods, and multiple pollutants into a multi-dimensional feature matrix. During the fusion process, the system correlates the spatial positions of sampling points, the characteristic weights of pollutants, and time dynamics based on a multi-source data weighting model. Specifically, the system assigns different weights to data points in space according to the migration and diffusion characteristics of pollutants. For example, in the analysis of heavy metal pollution, considering its slow migration speed, the system assigns lower weights to data from long-distance sampling points in space and higher weights to data from nearby points. Similarly, for pollutants with strong diffusivity such as volatile organic compounds, the system adopts different weight distribution strategies. This weighted fusion method can effectively balance the characteristic differences of different pollutants, enabling the fused data matrix to reflect both the overall distribution law of pollutants and retain local dynamic change characteristics. Based on the data fusion, the module further performs high-dimensional sparse feature extraction. Since pollutant data usually has high dimensionality and redundancy, directly using the original fused data for analysis may lead to too high computational complexity and the risk of information redundancy. Therefore, the system constructs a two-layer feature dictionary through sparse feature extraction technology to achieve hierarchical dimensionality reduction of data and extraction of key features. The first-layer feature dictionary is used to extract the basic features of each pollutant, such as its average concentration, spatial distribution trend, etc.; the second-layer feature dictionary extracts the coupling features and correlation patterns between multiple pollutants through cross-pollutant association analysis. This two-layer feature extraction mechanism can not only reduce the data dimension and improve computational efficiency but also retain the potential connections between pollutants, providing more accurate feature inputs for subsequent spatial analysis and risk modeling.
[0035] The principle of the spatial correlation analysis module is first based on the spatial weight reconstruction modeling of pollutant concentration data at multiple sampling points within the target area. In this process, the system constructs a spatial weight matrix based on the geographical locations of the sampling points to measure the degree of mutual correlation between different sampling points. In specific operations, the setting of spatial weights is usually based on the distance decay principle, that is, the spatial correlation between two sampling points weakens as the distance increases. The system uses the Gaussian function or the inverse distance weight model to assign weights to the geographical distances between sampling points, thereby constructing a correlation network that reflects the characteristics of spatial diffusion. For example, if the spatial distances between two sampling points are close and the pollutant concentration trends are similar, their spatial weights are higher, indicating that there may be a diffusion or transfer relationship between them; conversely, lower weights indicate weaker correlation. After constructing the spatial weight matrix, the module further calculates the spatial gradient change of the pollutant concentration to characterize the diffusion direction and intensity of the pollutant within the target area. The spatial gradient is an important indicator reflecting the change of pollutant concentration. It reveals the trend of pollutant diffusion from high-concentration areas to low-concentration areas by calculating the change rate of concentration with respect to spatial coordinates. To more accurately describe this change, the module introduces the Laplace operator to analyze the characteristics of the diffusion field of pollutants within the area. The result of the Laplace operator can intuitively reflect the diffusion intensity of pollutants. For example, when the Laplace value of a certain area is large, it indicates that the pollutant diffusion in this area is relatively intense; while a small or near-zero Laplace value indicates that the pollutant concentration in this area is relatively uniform and the diffusion trend is not obvious. In addition, to more comprehensively quantify the spatial diffusion law, the spatial correlation analysis module also integrates the potential field model. The potential field is used to describe the diffusion potential and accumulation trend of pollutants in the soil, especially having significant advantages in cases of complex terrain or uneven distribution of pollution sources. In practical applications, the module generates a pollution potential field based on the pollutant concentration data of the sampling points, and determines the main diffusion paths and boundary effects of pollutants by calculating the gradient and divergence of the potential field. For example, when the gradient value of a certain potential field is large, it indicates that the pollutant has a strong diffusion ability; while in areas with a small gradient, there may be an accumulation effect of pollutants. Potential field analysis can not only reveal the overall diffusion trend of pollutants within the area, but also help identify the source or convergence point of diffusion, providing specific guidance for subsequent pollution control.
[0036] The core of the risk assessment module lies in integrating various data output by the system front-end module into a unified risk prediction model. In this process, it is first necessary to calculate the environmental sensitivity, which is an important basis for evaluating regional risks. The calculation of environmental sensitivity depends on soil physical and chemical properties (such as pH value, organic matter content, cation exchange capacity, etc.) and regional ecological factors (such as vegetation coverage, biodiversity, etc.). Soil pH value has a significant impact on the bioavailability and toxicity of heavy metals. For example, under acidic conditions, the solubility of cadmium and lead will increase significantly, resulting in higher mobility and biotoxicity. The organic matter content determines the adsorption capacity of the soil for pollutants. High organic matter content can usually fix more heavy metals or organic pollutants, reducing their migration risk. By weighted combination of these physical and chemical parameters, the system can generate a sensitivity index that comprehensively reflects the regional soil environmental status, indicating the pollutant carrying capacity and potential ecological threats of the region. Based on the environmental sensitivity, the risk assessment module further incorporates the spatial correlation index provided by the spatial association analysis module into the model to reveal the diffusion range and dynamic characteristics of pollutants. The spatial correlation index not only reflects the correlation strength between pollutants at different sampling points, but also can quantify the pollution diffusion path and the concentration degree of high-risk areas. For example, through the spatial weight matrix and gradient analysis results, the system can identify the main diffusion trend of pollutants from high-concentration areas to low-concentration areas. Combining this information, the system can evaluate whether the pollutant concentration in the target area is likely to exceed the environmental threshold, or whether there is a cumulative effect that exerts long-term pressure on the ecosystem.
[0037] Example 2: The area range of the target soil region is 50 to 500 square kilometers; in the target soil region, sampling points are set using a uniform grid, and the range of the grid side length is 10 to 50 meters; the sensor group at each sampling point conducts sampling once a day to obtain pollutant data.
[0038] The target area has a relatively large area range, covering different types of soil environments and potential pollution source distributions. To achieve the uniformity and representativeness of data collection, the monitoring plan adopts a grid-based sampling method. Specifically, according to the complexity of the target area and the requirements of evaluation accuracy, a uniform grid with a grid side length ranging from 10 to 50 meters is selected. The setting of this range mainly considers the following factors: when the pollutant concentration gradient in the area is large or the pollution source distribution is dense, a smaller grid side length (such as 10 meters) is used to capture the fine distribution characteristics of pollutants; while in areas where the pollutant concentration is relatively uniform, the grid side length is appropriately increased (such as 50 meters) to improve the monitoring efficiency and reduce the sampling cost. Through this flexible grid sampling method, the accuracy and economy of data collection can be effectively balanced. At each sampling point, a set of multi-functional sensor groups are deployed to collect data on different types of pollutants in the soil. The design of the sensor group can simultaneously monitor the concentrations of various pollutants such as heavy metals, organic pollutants, volatile organic compounds, and salts. Heavy metal sensors usually adopt inductively coupled plasma mass spectrometry (ICP-MS) or spectroscopic sensors, which can detect elements such as cadmium, lead, and arsenic with high sensitivity; organic pollutant sensors are based on gas chromatography or electrochemical detection technologies to accurately obtain indicators such as polycyclic aromatic hydrocarbons or organochlorine pesticides. The sampling frequency of once a day fully considers the time characteristics of pollutant concentration changes and can balance the timeliness of data and the durability of sensors. For pollutants with large transient changes such as volatile pollutants (such as benzene, toluene, etc.), this sampling frequency can meet the general environmental assessment requirements; while for pollutants with slower concentration changes such as heavy metals, the daily sampling frequency is sufficient to cover their temporal variation patterns. After each sampling, the sensor group will transmit the collected pollutant data to the regional gateway through wireless or wired networks and then upload it to the cloud big database. During the transmission process, the system will perform preliminary processing on the data, such as removing noise, detecting outliers, and data compression, to ensure the quality of the uploaded data. The cloud database uniformly stores and manages the daily data collected, providing efficient data support for subsequent data processing, spatial analysis, and risk assessment.
[0039] Example 3: The data of the pollutants include: cadmium concentration, lead concentration, arsenic concentration, mercury concentration, copper concentration, zinc concentration, nickel concentration, chromium concentration, polycyclic aromatic hydrocarbon concentration, polychlorinated biphenyl concentration, nitrate concentration, sulfate concentration, chloride concentration, ammonium nitrogen concentration, and total petroleum hydrocarbon concentration; the environmental parameters are averages, including: average organic matter content, average cation exchange capacity, average pH value, average soil bulk density, average soil permeability coefficient, average soil conductivity, average clay content, average silt content, and the average value of the terrain slope of the target soil area.
[0040] Specifically, the composition of pollutant data covers a variety of key pollutant types, which are divided into three categories: heavy metals, organic pollutants, and inorganic pollutants. In the heavy metal data, the concentrations of cadmium (Cd), lead (Pb), arsenic (As), mercury (Hg), copper (Cu), zinc (Zn), nickel (Ni), and chromium (Cr) are included. These heavy metals are common soil pollutants, and their sources include industrial waste, mining emissions, pesticide residues, etc. Each heavy metal has different behavioral characteristics. For example, cadmium and lead have strong mobility in soil, while chromium has relatively weak mobility but significant toxicity. These differences need to be quantified through feature analysis and modeling processes in the system to reveal their distribution patterns and ecological risks in the target area. Among organic pollutants, the concentrations of polycyclic aromatic hydrocarbons (PAHs) and polychlorinated biphenyls (PCBs) are included. Both types of pollutants are persistent organic pollutants with high toxicity and long-term environmental residual characteristics. The main sources of PAHs are fossil fuel combustion and industrial emissions, while PCBs come from waste industrial materials and electronic waste. Their adsorption and degradation characteristics in soil are extremely important to the environmental impact, and their spatial distribution trends need to be quantified through the dynamic analysis module in the system.
[0041] For inorganic pollutants, the system monitored the concentrations of nitrate (NO3⁻), sulfate (SO4²⁻), chloride (Cl⁻), ammonium nitrogen (NH4⁺), and total petroleum hydrocarbons (TPH). These indicators can reflect the residues of nutrients in the soil and the potential pollution caused by industrial emissions. In particular, the concentrations of nitrate and ammonium nitrogen can be used to evaluate the impact of agricultural activities on soil quality, while total petroleum hydrocarbons are the core indicators for measuring petroleum pollution, commonly found in oil fields, transportation leaks, or industrial waste sites. In addition to pollutant data, the system also defined a series of average environmental parameters, which can help quantify the basic state of the soil environment and its impact on pollutant behavior. First, the average organic matter content reflects the adsorption capacity of the soil. High organic matter content usually enhances the fixation of heavy metals and organic pollutants in the soil. The average cation exchange capacity (CEC) measures the ability of the soil to fix metal ions. Soils with higher CEC have stronger adsorption capacity for cations, which helps reduce the mobility of heavy metals. The average pH value is an important parameter in the physical and chemical properties of the soil, directly affecting the solubility and bioavailability of pollutants. For example, in acidic soils, the solubility of cadmium and lead will increase significantly, thus enhancing their ecological risks, while in alkaline soils, the mobility of most heavy metals will decrease significantly. The bulk density and permeability coefficient of the soil reflect the compactness and water infiltration capacity of the soil, affecting the horizontal diffusion and vertical migration rates of pollutants. The average electrical conductivity (EC) is used to characterize the distribution of soil salinity. Soils with higher electrical conductivity may indicate higher salt concentrations, which are closely related to the distribution of nitrate and chloride. In terms of physical properties, the average clay content and average silt content reflect the distribution characteristics of soil particle composition. Higher clay content is usually accompanied by higher adsorption capacity and lower permeability, while soils with higher silt content are more prone to pollutant migration. In addition, the average terrain slope of the target soil area is a key spatial parameter, determining the intensity of surface runoff and the possibility of pollutant loss. In areas with larger slopes, pollutants may be more likely to enter adjacent water bodies or downstream soil areas with surface runoff, increasing the diffusion range and the risk of secondary pollution.
[0042] Example 4: Standardize and fuse the historical data through the following formula:
[0043] ;
[0044] where is the preprocessed data of the th sampling point; is the number of pollutant types; is the concentration of the th sampling point at the previous sampling time before the current sampling time for the th pollutant; is the The historical average concentration of a certain pollutant; is the historical standard deviation of the th pollutant; is the historical highest concentration of the th pollutant obtained at the th sampling point; is the
[0045] Specifically, the core of the formula lies in normalizing and weighted-fusing the concentration data of multiple pollutants at each sampling point, and finally generating a comprehensive preprocessed data , used to characterize the pollutant state of the th sampling point. In the specific implementation, first, the historical data is normalized, that is, the concentration value of each pollutant is converted into a relative value with respect to the historical highest value and the lowest value at this sampling point. The purpose of this step is to eliminate the numerical bias caused by the magnitude difference of different pollutants, so that the data of all pollutants can be compared and analyzed on a unified scale. Through the formula , the concentration value of each pollutant is normalized to the interval , thus ensuring that the weights of each pollutant in the calculation will not be affected by the dimension difference. After normalization, the system further performs weighted processing on the data in combination with the historical fluctuation characteristics of the pollutants. Specifically, the Gaussian distribution function is used to adjust the normalized concentration value of each pollutant, and calculate its deviation weight from the historical average concentration . Through , the system can highlight those pollutants whose concentration values are close to the historical average level, and at the same time attenuate the outliers with a large degree of dispersion. This processing method is based on the statistical characteristics of the pollutant distribution, which can effectively reduce the interference of abnormal points on the comprehensive data, and at the same time retain the characteristic information of the main pollutants. For each sampling point, the system performs weighted summation on the data of all pollutants to generate the comprehensive preprocessed data and in the formula, a comprehensive expression that takes into account both the concentration level and the historical volatility is constructed. The number of pollutant types determines the complexity and diversity of the comprehensive data, and can adapt to the soil environment analysis scenarios with multiple pollutants existing simultaneously.
[0046] Example 5: Through the following formula, by constructing a two-layer feature dictionary, redundant information is removed layer by layer to extract the sparse features of the preprocessed data:
[0047] ;
[0048] By solving the minimum optimization problem represented by this formula, we obtain 、 and , and calculate the sparse features of the preprocessed data as: ; where is the preprocessing vector composed of the preprocessed data of all sampling points; where is the number of sampling points; The first-layer feature dictionary; is the second-layer feature dictionary; is the sparse coefficient matrix, which is a preset diagonal matrix; represents the L2 norm; represents the L1 norm; is the Laplace operator; is the absolute value operator; is the F norm.
[0049] Specifically, the first term of the formula is the core part of the objective function, representing the error between the preprocessed data and the reconstructed data jointly affected by the two-layer feature dictionary and the sparse coefficient matrix . Here, the norm is used for error measurement, aiming to minimize the Euclidean distance between the original data and the reconstructed data, ensuring that the reconstruction result can reflect the feature distribution of the original data as realistically as possible. Through the optimization of this term, the system can dynamically adjust the content of the feature dictionary and the sparsity of the sparse coefficient matrix, so that the feature extraction result not only maintains the integrity of the data but also effectively compresses the redundant information. The second term of the formula is the sparse regularization term in the objective function, and its main role is to constrain the sparsity of the sparse coefficient matrix , ensuring that each feature dictionary represents the data sparsely and efficiently. Here, the norm is used to constrain the sparse coefficients, which can effectively suppress unimportant feature variables and focus the feature extraction on the dimensions that are crucial for data description. The introduction of sparse regularization is particularly suitable for multi-dimensional pollutant data analysis because the impact of some pollutants on the overall soil risk is relatively small. Through sparse regularization, these secondary information can be automatically filtered, thereby improving the computational efficiency and result interpretability of the system. The third term is a dictionary smoothing regularization term, which is unique in introducing the Laplace operator , which is used to measure the smoothness and consistency of the two-layer feature dictionary. The Laplace operator can suppress high-frequency noise or abnormal data in the feature dictionary, ensuring the smoothness and stability of the feature extraction results. By calculating the difference in the Laplace values of the first-layer dictionary and the second-layer dictionary and constraining it, the system can dynamically balance the complexity and detail retention of the two-layer features, avoiding excessive deviation in feature representation between the two-layer dictionaries. The optimization of this term ensures the multi-level effect of feature extraction, while improving the adaptability of the model to complex pollution scenarios. The fourth term is the feature dictionary hierarchical weight constraint term. By constraining the ratio of the Frobenius norms of the two-layer dictionaries, the system can dynamically adjust the weight distribution of the first-layer and second-layer feature dictionaries. The Frobenius norm measures the overall size of the dictionary matrix, and its ratio reflects the relative importance of the two-layer feature dictionaries in data representation. When the ratio is large, it indicates that the second-layer dictionary makes a more significant contribution to feature extraction, which is suitable for scenarios with strong correlation of complex pollutants; while when the ratio is small, the first-layer dictionary will dominate feature extraction, which is suitable for scenarios where the pollutant characteristics are relatively independent. This design introduces an adaptive weight allocation mechanism for feature extraction, ensuring that the model can optimize the feature extraction process according to the characteristics of different pollution environments. By solving the minimum value of the above objective function, the system can obtain the two-layer feature dictionaries and and the sparse coefficient matrix of the optimal solution. Further calculate the sparse feature , and map the pollutant preprocessing data to the sparse feature space. The sparse feature can highly summarize the internal patterns of the pollutant data in the target area, including global distribution characteristics and local features. By introducing the hierarchical structure of the two-layer dictionary, the system can not only effectively reduce the data dimension, but also deeper explore the complex coupling relationships between multiple pollutants, providing scientific support for pollutant behavior modeling and risk assessment.
[0050] Example 6: Calculate the spatial correlation index through the following formula :
[0051] ;
[0052] where represents the Euclidean distance between the th sampling point and the th sampling point; represents the preprocessing data of the th sampling point.
[0053] Specifically, the calculation of the spatial correlation index is based on the standardized data of pollutant concentrations , which are obtained from the normalization of the historical concentrations at the sampling points in the previous module. The significance of standardization is to eliminate the comparison bias caused by the dimensional or concentration differences of different pollutants, enabling all pollutant data to be analyzed on the same scale during the calculation process. Meanwhile, the formula focuses on the degree to which the concentration at each sampling point deviates from the historical average concentration by subtracting the concentration value from the historical average concentration . This way of analyzing the deviation can reveal concentration outliers, such as regions above or below the average level, thereby quantifying the heterogeneity of pollutant distribution. When analyzing the spatial correlation of pollutant concentrations between sampling points, distance is a crucial factor. The formula takes into account the attenuation effect of distance on correlation by introducing the Euclidean distance between sampling points . Generally, the spatial correlation between two sampling points weakens as the geographical distance increases, which is a natural law based on the diffusion characteristics of pollutants. Therefore, when the distance between two sampling points is relatively close, the contribution of their concentration correlation to the overall spatial correlation index is relatively large; while when the distance is far, its impact on the overall index is significantly weakened. This distance-based weight design enables the calculation result to more realistically reflect the actual diffusion range and spatial connection of pollutants
[0054] In addition to the concentration and distance relationships between sampling points, the introduction of the sparse feature matrix further enhances the global interpretability of the spatial correlation index. The sparse feature matrix is the result of dimensionality reduction and feature representation of the preprocessed data by the previous feature extraction module, and it contains the key patterns and internal structures of the pollutant data in the target area. By calculating the Frobenius norm of the sparse feature matrix , the system can globally normalize the spatial correlation index, ensuring the comparability of the correlation indexes under different regions or different pollutant combinations. This normalization method also effectively balances the relationship between local concentration correlation and global feature characteristics, making the final spatial correlation index not only applicable to the pollution analysis of a single region but also capable of comparing results across regions or different pollutant types. During the calculation process, the covariance formula in the numerator part is the key to revealing the concentration correlation between sampling points. Covariance describes whether the concentration deviations of two sampling points show the same trend of change. For example, when and When the deviations of both are higher than the average concentration, their product is positive, indicating the consistency of the concentration changes at the two sampling points; conversely, when the concentration deviation at one point is higher than the average value while that at the other point is lower than the average value, the product is negative, indicating an inverse correlation in their concentration changes. By summing up the covariances of all pairs of sampling points, the overall correlation pattern of pollutants within the entire target area can be quantified. The finally calculated indicator value can directly reflect the spatial characteristics of the pollutant distribution within the target area. When the value is high, it indicates that the concentration distribution of pollutants within the area has strong spatial aggregation, and high-concentration points tend to be close to each other, and a similar trend is also shown among low-concentration points; when the value is low or even negative, it means that the pollutant distribution within the area has significant discreteness or disorder. This indicator is of great significance for identifying the diffusion centers, high-risk areas, and potential pollution source locations of pollutants. In addition, by combining the calculation results with geographical information, the system can generate an intuitive spatial correlation distribution map to provide visual support for pollution control decision-making.
[0055] Example 7: According to the following formula, calculate the environmental sensitivity of the target soil area based on the environmental parameters of the target soil area :
[0056] ;
[0057] where is the average organic matter content; is the average cation exchange capacity; is the average pH value; is the average soil bulk density; is the average soil permeability coefficient; is the average soil electrical conductivity; and are the average clay content and average silt content of the soil respectively; is the average value of the terrain slope; is the gamma function.
[0058] Specifically, the first part of the formula mainly describes the contribution of the chemical properties of the soil to the environmental sensitivity. The numerator of this part is where represents the organic matter content of the soil, which is a key indicator for measuring the pollutant adsorption capacity of the soil. Organic matter reduces the mobility of pollutants through physical or chemical interactions with pollutant molecules through its surface charge and pore structure. Therefore, a higher organic matter content usually means a lower environmental sensitivity. And The (cation exchange capacity) is a measure of the soil's ability to retain and exchange cations, which reflects the potential of the soil to immobilize heavy metal ions (such as lead, cadmium, etc.). Soils with high CEC values have a stronger ability to immobilize pollutants, thereby reducing their bioavailability and mobility. The formula captures the synergistic effect of organic matter and cation exchange capacity in enhancing the soil's pollution buffering performance by multiplying and . The denominator part plays a role in suppressing or amplifying the environmental sensitivity. The pH value of the soil is an important factor affecting the solubility and chemical form of pollutants. For example, under acidic conditions, many heavy metals (such as lead and cadmium) become more soluble and mobile, so low pH will significantly increase the environmental sensitivity. When the pH is slightly alkaline, pollutants tend to form insoluble compounds, thus reducing the sensitivity. The formula non-linearly amplifies the pH value in a squared form, indicating that extreme changes in pH (too high or too low) have a more significant impact on sensitivity. The permeability coefficient is a measure of the water migration ability in the soil, which has a direct impact on the vertical migration of pollutants. The formula non-linearly processes through logarithmic and square root functions, reflecting the important role of soils with higher permeability in pollutant diffusion. When the permeability coefficient is high, pollutants are more likely to enter the groundwater layer, significantly increasing the environmental sensitivity.
[0059] The second part of the formula focuses on the effects of soil particle composition, slope, and bulk density. The numerator part represents the clay content of the soil. Clay is the smallest particle size component in the soil, with a high specific surface area and adsorption capacity. A higher clay content usually means a stronger immobilization effect of the soil on pollutants, lower pollutant mobility, and thus reduced environmental sensitivity. The denominator part includes the combined effects of slope , bulk density and conductivity . The sine value of the slope is used to quantify the impact of terrain on the intensity of surface runoff. The steeper the slope, the stronger the surface runoff, thus increasing the risk of pollutant loss and enhancing the sensitivity. At the same time, the non-linear treatment of the slope smooths the sensitivity change through the sine function, enabling the model to better adapt to different terrain conditions. Soil bulk density is an important parameter for measuring the compactness of the soil, and its restrictive effect on the water infiltration path directly affects the migration behavior of pollutants. A higher bulk density usually reduces the diffusion range of pollutants, thus reducing the sensitivity. However, in the formula, the effect of bulk density is corrected by the gamma function , where (Conductivity) is an indicator for measuring the soil salt concentration. Soils with high salt content usually have low permeability, but the adsorption capacity of the soil for metal pollutants may be weakened due to the ion competition effect. The introduction of the gamma function can describe the non-linear regulation of the effect of salt on the bulk density, enabling the model to more accurately capture the combined effect of salt and bulk density. The third part of the formula highlights the soil silt content and slope jointly. When the silt content is relatively high, the soil structure is relatively loose, the particle stability is poor, and the adsorption capacity for pollutants is relatively low. Therefore, the environmental sensitivity will increase significantly. The exponential function amplifies the square value of the silt content, further highlighting the influence on the sensitivity when the silt content is high. The denominator part smoothes the sharp changes in the slope through the hyperbolic cosine function, avoiding sharp fluctuations in the sensitivity in the extreme slope regions. This design ensures the stability and continuity of the sensitivity calculation and at the same time enhances the adaptability of the model to complex terrain conditions.
[0060] Example 8: Through the following formula,, combined with enhanced Lasso regression modeling, the spatial correlation index, environmental sensitivity and sparse features are associated to construct a risk prediction model:
[0061] ;
[0062] wherein, is the data vector composed of the pollutant data obtained at the th sampling point at the current time; this risk prediction model is a minimum optimization problem. By solving the risk prediction model, the risk prediction value is obtained; is the determinant value of the calculation matrix; represents the modulus of the calculation vector.
[0063] Specifically, the first part of the formula describes the direct influence of the environmental sensitivity and the spatial correlation index on the risk prediction value . The core of this part lies in quantifying the relationship between the environmental parameters and the current pollutant state, and revealing the spatial distribution characteristics of the pollution risk in the target area in combination with the spatial correlation. In this part, the data vector represents the pollutant concentration at the th sampling point at the current time. The formula performs a modulus operation Calculate the amplitude characteristics of pollutants in multi-dimensional space. This amplitude calculation can comprehensively reflect the overall intensity of the current pollutant concentration, rather than the changes in a single dimension, thus capturing the global characteristics of the current pollution state more comprehensively. Environmental sensitivity is used as a regulatory factor and multiplied by to describe the response ability of the soil to the current state of pollutants. The introduction of environmental sensitivity takes into account the regulatory effects of the physical and chemical properties of the soil (such as pH value, organic matter content, and cation exchange capacity) on pollutant behavior. For example, in soils with high pH, the solubility and mobility of heavy metals are low, and the risk prediction value should be reduced accordingly; while in soils with low pH, the solubility of heavy metals increases, and the predicted risk value should increase significantly.
[0064] On the other hand, the formula incorporates the influence of the spatial correlation index into the risk prediction model through . The determinant operation is an important tool for measuring the global characteristics of a matrix, and its value can reflect the strength of the pollutant correlation between the sampling points described by . For example, when the index indicates that the pollutant concentration has strong aggregation (such as the pollutants in the high-concentration area are closely related to each other), the determinant value is large, and its influence on the risk prediction value also increases. This design reflects that high-aggregation areas are usually accompanied by higher environmental risks, because the local accumulation of high-concentration pollutants will cause greater pressure on the carrying capacity of the soil system and the surrounding ecological systems. By minimizing , the system can adjust the prediction value to match the current environmental characteristics and spatial distribution patterns, thus achieving accurate quantification of risks. The second part of the formula further introduces the influence of the sparse feature matrix on risk prediction. The sparse feature matrix is the result generated by the previous feature extraction module and contains the deep feature patterns of the pollutant data in the target area. By removing redundant information and mining key features, the sparse feature matrix can reflect the complex distribution relationships and dynamic evolution characteristics of pollutants in the target area. In this part, the formula captures the deviation degree between the environmental sensitivity and the risk prediction value through the quadratic difference term . When the soil environment shows high sensitivity to pollutants (i.e., the value is large), the system tends to increase the predicted risk value to reflect the low carrying capacity of the environment for pollutants. In the design of the formula, the quadratic operation further amplifies the deviation between and The impact of deviation on risk prediction is ensured, and the risk values in highly sensitive areas can more truly reflect the vulnerability of the soil system. The formula also adjusts the weight of sparse features in risk prediction by introducing a sparse feature matrix and the ratio with the current pollutant data. The Frobenius norm is a measure of the overall strength of the sparse feature matrix, representing the combined influence of important patterns contained in the sparse features. By comparing it with the Frobenius norm of the current pollutant data , the formula can dynamically adjust the contributions of sparse features and current data to risk prediction. For example, when the deep patterns in the sparse feature matrix are highly consistent with the current data, the influence of the sparse features on risk prediction will be amplified; while when the strength of the sparse features is low, the influence of the current pollutant data will dominate. In addition, the determinant terms and further correlate the overall characteristics of the sparse features with the spatial correlation index. This design captures the internal connections between different levels of features through multiple matrix operations, such as how sparse feature patterns interact with spatial distribution characteristics to affect risk prediction. Finally, by minimizing , the system can optimize the risk prediction value to make it closer to the actual pollutant distribution and environmental characteristics.
[0065] Example 9: The risk assessment module compares the risk prediction value with a preset threshold. If it exceeds the set threshold, it is determined that the target soil area has risks and a warning is issued.
[0066] Specifically, the risk prediction value is the core output of the previous model calculation, which synthesizes multi-dimensional information such as the current concentration data, spatial distribution characteristics, environmental sensitivity, and sparse features of pollutants in the target area. The magnitude of this value directly reflects the comprehensive pollution risk level of the target area: a higher value indicates that the concentration, diffusivity, or environmental sensitivity of pollutants in the area is at a high level, while a lower The value indicates a relatively low regional pollution risk. In this embodiment, the system classifies and determines the predicted risk level by setting a scientific and reasonable risk threshold. The setting of the threshold is the key to the entire early warning mechanism. The size of the threshold depends on the specific environmental characteristics of the target soil area, the historical pollution level, and the risk tolerance of regional management. For example, in areas with high environmental sensitivity (such as areas with high groundwater levels or soil areas with sparse vegetation cover), the threshold can be set lower to detect potential risks earlier. For areas with a high background value of pollutants but a strong ecosystem recovery ability, the threshold can be appropriately increased to avoid overly frequent false alarms. In addition, the system also supports setting different thresholds according to specific pollutant types. For example, for heavy metal pollution, a lower threshold may be used because of its slow mobility but significant long-term toxicity; for volatile organic compounds, the threshold may be higher to cope with the instantaneous fluctuations in their concentrations.
[0067] When the risk prediction value exceeds the preset threshold, the system will automatically determine that there is a potential pollution risk in the target area and trigger the early warning function. The early warning signal can be transmitted in various ways. For example, a risk alert is displayed on the system interface, the manager is notified by text message or email, or the operation of pollution control equipment is directly triggered. The early warning mechanism can not only provide qualitative risk prompts, but also further quantify the risk level according to the size of the value. For example, when is only slightly higher than the threshold, the system can determine it as "low risk" and recommend carrying out basic monitoring; while when is significantly higher than the threshold, the system will determine it as "high risk" or "extremely high risk" and recommend taking immediate emergency treatment measures. To improve the accuracy and applicability of the early warning, the system also supports dynamic adjustment of the threshold. The dynamic threshold adjustment mechanism is based on historical risk data, changes in environmental conditions, and the fluctuation trend of pollutant concentrations. For example, in the rainy season with frequent rainfall, the system can automatically lower the threshold to cope with the risk of pollutant leaching and diffusion caused by rainwater; while in the dry season, the threshold can be appropriately increased to reduce unnecessary false alarms. In addition, the dynamic threshold can also be adjusted in combination with the change rate of the risk prediction value. If the value grows rapidly in a short period of time, which usually means the spread or sharp increase of pollutants, the system can actively lower the threshold, improve the sensitivity of the early warning, and ensure timely detection of potential problems. Another important feature of the early warning mechanism is visual display. When the risk prediction value exceeds the threshold, the system will generate a detailed risk analysis report and present it in an intuitive form. For example, the system can draw a risk distribution heat map to show the spatial distribution of high-risk points in the target area; or generate a time series graph to display The changing trend of values over time. These visualization results can help managers quickly understand the sources, scope, and evolution laws of risks, thereby formulating more targeted governance plans.
[0068] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A big data analysis system for soil environmental damage and risk assessment, characterized in that, The system includes: a data acquisition module, a data processing and fusion module, a spatial correlation analysis module, and a risk assessment module; the data acquisition module is configured to set multiple sampling points in a target soil area, deploy the same sensor group at each sampling point, where the sensor group includes multiple types of sensors, and each type of sensor respectively acquires data of a type of pollutant at the sampling point, and sends the acquired data to the cloud big database; the data processing and fusion module is configured to obtain historical data acquired at a sampling time before the current sampling time from the cloud big database, perform standardization and fusion processing on the historical data to obtain preprocessed data for each sampling point, perform high-dimensional sparse feature extraction on the preprocessed data, and extract redundant information layer by layer by constructing a two-layer feature dictionary to extract the sparse features of the preprocessed data; the spatial correlation analysis module is configured to analyze the variation law of pollutant concentration over time and space based on the preprocessed data of each sampling point, and reveal the spatial diffusion characteristics of pollutants between sampling points by constructing a spatial weight matrix and a time variation model to obtain a spatial correlation index; the risk assessment module is configured to calculate the environmental sensitivity of the target soil area according to the environmental parameters of the target soil area, combine enhanced Lasso regression modeling, associate the spatial correlation index, environmental sensitivity with the sparse features, and construct a risk prediction model; and use the data vector composed of the data of the pollutants acquired at the sampling points at the current time as input, input it into the risk prediction model to obtain the current soil environmental risk value corresponding to each sampling point; according to the following formula, calculate the environmental sensitivity of the target soil area according to the environmental parameters of the target soil area : ; Among them, is the average organic matter content; is the average cation exchange capacity; is the average pH value; is the average soil bulk density; is the average soil permeability coefficient; is the average soil conductivity; and are the average clay content and average silt content of the soil, respectively; is the average value of the terrain slope; is the gamma function.
2. The big data analysis system for soil environmental damage and risk assessment according to claim 1, characterized in that, The area of the target soil area ranges from 50 to 500 square kilometers; in the target soil area, sampling points are set using a uniform grid, and the range of the grid side length is from 10 to 50 meters; the sensor group at each sampling point samples once a day to obtain pollutant data.
3. The big data analysis system for soil environmental damage and risk assessment according to claim 2, wherein The pollutant data includes: cadmium concentration, lead concentration, arsenic concentration, mercury concentration, copper concentration, zinc concentration, nickel concentration, chromium concentration, polycyclic aromatic hydrocarbon concentration, polychlorinated biphenyl concentration, nitrate concentration, sulfate concentration, chloride concentration, ammonium nitrogen concentration, and total petroleum hydrocarbon concentration; the environmental parameters are averages, including: average organic matter content, average cation exchange capacity, average pH value, average soil bulk density, average soil permeability coefficient, average soil electrical conductivity, average clay content, average silt content, and the average of the terrain slope of the target soil area.
4. The big data analysis system for soil environmental damage and risk assessment according to claim 3, characterized in that The historical data is standardized and fused through the following formula: ; wherein, is the preprocessed data of the th sampling point; is the number of pollutant types; is the original concentration of the th pollutant obtained at the previous sampling time before the current sampling time for the th sampling point; is the historical average concentration of the th pollutant; is the historical standard deviation of the th pollutant; is the historical maximum concentration of the th pollutant obtained at the th sampling point; is the historical minimum concentration of the th pollutant obtained at the th sampling point.
5. The soil environmental damage and risk assessment big data analysis system according to claim 4, wherein Through the following formula, by constructing a two-layer feature dictionary, redundant information is removed layer by layer, and the sparse features of the preprocessed data are extracted: ; By solving the minimum optimization problem represented by this formula, we obtain , and , and calculate the sparse features of the preprocessed data as follows: ; where is the preprocessing vector composed of the preprocessed data of all sampling points; where is the number of sampling points; is the first-layer feature dictionary; is the second-layer feature dictionary; is the sparse coefficient matrix, which is a preset diagonal matrix; represents the L2 norm; represents the L1 norm; is the Laplace operator; is the absolute value operator; is the F norm.
6. The big data analysis system for soil environmental damage and risk assessment according to claim 5, characterized in that Calculate the spatial correlation index using the following formula :[[]]END]] ; Among them, represents the Euclidean distance between the th sampling point and the th sampling point; represents the preprocessed data of the 7. The big data analysis system for soil environmental damage and risk assessment according to claim 6, characterized in that Through the following formula, combined with enhanced Lasso regression modeling, the spatial correlation index, environmental sensitivity, and sparse features are associated to construct a risk prediction model: ; Among them, is the data vector composed of the data of pollutants obtained at the th sampling point of the current time; the risk prediction model is a minimum optimization problem, and by solving the risk prediction model, the risk prediction value is obtained; is the determinant value of the calculation matrix; represents the modulus of the calculation vector.
8. The big data analysis system for soil environmental damage and risk assessment according to claim 7, wherein, The risk assessment module compares the risk prediction value with a preset threshold. If the threshold is exceeded, it is determined that the target soil area has risks and a warning is issued.
Citation Information
Patent Citations
Large-area soil heavy metal detection and spatial and temporal distribution characteristic analysis method and system
CN112986538A
Plain river network water engineering cluster multi-target scheduling rule making method and system
CN117035201A
High-rise building construction environment monitoring system based on distributed edge calculation
CN118966744A