Soil remediation process intelligent control system based on artificial intelligence
Through the AI-based intelligent control system for soil remediation, the problems of high cost and poor timeliness of traditional soil remediation monitoring have been solved, accurate pollutant concentration prediction and intelligent decision-making have been achieved, and the soil remediation process has been optimized.
Patent Information
- Application Number
- CN202511128298.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional soil remediation monitoring relies on manual sampling and analysis, which is costly and inefficient, making it difficult to achieve accurate and real-time monitoring of pollutant concentrations and intelligent decision-making.
An AI-based intelligent control system for the soil remediation process is adopted. Through data acquisition, preprocessing, state encapsulation and pollutant concentration prediction modules, an alignment feature matrix is constructed. Sequence codecs are used for soft measurement prediction of pollutant concentrations, and active sampling decisions are generated to optimize sampling resource deployment.
It has reduced monitoring costs, improved timeliness, achieved closed-loop control from data acquisition to decision-making, and improved the intelligence level and efficiency of the soil remediation process.
Smart Images

Figure CN120631072A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent control, and more specifically, to an intelligent control system for soil remediation process based on artificial intelligence. Background Art
[0002] Soil is a critical natural resource for maintaining ecosystem balance and human survival and development. Therefore, effective soil remediation of contaminated sites is crucial. During the soil remediation process, continuous monitoring of soil pollutant concentrations is necessary to monitor remediation results in real time and dynamically adjust remediation strategies.
[0003] Traditional soil remediation monitoring relies primarily on manual on-site sampling and laboratory chemical analysis. While this method can provide accurate pollutant concentration data, its drawbacks are also significant: First, the sampling and analysis process is time-consuming, labor-intensive, and costly, limiting monitoring frequency and preventing timely reflection of dynamic changes in soil conditions. Second, the selection of sampling points often relies on experience and is highly subjective, potentially leading to omissions of contaminated hotspots or over-monitoring of remediated areas, resulting in a waste of resources. Finally, the remediation process involves complex physical, chemical, and biological reactions, and changes in pollutant concentrations are affected by the coupling of multiple factors such as temperature, humidity, and pH. Therefore, it is difficult to fully and deeply understand the remediation mechanism and predict future trends based solely on discrete concentration data. Therefore, how to reduce monitoring costs, improve monitoring timeliness, and achieve intelligent decision-making in the remediation process while ensuring monitoring accuracy has become a pressing technical issue.
[0004] Therefore, an optimized intelligent control system for soil remediation process is desired. Summary of the Invention
[0005] In order to solve the above technical problems, the present application is proposed. The embodiments of the present application provide an intelligent control system for soil remediation process based on artificial intelligence.
[0006] According to one aspect of the present application, there is provided an artificial intelligence-based intelligent control system for a soil remediation process, which includes: Data acquisition module, used to obtain raw sensor data and control logs; Data preprocessing module, used to preprocess the raw sensor data and control logs to obtain the alignment feature matrix; A state encapsulation module is used to perform feature vector serialization and state encapsulation on the alignment feature matrix to obtain a state sequence at a current moment; A pollutant concentration prediction module is used to perform soft measurement prediction of pollutant concentration based on a sequence codec on the state sequence at the current moment to obtain a prediction result, wherein the prediction result includes a predicted pollutant concentration value and its confidence level; An active sampling decision module is used to make an active sampling decision based on uncertainty based on the prediction result to obtain a sampling plan.
[0007] In the above-mentioned AI-based intelligent control system for the soil remediation process, the raw sensor data include sensor ID, timestamp, soil temperature, soil moisture, soil pH, redox potential, electrical conductivity, dissolved oxygen concentration, photoionization detector reading, flame ionization detector reading, and ion selective electrode reading.
[0008] In the above-mentioned artificial intelligence-based intelligent control system for soil remediation process, the data preprocessing module is used to: normalize and clean the original sensor data and control log to obtain a cleaned sensor data table and a cleaned control log table; perform data aggregation and alignment based on a unified time grid on the cleaned sensor data table and the cleaned control log table to obtain unified time grid data, wherein each row of the unified time grid data represents a time window, and each column of the unified time grid data represents the aggregated statistical features of all sensors and control quantities; perform multidimensional feature derivation on the unified time grid data to obtain the aligned feature matrix, wherein each row of the aligned feature matrix represents a time step, and each column of the aligned feature matrix represents the aggregated statistical features and multidimensional derived features of all sensors and control quantities.
[0009] In the above-mentioned artificial intelligence-based intelligent control system for the soil remediation process, the state encapsulation module is used to: based on the current timestamp, extract the feature data frame based on the preset time window of the alignment feature matrix to obtain a time data frame; perform eigenvalue normalization processing on the time data frame to obtain a normalized data frame; and perform feature serialization on the normalized data frame to obtain the state sequence at the current moment.
[0010] In the above-mentioned artificial intelligence-based intelligent control system for soil remediation process, the pollutant concentration prediction module includes: a soil state sequence encoding unit, which is used to input the state sequence at the current moment into a sequence encoder to obtain a time series of soil state feature implicit coding vectors; a soil state information transmission unit, which is used to input the time series of soil state feature implicit coding vectors into a soil state information transmission layer to obtain a soil state feature time series context coding vector; and a soil state time series prediction unit, which is used to input the soil state feature time series context coding vector into a sequence decoder to obtain the prediction result.
[0011] In the above-mentioned intelligent control system of soil remediation process based on artificial intelligence, the soil state information transmission unit includes: a soil state contrast quantization subunit, which is used to quantify the soil state uncertainty of each soil state feature implicit coding vector in the time series of the soil state feature implicit coding vector to obtain the sequence distribution of the soil state feature time series entropy gain; a soil state time series weight calculation subunit, which is used to determine the sequence distribution of the soil state feature time series modulation weight based on the sequence distribution of the soil state feature time series entropy gain; a soil state time series feature modulation subunit, which is used to perform weighted modulation on the sequence distribution of the soil state feature implicit coding vector based on the sequence distribution of the soil state feature time series modulation weight to obtain the sequence distribution of the modulated soil state feature vector; a soil state time series context encoding subunit, which is used to input the sequence distribution of the modulated soil state feature vector into a forward LSTM-based sequence encoder to obtain the soil state feature time series context encoding vector.
[0012] In the above-mentioned artificial intelligence-based intelligent control system for soil remediation process, the sequence encoder is a trained LSTM model, and the sequence decoder is a fully connected layer.
[0013] In the above-mentioned AI-based intelligent control system for the soil remediation process, the active sampling decision module is used to: convert the confidence level in the prediction results of each grid point into basic uncertainty to obtain an uncertainty map; based on a minimum spatial distance threshold, generate a candidate sampling point pool based on the uncertainty score for the uncertainty map to obtain a ranked candidate pool; and select the top N candidate points from the ranked candidate pool to obtain the sampling plan, where N is the sampling budget number.
[0014] Compared to existing technologies, this application provides an AI-based intelligent control system for soil remediation processes. This system performs deep preprocessing and feature engineering on multi-source, heterogeneous sensor data and control logs to construct an aligned feature matrix that comprehensively characterizes the dynamic process of soil remediation. Subsequently, to address the difficulty of obtaining pollutant concentrations in real time, a soft sensing model based on a sequence codec is employed to map the processed state sequence into pollutant concentrations and their confidence levels. Furthermore, a soil state information transfer layer is introduced into the soft sensing prediction process for pollutant concentrations. This layer dynamically adjusts information flow weights by quantifying and leveraging the uncertainty of time series information, effectively suppressing noise interference and capturing complex time series dependencies, ensuring the accuracy and robustness of predictions. Finally, rather than blindly accepting prediction results, the system utilizes the confidence information in the prediction results, converting them into uncertainty maps and generating an optimized active sampling plan based on these maps. This allows limited sampling resources to be precisely deployed to areas most in need of verification and with the highest uncertainty, reducing monitoring costs while achieving closed-loop control from data acquisition, intelligent prediction, to optimized decision-making. This in turn improves the intelligence level and efficiency of the soil remediation process. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0016] Figure 1 This is a system block diagram of an artificial intelligence-based intelligent control system for soil remediation process according to an embodiment of the present application.
[0017] Figure 2 Schematic diagram of data flow of an intelligent control system for soil remediation process based on artificial intelligence according to an embodiment of the present application.
[0018] Figure 3 This is a block diagram of a pollutant concentration prediction module in an artificial intelligence-based soil remediation process intelligent control system according to an embodiment of the present application.
[0019] Figure 4 4 is a block diagram of a soil state information transmission unit in an artificial intelligence-based soil remediation process intelligent control system according to an embodiment of the present application. DETAILED DESCRIPTION
[0020] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.
[0021] As used in this application and the claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.
[0022] Although the present application makes various references to certain modules in the system according to embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are illustrative only, and different aspects of the system and method can use different modules.
[0023] Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the various steps may be processed in reverse order or simultaneously, as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0024] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.
[0025] To address the technical challenges in traditional soil remediation monitoring, as outlined in the background technology, namely, how to overcome the limitations of high costs and poor timeliness of manual sampling and analysis, and leverage readily available process data to accurately predict pollutant concentrations and intelligently inform remediation decisions, the technical solution of this application proposes an AI-based intelligent control system for the soil remediation process. First, the system comprehensively collects and integrates raw data streams and equipment control logs generated by various sensors deployed at the remediation site (such as temperature, humidity, pH, PID, etc.). Through a series of refined preprocessing operations, including standardized cleaning, unified time grid-based aggregation alignment, and multidimensional feature derivation, this multi-source, heterogeneous information is transformed into a structurally unified, information-rich aligned feature matrix, laying a solid data foundation for subsequent in-depth analysis.
[0026] Next, to accurately predict key pollutant concentrations, the system serializes the current feature matrix into a state sequence and inputs it into an artificial intelligence model based on a sequence codec. The core of this model lies in its unique soil state information transmission layer, which quantifies the uncertainty of information at each time node by calculating the temporal entropy gain during the encoding process, and dynamically adjusts the weight of information transmission along the time series based on this. This mechanism enables the model to intelligently focus on high-value information and suppress noisy data, thereby capturing the deep, nonlinear temporal dependencies between pollutant concentrations and numerous environmental variables, ultimately outputting highly accurate pollutant concentration predictions and their reliability assessments.
[0027] Finally, this solution does not stop at prediction, but transforms the prediction results into intelligent decisions to guide practice. The system uses the confidence information in the prediction results to generate a spatial uncertainty map covering the entire remediation site. Based on this map, the system automatically plans the optimal sampling point plan through an active sampling decision algorithm, taking into account the uncertainty score and spatial distribution. This solution can accurately guide limited and expensive physical sampling resources to the key areas where the model is most uncertain and most in need of data verification, thereby obtaining the highest value feedback information at the lowest cost, realizing dynamic optimization and closed-loop intelligent control of the remediation process, and ultimately achieving the goal of efficient, economical and reliable soil remediation.
[0028] Figure 1 This is a system block diagram of an artificial intelligence-based intelligent control system for soil remediation process according to an embodiment of the present application. Figure 2 Schematic diagram of data flow of the intelligent control system of soil remediation process based on artificial intelligence according to the embodiment of the present application. Figure 1 and Figure 2 As shown, according to an embodiment of the present application, an artificial intelligence-based intelligent control system 100 for soil remediation process includes: a data acquisition module 110 for acquiring raw sensor data and control logs; a data preprocessing module 120 for preprocessing the raw sensor data and control logs to obtain an alignment feature matrix; a state encapsulation module 130 for performing feature vector serialization and state encapsulation on the alignment feature matrix to obtain a state sequence at the current moment; a pollutant concentration prediction module 140 for performing soft measurement prediction of pollutant concentration based on a sequence codec on the state sequence at the current moment to obtain a prediction result, wherein the prediction result includes a predicted pollutant concentration value and its confidence level; and an active sampling decision module 150 for performing uncertainty-based active sampling decision based on the prediction result to obtain a sampling plan.
[0029] In the aforementioned AI-based intelligent control system 100 for soil remediation, the data acquisition module 110 and data preprocessing module 120 are used to acquire raw sensor data and control logs, and preprocess these data to obtain an aligned feature matrix. It is worth noting that the raw sensor data here includes sensor ID, timestamp, soil temperature, soil moisture, soil pH, redox potential, conductivity, dissolved oxygen concentration, photoionization detector readings, flame ionization detector readings, and ion-selective electrode readings. It should be understood that in actual soil remediation scenarios, raw sensor data and control logs collected directly from the field come from diverse sources and formats, and inevitably suffer from issues such as missing data, abnormal noise, misaligned timestamps, and inconsistent sampling frequencies. The quality of these raw data varies greatly and cannot be directly used as input for AI models. If left unprocessed, they will severely impact the training and prediction accuracy of subsequent pollutant concentration prediction models, and may even cause the models to fail to converge. Therefore, in the technical solution of this application, the raw sensor data and control logs are further preprocessed to transform the original, chaotic data into structured, high-quality, and information-rich feature data, providing a standardized and reliable data foundation for subsequent time series modeling. This can improve the stability and performance of the model and ensure the accuracy of the decision-making of the entire intelligent control system.
[0030] Specifically, in a specific example of this application, the process of acquiring raw sensor data and control logs and preprocessing them to obtain an aligned feature matrix is as follows: First, normalization and cleaning are performed. The system obtains raw data from a database or data interface, including sensor data such as soil temperature, moisture, pH, redox potential, conductivity, dissolved oxygen concentration, photoionization detector readings, flame ionization detector readings, and ion-selective electrode readings from various sensors, as well as control logs recording the operating status of remediation equipment (such as injection pumps and ventilators).
[0031] Then, the raw sensor data and control log are preprocessed. In an embodiment of the present application, the data preprocessing module is used to: normalize and clean the raw sensor data and control log to obtain a cleaned sensor data table and a cleaned control log table; perform data aggregation and alignment based on a unified time grid on the cleaned sensor data table and the cleaned control log table to obtain unified time grid data, wherein each row of the unified time grid data represents a time window, and each column of the unified time grid data represents the aggregated statistical features of all sensors and control quantities; perform multidimensional feature derivation on the unified time grid data to obtain the aligned feature matrix, wherein each row of the aligned feature matrix represents a time step, and each column of the aligned feature matrix represents the aggregated statistical features and multidimensional derived features of all sensors and control quantities.
[0032] More specifically, the system normalizes these raw sensor data and control logs, unifying the data format (e.g., timestamp format and data units), and performs data cleansing. By setting thresholds or using interpolation algorithms (such as linear or spline interpolation) to address obvious outliers and missing values, the system generates cleaned sensor data tables and cleaned control log tables with uniform formats and complete data. Next, the system performs data aggregation and alignment based on a unified time grid. Because the timestamps of the cleaned data sources are still discrete and asynchronous, the system sets a unified time grid, such as a 5-minute or 10-minute time window. The system then aggregates all sensor readings and control variable data falling within each time window. Aggregation involves calculating statistical features of the data within that time window, such as the mean, maximum, minimum, and variance. This operation aligns data of varying frequencies to a unified time base, generating unified time-grid data. In this data table, each row represents a fixed time window, and each column represents the aggregated statistical features of a sensor or control variable, addressing the heterogeneity of the data along the temporal dimension. Finally, multidimensional feature derivation is performed. In order to more deeply explore the implicit correlation information in the data, the system performs multi-dimensional feature derivation based on the unified time grid data. This includes calculating the time difference of the features to capture their rate of change; calculating the cross-combination between features to characterize the interaction between different physical and chemical parameters; and calculating the moving average of the features at different time scales to reflect their short-term and long-term trends. These newly derived features and the original aggregated statistical features together constitute the final aligned feature matrix. Each row of this matrix represents the time step required for a model, and each column contains all the original and derived features related to the soil remediation status, which greatly enriches the input information dimension and provides sufficient feature support for subsequent models to accurately capture the complex dynamic changes in soil status.
[0033] In the aforementioned AI-based intelligent control system 100 for soil remediation, the state encapsulation module 130 is used to perform feature vector serialization and state encapsulation on the aligned feature matrix to obtain the state sequence at the current moment. It should be understood that, although the aligned feature matrix obtained after preprocessing has a regular structure, it is essentially still a large static data table. The subsequent sequence model (such as LSTM) used to predict pollutant concentrations requires input data to be a time series with a specific time length and fixed format. Furthermore, the dimensions and numerical ranges of different features in the aligned feature matrix vary significantly. For example, temperature values may range from 0-40, while conductivity may range from 0-2000. This difference can severely impact the convergence speed and final performance of model training, causing model weights to favor features with larger values. Therefore, in the technical solution of the present application, the aligned feature matrix is further subjected to feature vector serialization and state encapsulation to transform the static, global feature matrix into a standardized, dynamic state sequence containing historical information that meets the input requirements of the time series model. In this way, the subsequent sequence encoder-decoder model can be provided with input data with correct format and stable values, thus ensuring that the model can efficiently and accurately learn the temporal evolution law of soil state.
[0034] Specifically, in an embodiment of the present application, the state encapsulation module is used to: based on the current timestamp, extract the feature data frame based on the preset time window of the alignment feature matrix to obtain a time data frame; perform eigenvalue normalization on the time data frame to obtain a normalized data frame; perform feature serialization on the normalized data frame to obtain the state sequence at the current moment.
[0035] Specifically, the process for serializing the feature vectors and encapsulating the state of the aligned feature matrix to obtain the current state sequence is as follows: First, the system extracts a continuous segment of data from the complete aligned feature matrix based on the timestamp of the current prediction and a preset time window length (e.g., looking back 24 time steps). This segment of data constitutes a time data frame. This data frame contains a series of feature vectors from a certain moment in the past to the current moment, fully recording the historical evolution of the recent soil state and providing the necessary time dimension data for the model to understand the context of the current state. Second, the system normalizes all eigenvalues within the extracted time data frame. In a specific example of this application, min-max normalization is used to linearly scale all values of each feature to a fixed interval, such as [0, 1] or [-1, 1]. This operation eliminates the impact of different dimensions and value ranges between different features, ensuring that each feature has an equal contribution to model training, thereby accelerating model convergence and improving its stability, resulting in a normalized data frame. Finally, the system encapsulates the normalized data frame, which is already a two-dimensional array (time step × feature dimension), into a sequence object. This sequence object represents the state sequence at the current moment. It contains a specific history length, all features are normalized, and can be directly fed into the sequence encoder / decoder model for processing. This step completes the final conversion from static data to dynamic time series input, paving the way for subsequent soft-sensing predictions of pollutant concentrations.
[0036] In the aforementioned AI-based intelligent control system 100 for soil remediation, the pollutant concentration prediction module 140 is used to perform soft-sensing prediction of pollutant concentration based on a sequence codec on the current state sequence to obtain a prediction result, which includes the predicted pollutant concentration value and its confidence level. It should be understood that since the change in pollutant concentration in soil is a complex dynamic process, it is affected by the nonlinear and temporal coupling of multiple environmental factors such as temperature, humidity, and pH. Using traditional linear or static models makes it difficult to accurately capture this complex temporal dependency, resulting in insufficient prediction accuracy. At the same time, the reliability of the prediction result is crucial for subsequent remediation decisions. A single prediction value without confidence cannot provide a basis for risk assessment and proactive sampling. Therefore, in the technical solution of the present application, soft-sensing prediction of pollutant concentration based on a sequence codec is further performed on the current state sequence to deeply explore the temporal correlation between multidimensional sensor data and pollutant concentration and quantify the uncertainty of the prediction result. In this way, it is possible to provide not only an accurate pollutant concentration prediction value but also its confidence level, providing core decision-making information for more intelligent and reliable remediation process control.
[0037] Figure 3 FIG is a block diagram of a state encapsulation module in an intelligent control system for soil remediation process based on artificial intelligence according to an embodiment of the present application. Figure 3 As shown, in an embodiment of the present application, the pollutant concentration prediction module 140 includes: a soil state sequence encoding unit 141, which is used to input the state sequence at the current moment into a sequence encoder to obtain a time series of soil state feature implicit coding vectors; a soil state information transmission unit 142, which is used to input the time series of soil state feature implicit coding vectors into a soil state information transmission layer to obtain a soil state feature time series context coding vector; a soil state time series prediction unit 143, which is used to input the soil state feature time series context coding vector into a sequence decoder to obtain the prediction result.
[0038] Specifically, the soil state sequence encoding unit 141 is used to input the state sequence at the current moment into a sequence encoder to obtain a time series of soil state feature implicit encoding vectors. It should be understood that since the input state sequence at the current moment is a high-dimensional feature set containing multiple time steps, it directly reflects the changes in physical and chemical parameters of the soil over a period of time, but there are complex, nonlinear temporal dependencies between these original features. If these original sequences are directly used for prediction without effective feature extraction, it is difficult to capture the key dynamic patterns that determine the changes in pollutant concentrations. In addition, the original sequence data has high dimensions and redundancy, and direct processing will increase the complexity of subsequent calculations. Therefore, in the technical solution of the present application, the state sequence at the current moment is further input into a sequence encoder to perform deep feature extraction and information compression on the original time series data, converting the high-dimensional, explicit state sequence into a low-dimensional, abstract representation containing core time series dynamic information. In this way, the dynamic features most relevant to the changes in pollutant concentrations can be effectively extracted from the complex sensor data stream, laying the foundation for subsequent more accurate analysis and prediction.
[0039] More specifically, in one example of this application, the process for inputting the current state sequence into a sequence encoder to obtain a time series of implicitly encoded soil state feature vectors is as follows: First, the normalized current state sequence generated in the previous stage is used as input. This sequence is structurally a two-dimensional tensor with dimensions [sequence length × number of features], where the sequence length is the preset backtracking time step and the number of features is the total number of features after preprocessing and derivation. Next, a pretrained long short-term memory (LSTM) model is loaded as the sequence encoder. This LSTM model has been fully trained using historical data (including sensor data sequences and corresponding true pollutant concentration labels), and its internal weight parameters have learned the mapping rules from the soil environmental parameter sequence to its internal state changes. Finally, the encoding process is performed. The prepared state sequence is fed into the LSTM model. The model processes each feature vector in the sequence in order of time steps. At each time step, the LSTM unit uses its internal input, forget, and output gate structures to determine which new information to remember and which old information to forget, and updates its internal cell state and hidden state. The hidden state output at each time step is the implicit encoding vector of the soil state characteristics at that moment. When the entire input sequence is processed, the hidden states generated by all time steps form a new time series: the time series of the implicit encoding vectors of the soil state characteristics. This sequence has a dimension of [sequence length × hidden layer dimension] and represents the entire temporal dynamic information in the original input sequence in a more concise and abstract manner.
[0040] Specifically, the soil state information transmission unit 142 is configured to input the time series of soil state feature implicit encoding vectors into the soil state information transmission layer to obtain a soil state feature temporal context encoding vector. It should be understood that, although the time series of soil state feature implicit encoding vectors output by the sequence encoder has compressed the raw sensor data, it may still contain a large amount of redundant or noisy information and treat each time step equally. In real soil remediation scenarios, not all soil state changes are equally important; certain critical moments, such as the intense reaction period after remediation agent injection, contain state changes that are far more valuable than those during stable periods. Traditional sequence processing approaches passively aggregate all information and struggle to actively identify these key events that determine the success or failure of remediation, resulting in the model's inaccurate capture of the system's dynamic evolutionary logic. Therefore, in the technical solution of this application, the time series of soil state feature implicit encoding vectors is further input into the soil state information transmission layer. This shifts the temporal information processing paradigm from passive aggregation to active identification and guided encoding, assigning each time step an endogenous, dynamic importance score, and thus nonlinearly reshaping the entire temporal information flow. In this way, subsequent models can be freed from complex raw data and focus on in-depth modeling of the temporal dependencies between key events that have been identified as high-value, thereby achieving more accurate and efficient capture of the dynamic evolution logic of the soil remediation process.
[0041] Figure 4 FIG is a block diagram of a soil state information transmission unit in an intelligent control system for soil remediation process based on artificial intelligence according to an embodiment of the present application. Figure 4 As shown, in an embodiment of the present application, the soil state information transmission unit 142 includes: a soil state contrast quantization subunit 1421, which is used to quantify the soil state uncertainty of each soil state feature implicit coding vector in the time series of the soil state feature implicit coding vector to obtain the sequence distribution of the soil state feature time series entropy gain; a soil state time series weight calculation subunit 1422, which is used to determine the sequence distribution of the soil state feature time series modulation weight based on the sequence distribution of the soil state feature time series entropy gain; a soil state time series feature modulation subunit 1423, which is used to perform weighted modulation on the sequence distribution of the soil state feature implicit coding vector based on the sequence distribution of the soil state feature time series modulation weight to obtain the sequence distribution of the modulated soil state feature vector; a soil state time series context encoding subunit 1424, which is used to input the sequence distribution of the modulated soil state feature vector into a forward LSTM-based sequence encoder to obtain the soil state feature time series context encoding vector.
[0042] Accordingly, the soil state contrast quantification subunit 1421 is configured to quantify soil state uncertainty of each soil state feature implicit coding vector in the time series of the soil state feature implicit coding vector to obtain a sequence distribution of soil state feature temporal entropy gain, which is expressed by the following formula:
[0043]
[0044]
[0045] in, is the first in the time series of the soil state feature implicit coding vector The soil state feature implicit coding vector at the time node, is the one-norm of the vector, is the density matrix, is the eigenvalue of each position in the density matrix, is the number of vectors in the time series of the soil state feature implicit coding vector, for Soil state characteristic information, for Soil state characteristic information, is the sequence distribution of soil state characteristic temporal entropy gain and The soil state characteristic time series entropy gain between for activation function, is an exponential function with the natural constant e as the base, and are the mean and standard deviation of the temporal entropy gain, is the first in the sequence distribution of soil state characteristic time series entropy gain The time series entropy gain of soil state characteristics.
[0046] It should be understood that, in the process of soil remediation, the state information at different time points is not equally important for predicting the future. Although an implicit coding vector sequence contains time series information, it does not explicitly distinguish between "things that mark key changes in the system" and "things that are just regular, stable evolutions." If processed without distinction, the model may be disturbed by a large amount of low-value information generated by stable periods, making it difficult to focus on those mutation moments that truly drive changes in pollutant concentrations, such as the most intense chemical reaction stage after the injection of the remediation agent. Therefore, in the technical solution of the present application, the soil state uncertainty of each soil state feature implicit coding vector in the time series of the soil state feature implicit coding vector is further quantified, so as to draw on the concept of entropy in information theory, abstract the characteristic distribution form of the soil state at each moment into a measurable state certainty, and calculate the time series entropy gain between adjacent time steps. In this way, it is possible to go beyond the direct dependence on the eigenvalue itself and focus on the intrinsic mutation of the feature evolution pattern, introducing a high-order meta-information into the system, namely the unexpectedness or information content of the event, thereby providing a clear and physically meaningful basis for the subsequent dynamic weight allocation.
[0047] Accordingly, the soil state time series weight calculation subunit 1422 is used to determine the sequence distribution of soil state feature time series modulation weights based on the sequence distribution of soil state feature time series entropy gain, which is expressed by the following formula:
[0048] in, and are the learnable scaling parameters, is the normalized activation term, is the historical entropy decay term, is the sequence distribution of soil state characteristic temporal modulation weights The corresponding soil state characteristic time series modulation weight.
[0049] It should be understood that simply calculating the entropy gain of the soil state characteristics over time only yields a sequence of indicators representing the unexpectedness or information content of the soil state at each moment. This sequence itself cannot be directly used to adjust the information flow. For the model to proactively focus on critical moments with high entropy gain in the soil state, this unexpectedness must be transformed into actual control over the information flow. Traditional processing methods, such as fixed time decay models, are unable to allocate attention based on the dynamic evolution of the data itself, lacking flexibility and specificity. Therefore, in the technical solution of this application, the sequence distribution of the soil state characteristics over time modulation weights is further determined based on the sequence distribution of the entropy gain of the soil state characteristics over time. This is used to construct a dynamic, content-dependent weight generation mechanism, mapping the unexpectedness indicator, entropy gain, into influence weights through a nonlinear function. This achieves an endogenous attention mechanism, allowing the model to autonomously and nonlinearly amplify the influence of those moments that mark key phase transitions in soil remediation (such as chemical reaction peaks), while suppressing noise interference during regular, stable evolution periods. This ensures that the importance of the time dimension is no longer predetermined but determined by the dynamic evolution of the data itself.
[0050] Accordingly, the soil state time series feature modulation subunit 1423 is configured to perform weighted modulation on the sequence distribution of the soil state feature implicit coding vector based on the sequence distribution of the soil state feature time series modulation weights to obtain the sequence distribution of the modulated soil state feature vector, which is expressed by the following formula:
[0051] in, is the sequence distribution of the modulated soil state feature vector, are the first, second, and third order in the sequence distribution of the modulated soil state feature vector. and The sequence distribution of the modulated soil state feature vectors.
[0052] It should be understood that although the timing modulation weight sequence generated in the previous step has quantified the importance of each time point, it is still an independent control signal and has not yet had an effect on the actual soil state information flow. The original soil state feature implicit coding vector sequence is still an unscreened data stream with uniform information density, in which the influence of key events is mixed with the influence of noise in the stable period, which will reduce the efficiency and accuracy of subsequent encoder learning. Therefore, in the technical solution of the present application, the sequence distribution of the soil state feature implicit coding vector is further weighted modulated based on the sequence distribution of the soil state feature timing modulation weights, so as to implement a kind of information gating, and directly apply the importance weights calculated in the previous step to the original feature data stream, substantially reshaping the time series input to the downstream encoder. In this way, the modulated sequence is no longer a uniform time record, but a sequence that has been preprocessed by importance and has dynamically changing information density. The soil state characteristics that are judged to be critical and high-entropy gain moments (such as the period of intense remediation reaction) will be completely retained or even amplified, while the influence of the soil state characteristics at those unremarkable moments will be weakened, thereby highlighting the historical turning points that are most valuable for predicting future pollutant concentrations.
[0053] Accordingly, the soil state temporal context encoding subunit 1424 is configured to input the sequence distribution of the modulated soil state feature vector into a forward LSTM-based sequence encoder to obtain the soil state feature temporal context encoding vector, which is expressed by the following formula:
[0054] in, is a forward LSTM-based sequence encoder, is the soil state feature temporal context encoding vector.
[0055] It should be understood that although the soil state feature vector sequence after weighted modulation has highlighted the information of key time nodes, it is essentially still a reshaped time series, in which the temporal correlation and evolutionary logic between each key event have not been systematically modeled and refined. To generate a final code that can comprehensively summarize the dynamic evolutionary nature of the entire remediation process, a mechanism that can effectively capture long-range temporal dependencies is needed to integrate this key-marked information sequence. Therefore, in the technical solution of the present application, the sequence distribution of the modulated soil state feature vector is further input into a sequence encoder based on the forward LSTM, so as to utilize the powerful ability of LSTM to capture long-range temporal dependencies and perform the final deep encoding on this sequence that has been pre-processed by the upstream module and whose information density changes dynamically. In this way, the function of the sequence encoder can be further focused and strengthened, so that it no longer needs to learn patterns from the original sequence that may be full of redundant soil state information, but directly operates on the historical event sequence that has been proven to be crucial, thereby generating a highly condensed soil state feature temporal context encoding vector that deeply reflects the dynamic evolutionary nature of soil remediation.
[0056] Specifically, the soil state time series prediction unit 143 is used to input the soil state feature time series context encoding vector into the sequence decoder to obtain the prediction result. It should be understood that since the soil state feature time series context encoding vector generated in the previous step is a highly condensed and abstract feature representation, although it contains the core time series dynamic information that is crucial to the change of pollutant concentration, it is not a directly readable prediction value in itself. This vector exists in a high-dimensional feature space and requires an effective mapping mechanism to convert it from this abstract representation space to a specific, physically meaningful target prediction space, that is, the pollutant concentration value and its uncertainty measurement. Therefore, in the technical solution of the present application, the soil state feature time series context encoding vector is further input into the sequence decoder to finally parse and convert this deep feature that contains the evolution logic of key historical events, and decode it into a specific prediction of the pollutant concentration at the current moment. In this way, the mapping from abstract features to the final prediction result can be completed, and the model's deep understanding of the soil remediation process can be converted into a quantitative decision-making basis that can be used to guide actual operations.
[0057] More specifically, in one example of this application, the process for inputting the soil state feature temporal context encoding vector into a sequence decoder to obtain the prediction result is as follows: First, the system uses the soil state feature temporal context encoding vector output by the soil state information transmission layer as the sole input to the sequence decoder. This input is a fixed-dimensional vector that condenses the dynamic information most relevant to the prediction target across the entire historical state sequence. Second, the soil state feature temporal context encoding vector is fed into one or more fully connected layers, which serve as the sequence decoder. The fully connected layer performs a linear transformation on the input encoding vector using its internal weight matrix and bias vector, and then processes it through a nonlinear activation function (such as a ReLU function). This process nonlinearly maps the input vector from its high-dimensional feature space to a new output space with a dimensionality matching the prediction target. Finally, the final layer of the decoder is designed with two independent output heads. The first output head is a neuron, whose output value directly corresponds to the predicted pollutant concentration value. The second output head is also composed of one or more neurons, whose output is used to calculate the confidence of the prediction result. A specific implementation involves the second output head predicting a variance value, which represents the degree of uncertainty in the prediction. The inverse of this variance value serves as the confidence level. These two outputs together constitute the final prediction result, providing complete and necessary information for subsequent active sampling decisions based on uncertainty.
[0058] In the aforementioned AI-based intelligent control system 100 for soil remediation, the active sampling decision module 150 is configured to make active sampling decisions based on uncertainty based on the prediction results to determine a sampling plan. It should be understood that comprehensive, high-frequency manual sampling at the soil remediation site is costly and impractical, while traditional fixed grid or random sampling strategies lack specificity, often resulting in insufficient sampling in key areas and excessive sampling in stable areas, making it impossible to efficiently obtain the most valuable data for calibrating and optimizing the prediction model. Model prediction results inherently carry uncertainty, especially in areas where pollutant concentrations fluctuate dramatically or where model knowledge is insufficient. This uncertainty itself is key information for guiding sampling and improving monitoring efficiency. Therefore, in the technical solution of the present application, active sampling decisions based on uncertainty are further made based on the prediction results. This uses the uncertainty of the model prediction as a core driving force to intelligently plan the next sampling location, precisely allocating limited sampling resources to areas where the model is most confused and where information gain is greatest. In this way, key information for model iteration and process validation can be maximized with minimal sampling cost, thereby achieving closed-loop optimization of remediation monitoring and significantly improving the timeliness and cost-effectiveness of monitoring.
[0059] Specifically, in an embodiment of the present application, the active sampling decision module is used to: convert the confidence in the prediction result of each grid point into basic uncertainty to obtain an uncertainty map; based on a minimum spatial distance threshold, generate a candidate sampling point pool based on the uncertainty score of the uncertainty map to obtain a ranked candidate pool; select the top N candidate points from the ranked candidate pool to obtain the sampling plan, where N is the sampling budget number.
[0060] Specifically, the implementation process for making uncertainty-based proactive sampling decisions based on the prediction results to determine a sampling plan is as follows: First, the prediction results generated for all grid points across the entire remediation site in the previous phase are obtained, and the confidence value for each grid point is extracted. Using a preset conversion function, such as the inverse of the confidence value, the confidence value for each grid point is converted into its corresponding base uncertainty score. The uncertainty scores for all grid points are visualized or digitized in two dimensions, forming an uncertainty map covering the entire remediation site that intuitively reflects the spatial distribution of the model's prediction reliability. Next, a pool of candidate sampling points is generated and ranked. The system traverses all grid points on the uncertainty map and initially selects those with uncertainty scores above a preset threshold as candidate sampling points. Next, to ensure a representative spatial distribution of sampling points and avoid redundancy, the system applies a non-maximum suppression algorithm based on a minimum spatial distance threshold. This algorithm starts with the point with the highest uncertainty, selects it into the candidate pool, and then suppresses all other candidate points within a preset spatial distance (e.g., the minimum borehole spacing) surrounding it. This process is iterated until all points have been processed. The points in the final pool of candidate sampling points not only have high uncertainty but are also spatially separated. The system then sorts the pool in descending order based on the uncertainty score of each candidate point to form a sorted candidate pool. Finally, based on the preset sampling budget N, that is, the maximum number of points allowed for this sampling, the top N candidate points are selected directly from the top of the sorted candidate pool. These N points with the highest uncertainty scores and that meet the spatial distribution requirements together constitute the final sampling plan. The plan is output in the form of a coordinate list, clearly indicating the specific locations where on-site sampling personnel need to conduct the next physical sampling.
[0061] In summary, the artificial intelligence-based intelligent control system for soil remediation processes according to the embodiments of the present application is illustrated. It constructs an aligned feature matrix that can comprehensively characterize the dynamic process of soil remediation by performing deep preprocessing and feature engineering on multi-source heterogeneous sensor data and control logs. Subsequently, to solve the problem of difficulty in obtaining pollutant concentrations in real time, a soft measurement model based on a sequence codec is adopted to map the processed state sequence into pollutant concentrations and their confidence levels. In addition, in the soft measurement prediction process of pollutant concentrations, a soil state information transmission layer is introduced to dynamically adjust the information flow weight by quantifying and utilizing the uncertainty of time series information, thereby effectively suppressing noise interference, capturing complex time series dependencies, and ensuring the accuracy and robustness of the prediction. Finally, the system does not blindly accept the prediction results, but uses the confidence information in the prediction results to convert them into an uncertainty map, and generates an optimized active sampling plan based on this. This enables limited sampling resources to be accurately deployed to areas that need verification the most and have the highest uncertainty, not only reducing monitoring costs but also achieving closed-loop control from data acquisition, intelligent prediction to optimized decision-making, thereby improving the intelligence level and remediation efficiency of the soil remediation process.
[0062] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. An intelligent control system for soil remediation process based on artificial intelligence, characterized in that: include: Data acquisition module, used to obtain raw sensor data and control logs; Data preprocessing module, used to preprocess the raw sensor data and control logs to obtain the alignment feature matrix; A state encapsulation module is used to perform feature vector serialization and state encapsulation on the alignment feature matrix to obtain a state sequence at a current moment; A pollutant concentration prediction module is used to perform soft measurement prediction of pollutant concentration based on a sequence codec on the state sequence at the current moment to obtain a prediction result, wherein the prediction result includes a predicted pollutant concentration value and its confidence level; an active sampling decision module, configured to make an active sampling decision based on uncertainty based on the prediction result to obtain a sampling plan; The raw sensor data includes sensor ID, timestamp, soil temperature, soil moisture, soil pH, redox potential, conductivity, dissolved oxygen concentration, photoionization detector readings, flame ionization detector readings, and ion selective electrode readings.
2. The soil remediation process intelligent control system based on artificial intelligence according to claim 1 is characterized in that: The data preprocessing module is used to: Normalize and clean the original sensor data and control log to obtain a cleaned sensor data table and a cleaned control log table; Performing data aggregation and alignment based on a unified time grid on the post-cleaning sensor data table and the post-cleaning control log table to obtain unified time grid data, wherein each row of the unified time grid data represents a time window, and each column of the unified time grid data represents the aggregated statistical features of all sensors and control quantities; Multidimensional feature derivation is performed on the unified time grid data to obtain the alignment feature matrix, where each row of the alignment feature matrix represents a time step, and each column of the alignment feature matrix represents the aggregated statistical features and multidimensional derivative features of all sensors and control quantities.
3. The soil remediation process intelligent control system based on artificial intelligence according to claim 1 is characterized in that: The state encapsulation module is used to: Based on the current timestamp, extracting a feature data frame based on a preset time window from the alignment feature matrix to obtain a time data frame; Performing feature value normalization on the time data frame to obtain a normalized data frame; Feature serialization is performed on the normalized data frame to obtain the state sequence at the current moment.
4. The soil remediation process intelligent control system based on artificial intelligence according to claim 3 is characterized in that: The pollutant concentration prediction module includes: a soil state sequence encoding unit, configured to input the state sequence at the current moment into a sequence encoder to obtain a time sequence of soil state feature implicit coding vectors; A soil state information transmission unit, configured to input the time series of the soil state feature implicit coding vector into the soil state information transmission layer to obtain a soil state feature temporal context coding vector; The soil state time series prediction unit is used to input the soil state feature time series context coding vector into the sequence decoder to obtain the prediction result.
5. The soil remediation process intelligent control system based on artificial intelligence according to claim 4 is characterized in that: The soil state information transmission unit includes: A soil state comparison and quantification subunit, configured to quantify soil state uncertainty of each soil state feature implicit coding vector in the time series of the soil state feature implicit coding vector to obtain a sequence distribution of soil state feature temporal entropy gain; A soil state time series weight calculation subunit, configured to determine a sequence distribution of soil state feature time series modulation weights based on a sequence distribution of soil state feature time series entropy gains; A soil state time series feature modulation subunit, configured to perform weighted modulation on the sequence distribution of the soil state feature implicit coding vector based on the sequence distribution of the soil state feature time series modulation weights to obtain a sequence distribution of the modulated soil state feature vector; The soil state temporal context encoding subunit is used to input the sequence distribution of the modulated soil state feature vector into the forward LSTM-based sequence encoder to obtain the soil state feature temporal context encoding vector.
6. The soil remediation process intelligent control system based on artificial intelligence according to claim 5 is characterized in that: The sequence encoder is a trained LSTM model, and the sequence decoder is a fully connected layer.
7. The soil remediation process intelligent control system based on artificial intelligence according to claim 1 is characterized in that: The active sampling decision module is used to: Convert the confidence level in the prediction results of each grid point into the basic uncertainty to obtain the uncertainty map; Based on a minimum spatial distance threshold, generating a candidate sampling point pool based on uncertainty scores on the uncertainty map to obtain a ranked candidate pool; The first N candidate points are selected from the sorted candidate pool to obtain the sampling plan, where N is the sampling budget.
Citation Information
Cited By
Closed-loop optimization method for carbon cycle regulation and control of soil organic matters and collaborative removal of pollutants
CN122172593A